【问题标题】:Implementing a C like program in racket using lex/yacc使用 lex/yacc 在球拍中实现类似 C 的程序
【发布时间】:2020-04-22 08:38:45
【问题描述】:

我正在查看这两个资源(https://github.com/racket/parser-tools/blob/master/parser-tools-lib/parser-tools/examples/calc.rkt 和 https://gist.github.com/gcr/1318240),虽然我还不完全理解主 calc 函数的工作原理,但我想知道是否可以将其扩展为适用于类似 c 的简单程序只是没有功能?因此它将对 if、while 和 print 语句进行 lex、解析和评估。所以像(define-empty-tokens op-tokens ( newline = OC CC (open-curly/closed-curly for block statements) DEL PRINT WHILE (WHILE exp S) S IF S1 S2 (IF exp S1 S2) OP CP + - * / || % or && == != >= <= > < EOF ))

到目前为止,我将它(第一个链接的代码)扩展为也可以使用布尔值:

所以在 calcl 中我添加了这两行:

[ (:= 2  #\|)   (token-||)]
[(:or "=" "+" "-" "*" "/" "%" "&&"      "==" "!=" ">=" "<=" ">" "<") (string->symbol lexeme)]

然后:

(define calcp
  (parser

   (start  start)
   (end newline EOF)
   (tokens value-tokens op-tokens)
   (error (lambda (a b c) (void)))

   (precs (right =)


          (left  ||)
          (left &&)
          (left == !=)
          (left <= >= < >)
          (left - +)
          (left * / %)

         )

   (grammar

    (start [() #f]
           [(error start) $2]
           [(exp) $1])

    (exp [(NUM) $1]
         [(VAR) (hash-ref vars $1 (lambda () 0))]
         [(VAR = exp) (begin (hash-set! vars $1 $3)
                             $3)]
         [(exp || exp) (if  (not(and (equal? $1 0) (equal? $3 0) ))  1 0) ]  
         [(exp && exp) (and $1 $3)]
         [(exp == exp) (equal? $1 $3)]
         [(exp != exp) (not(equal? $1 $3))]
         [(exp < exp) (< $1 $3)]
         [(exp > exp) (> $1 $3)]
         [(exp >= exp) (>= $1 $3)]
         [(exp <= exp) (<= $1 $3)]


         [(exp + exp) (+ $1 $3)]
         [(exp - exp) (- $1 $3)]
         [(exp * exp) (* $1 $3)]
         [(exp / exp) (/ $1 $3)]
         [(exp % exp) (remainder $1 $3)]


         [(OP exp CP) $2]))))

但我很难理解上面的代码以及下面的代码。如果可能的话,我会更改它以便它也适用于 ifs 和 whiles 等?

(define (calc ip)
   (port-count-lines! ip)
  (letrec ((one-line
        (lambda ()
              (let ((result (calcp (lambda () (calcl ip))  )))
                (when result (printf "~a\n" result)  (one-line))
                )
                 ) ))
    (one-line))
  )

另外,这家伙似乎依赖换行符来标记语句的结尾。即你不能在一行上有超过 1 个语句。我希望程序能够识别一行中的两个语句,并通过某种方式向前看并检查是否有新的未声明变量、特殊关键字或开/闭括号等来分别评估它们。

更新:

我通过以下规则设法为 arith 表达式构建了一个 AST,但我如何摆脱除重要括号之外的所有内容,以便我可以评估它?

例如:输入列表:(list (token 'NUM 17) '+ (token 'NUM 1) '* (token 'NUM 3) '/ 'OP (token 'NUM 6) '- (token 'NUM 5) 'CP)

我回来了:

'(exp (((((factor 17)))) + (((((factor 1))) * ((factor 3))) / ((((((factor 6)))) - (((factor 5))))))))

这是我的规则:

exp : add
/add : add ('+' mul)+  | add ('-'  mul)+ | mul  
/mul : mul ('*' atom)+  | mul ('/'  atom)+ | mul ('%'   atom)+ | atom
/atom :  /OP add /CP | factor
factor :  NUM | ID

【问题讨论】:

  • 更新:这是我目前所理解的: 1 - 单行是一个声明的函数,当被调用时执行一个 lambda: 2 - 为结果分配一个值,它将从中接收一个值计算。 3 - calcp 将通过执行它从一行接收到的 lambda 来产生这个值。 4 - 一个 lambda,它在执行时调用 calcl 从词法分析器传递它的“一个”值。 5 - 这样该值将被转发到 calcp。
  • 6 - calcp 将做它的事情并传回设置为结果的值。 7 - 然后打印结果 8 - 再次递归调用一行 9 - 在函数结束时,调用单行以使事情在开始时正确运行 所以在 5) 是我所在的位置困惑,当 calcp 只从 calcl 收到“一个”标记时,它如何设法评估整个表达式?
  • calcp 传递了一个函数,每次调用它都会返回一个令牌。所以calcp 可以使用它来获取它认为必要的尽可能多的代币。
  • 您可以通过更改语法中的归约操作让calcp 做任何您想做的事情。现在,这些动作评估并返回一个数字。要制作 AST,这些操作需要构建并返回一个 AST 节点(如果这是它们返回的,这就是它们将在 $1 中恢复的,...)
  • 您必须查看calcp 的定义才能看到它在哪里调用其参数,并且可能不容易看到定义,因为parser 可能会生成函数.你真的应该阅读 SICP :-)

标签: parsing racket interpreter yacc lexer


【解决方案1】:

您无法使用基于即时评估的评估器轻松实现具有条件和循环结构的语言。

至少对于循环来说应该很清楚。如果你有类似的东西(使用超级简化的语法):

repeat 3 { i = i + 1 }

如果您在解析过程中求值,i = i + 1 将只求值一次,因为字符串只解析一次。为了使其被多次评估,解析器需要将i = i + 1 转换为可以在评估repeat 时多次评估的东西。

这东西通常是一个抽象语法树(AST),或者可能是一个虚拟机操作的列表。使用 Scheme,您还可以将被解析的表达式转换为函数。

所有这些都是完全实用的,甚至不是特别困难,但是您确实需要准备阅读一些关于解析和生成可执行文件的内容。对于后者,我强烈推荐经典的Structure and Interpretation of Computer Programs (Abelson & Sussman)。

【讨论】:

  • 所以 yacc 不能先进行词法分析再单独解析?上面的代码是在前面的行中使用先前评估的变量来表示表达式。他通过实现某种记忆来做到这一点。我之前在序言中使用 Backus Naur 和构建 AST 完成了一个解释器,但我没有在网上找到任何资源来展示如何在 Racket 中执行此操作(我需要为作业执行此操作)。在 Racket 中,我能找到的所有内置支持都是 yacc,所以我很难相信上面代码中使用的这个库不能扩展到表达式评估之外。
  • 要实现您所描述的循环,我需要做的只是一些递归,当我处于解析的那个阶段时,并在检查 while 条件的同时不断更新该变量,直到例如,该变量的存储值没有通过条件,不是吗?
  • @nkatz:您可以使用自上而下的解析器来做到这一点,但这并不容易。但是,您有一个自底向上的解析器,它首先处理组件,因此您不知道 i = i + 1 在减少时是条件或循环的一部分。
  • @nkatz:如果您知道如何构建 AST,那么您当然可以在球拍中做到这一点。 AST 可以像 Scheme 程序本身一样简单,(node-type child ...)。评估 AST 只是自上而下的树遍历。
  • 哦,我明白你在说什么。为了能够重新评估,您需要保持表达式活动/扩展,并且上面的代码立即评估/折叠它?!
【解决方案2】:

根据我看到的here 和here,我设法为一个简单的类 c 程序实现了解释器的 lex 和 parse 阶段。它没有函数,但它有赋值、变量、条件、循环和打印语句(除了算术,它还包含逻辑表达式。)。我把它贴在下面,以防其他人觉得它有用(包括评估阶段和输入样本在内的整个事情是here):

(require parser-tools/yacc //provides you with the lexer, the parser and the lexeme tools eg. string-> symbol, string->number etc - In general, with the ability to map the literals in the input  
         parser-tools/lex
         (prefix-in : parser-tools/lex-sre))

(define-tokens value-tokens (NUM VAR  ))
(define-empty-tokens op-tokens ( newline  = OC CC DEL OP CP + - * / || %   or && == != >= <= > <  EOF PRINT WHILE IF ELSE  ))



(define vars (make-hash)) ;to store the values in the variables 

(define-lex-abbrevs
  (lower-letter (:/ "a" "z"))
  (upper-letter (:/ #\A #\Z))
  (digit (:/ "0" "9")))
 
(define calcl ;lexer is mapping the literals to their tokens or values
  (lexer
   [(eof) 'EOF]
   [(:or #\tab #\space  #\return #\newline  ) (calcl input-port)]
   [ (:= 2  #\|)   (token-||)]
   [(:or "=" "+" "-" "*" "/" "%" "&&"      "==" "!=" ">=" "<=" ">" "<") (string->symbol lexeme)]
   ["(" 'OP]
   [")" 'CP]
   ["{" 'OC]
   ["}" 'CC]
   [ "print" 'PRINT   ]
   [#\,  'DEL  ]
   [ "while" 'WHILE ]
   [ "if" 'IF ]
   [ "else" 'ELSE ] 
   [(:+ (:or lower-letter upper-letter)) (token-VAR (string->symbol lexeme))]
   [(:+ digit) (token-NUM (string->number lexeme))]
   ))


(define calcp ;defines how to parse the program and how the program structure is made up (recursively)
  (parser
   (start   start);refers to the block named 'start' below. Every parser has a start, end, tokens definition, optional error message when needed and operator precedence def.
   (end   EOF)
   (tokens value-tokens op-tokens  )
   (error (λ(ok? name value) (if (boolean? value) (printf "Couldn't parse: ~a\n" name) (printf "Couldn't parse: ~a\n" value)))) 
   (precs   ;sets the precedence of the operators in relation to each other - that is, to which operand they bind stronger
    (left DEL);from lowest to highest
    (right =)             
    (left  ||)
    (left &&)
    (left == !=)
    (left <= >= < >)
    (left - +)
    (left * / %)
    (right OP)
    (left CP)
    (right OC)
    (left CC)
    )
   
   (grammar    ;what is the grammar of my program?
    (start
     [() '()]  ; returns empty list when it matches onto nothing              
     [(statements) `(,$1)]
     [(statements start) `(,$1,$2)] ;we can have more than one statement - one example of the recursiveness 
     )

    (statements /what type of major statements we might have
                
     [(var = exp ) `(assign ,$1 ,$3)]
     [(IF ifState) $2]
     [(WHILE while) $2]
     [(PRINT printVals)  `(print ,$2)]
               
                 
                
              
     )
 
    (ifState ; It's assumed that you cannot have an if inside an if unless it's within curly braces
  
     [(OP exp CP statements) `(if ,$2 ,$4)] /combinations of different major statements
     [(OP exp CP block) `(if ,$2 ,$4)]       
     [(OP exp CP block ELSE statements ) `(if ,$2 ,$4 ,$6)]
     [(OP exp CP statements ELSE block ) `(if ,$2 ,$4 ,$6)]
     [(OP exp CP statements ELSE statements ) `(if ,$2 ,$4 ,$6)]       
     [(OP exp CP block ELSE block ) `(if ,$2 ,$4 ,$6)]
     )
    (while
     [(OP exp CP block) `(while ,$2, $4)]
     )

    (block
     [(OC start CC) $2] ;we can statements or entire program wrapped into curly braces - a block 
     )
   
     
    (var
     [(VAR) $1]
     )
     
    (printVals
          
     [(exp DEL printVals ) `(,$1 ,$3)]
     [(exp) $1]          
         
         
     )
     
        
    (exp [(NUM)  $1] ; smallest (most reducible) chunk in an expression when the chunk is an integer
         [(VAR)  $1] ; smallest (most reducible) chunk in an expression when the chunk is a variable
         [(exp || exp) `((lambda (a b) (or a b))  ,$1 ,$3) ]  
         [(exp && exp) `((lambda (a b) (and a b)) ,$1 ,$3)]
         [(exp == exp) `(equal? ,$1 ,$3)]
         [(exp != exp) `(not(equal? ,$1 ,$3))]
         [(exp < exp) `(< ,$1 ,$3)]
         [(exp > exp) `(> ,$1 ,$3)]
         [(exp >= exp) `(>= ,$1 ,$3)]
         [(exp <= exp) `(<= ,$1 ,$3)]
         [(exp + exp) `(+ ,$1 ,$3)]
         [(exp - exp) `(- ,$1 ,$3)]
         [(exp * exp) `(* ,$1 ,$3)]
         [(exp / exp) `(quotient ,$1 ,$3)]
         [(exp % exp) `(modulo ,$1 ,$3)]
         [(OP exp CP) $2]) ;when the expressions are wrapped parentheses
    )
  
   )
  )

Procedure 调用解析器,将其传递给它一个 lambda(调用 lexer 为其提供输入),它将用于从 lexer 获取值,一次它认为合适的值。 ;即根据上面的解析规则,规定

(define (calceval ip)
  (calcp (lambda () (calcl ip))))

要查看评估者(遗憾的是还没有多少 cmets)请参阅 here

【讨论】:

    猜你喜欢
    • 2010-10-07
    • 2013-10-17
    • 2011-11-05
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多