【问题标题】:Python Lex-Yacc(PLY): Not recognizing start of line or start of stringPython Lex-Yacc(PLY):无法识别行开头或字符串开头
【发布时间】:2014-05-29 04:44:31
【问题描述】:

我对@9​​87654321@ 还很陌生,而且对 Python 的了解还不止于此。我正在尝试使用PLY-3.4 和python 2.7 来学习它。请看下面的代码。我正在尝试创建一个令牌 QTAG,它是一个由零个更多空格组成的字符串,后跟“Q”或“q”,然后是“。”和一个正整数和一个或多个空格。例如 VALID QTAG 是

"Q.11 "
"  Q.12 "
"q.13     "
'''
   Q.14 
'''

无效的是

"asdf Q.15 "
"Q.  15 "

这是我的代码:

import ply.lex as lex

class LqbLexer:
     # List of token names.   This is always required
     tokens =  [
        'QTAG',
        'INT'
        ]


     # Regular expression rules for simple tokens

    def t_QTAG(self,t):
        r'^[ \t]*[Qq]\.[0-9]+\s+'
        t.value = int(t.value.strip()[2:])
        return t

    # A regular expression rule with some action code
    # Note addition of self parameter since we're in a class
    def t_INT(self,t):
    r'\d+'
    t.value = int(t.value)   
    return t


    # Define a rule so we can track line numbers
    def t_newline(self,t):
        r'\n+'
        print "Newline found"
        t.lexer.lineno += len(t.value)

    # A string containing ignored characters (spaces and tabs)
    t_ignore  = ' \t'

    # Error handling rule
    def t_error(self,t):
        print "Illegal character '%s'" % t.value[0]
        t.lexer.skip(1)

    # Build the lexer
    def build(self,**kwargs):
        self.lexer = lex.lex(debug=1,module=self, **kwargs)

    # Test its output
    def test(self,data):
        self.lexer.input(data)
        while True:
             tok = self.lexer.token()
             if not tok: break
             print tok

# test it
q = LqbLexer()
q.build()
#VALID inputs
q.test("Q.11 ")
q.test("  Q.12 ")
q.test("q.13     ")
q.test('''
   Q.14 
''')
# INVALID ones are
q.test("asdf Q.15 ")
q.test("Q.  15 ")

我得到的输出如下:

LexToken(QTAG,11,1,0)
Illegal character 'Q'
Illegal character '.'
LexToken(INT,12,1,4)
LexToken(QTAG,13,1,0)
Newline found
Illegal character 'Q'
Illegal character '.'
LexToken(INT,14,2,6)
Newline found
Illegal character 'a'
Illegal character 's'
Illegal character 'd'
Illegal character 'f'
Illegal character 'Q'
Illegal character '.'
LexToken(INT,15,3,7)
Illegal character 'Q'
Illegal character '.'
LexToken(INT,15,3,4)

请注意,只有第一个和第三个有效输入被正确标记。我无法弄清楚为什么我的其他有效输入没有被正确标记。在 t_QTAG 的文档字符串中:

  1. 将'^' 替换为'\A' 无效。
  2. 我尝试删除 '^' 。然后所有有效输入都被标记化,然后是第二个 无效输入也会被标记化。

提前感谢任何帮助!

谢谢

PS:我加入了 google-group ply-hack 并尝试在那里发帖,但我无法直接在论坛或通过电子邮件发帖。我不确定该组是否已处于活动状态。 Beazley 教授也没有回应。有任何想法吗?

【问题讨论】:

  • 嗯,你明确指出应该通过t_ignore 标志忽略空格,但是你需要在你的正则表达式中使用它们。我确实相信您的第二个“错误输入”得到验证的原因是因为 t_ignore 正在消耗内部空格,使其看起来像 Q.15 您可以尝试将 T_QTAG 正则表达式设置为此吗? [Qq]\.[0-9]+ 另外你能检查一下当你从 t_ignore 中删除空格时会发生什么吗?
  • 我应该提到我从未使用过 PLY,但过去使用过 JFLEX/CUP。另外,为什么在 [ \t] 和 \s 之间跳转。请注意, \s 涵盖: [ \t\n\r\f\v] 这意味着此正则表达式将使用换行符,而不会让换行符检查器找到它们
  • 您还说followed by '.' and a positive integer and zero or more spaces or tab,而您的正则表达式最后使用\s+,表明您最后需要至少一个空格。不知道是不是故意的。
  • @Tadgh Re.您的第一条评论:感谢t_ignore 占用空间。是真的。 @Tadgh Re。您的第二条评论:您是对的。开头可能只是\s。 @Tadgh Re。您的最后一条评论:正则表达式是正确的,但我的措辞是错误的!它应该是:“......后跟'。'以及一个正整数和一个或多个空格,即 [ \t\n\r\f\v]。”通过阅读一些文档,我找到了正确的答案并发布在下面。感谢@Tadgh 的指点!

标签: python regex tokenize lexer ply


【解决方案1】:

最后我自己找到了答案。发布它,以便其他人可能会发现它有用。

正如@Tadgh 正确指出的那样,t_ignore = ' \t' 消耗了空格和制表符,因此我将无法按照上面的正则表达式匹配t_QTAG,结果是第二个有效输入没有被标记化。通过仔细阅读 PLY 文档,我了解到如果要维护令牌的正则表达式的顺序,那么它们必须在函数中定义,而不是像 t_ignore 那样在字符串中定义。如果使用字符串,则 PLY 会自动按从最长到最短的长度对它们进行排序,并将它们附加到函数之后。我猜这里t_ignore 很特别,它以某种方式在其他任何事情之前执行。这部分没有明确记录。解决此问题的方法是使用新标记定义函数,例如 t_SPACETAB、after t_QTAG,并且不返回任何内容。有了这个,所有 valid 输入现在都被正确标记,除了带有三引号的输入(包含"Q.14" 的多行字符串)。此外,根据规范,无效的未标记化。

多行字符串问题:原来PLY内部使用re模块。在该模块中,^ 仅在 string 的开头被解释,而不是在 每一行 的开头,默认情况下。要改变这种行为,我需要打开多行标志,这可以使用(?m) 在正则表达式中完成。因此,要正确处理我的测试中的所有有效和无效字符串,正确的正则表达式是:

r'(?m)^\s*[Qq]\.[0-9]+\s+'

这是添加了更多测试的更正代码:

import ply.lex as lex

class LqbLexer:
    # List of token names.   This is always required

    tokens = [
        'QTAG',
        'INT',
        'SPACETAB'
        ]


    # Regular expression rules for simple tokens

    def t_QTAG(self,t):
        # corrected regex
        r'(?m)^\s*[Qq]\.[0-9]+\s+'
        t.value = int(t.value.strip()[2:])
        return t

    # A regular expression rule with some action code
    # Note addition of self parameter since we're in a class
    def t_INT(self,t):
        r'\d+'
        t.value = int(t.value)    
        return t

    # Define a rule so we can track line numbers
    def t_newline(self,t):
        r'\n+'
        print "Newline found"
        t.lexer.lineno += len(t.value)

    # A string containing ignored characters (spaces and tabs)
    # Instead of t_ignore  = ' \t'
    def t_SPACETAB(self,t):
        r'[ \t]+'
        print "Space(s) and/or tab(s)"

    # Error handling rule
    def t_error(self,t):
        print "Illegal character '%s'" % t.value[0]
        t.lexer.skip(1)

    # Build the lexer
    def build(self,**kwargs):
        self.lexer = lex.lex(debug=1,module=self, **kwargs)

    # Test its output
    def test(self,data):
        self.lexer.input(data)
        while True:
             tok = self.lexer.token()
             if not tok: break
             print tok

# test it
q = LqbLexer()
q.build()
print "-============Testing some VALID inputs===========-"
q.test("Q.11 ")
q.test("  Q.12 ")
q.test("q.13     ")
q.test("""


   Q.14
""")
q.test("""

qewr
dhdhg
dfhg
   Q.15 asda

""")

# INVALID ones are
print "-============Testing some INVALID inputs===========-"
q.test("asdf Q.16 ")
q.test("Q.  17 ")

这是输出:

-============Testing some VALID inputs===========-
LexToken(QTAG,11,1,0)
LexToken(QTAG,12,1,0)
LexToken(QTAG,13,1,0)
LexToken(QTAG,14,1,0)
Newline found
Illegal character 'q'
Illegal character 'e'
Illegal character 'w'
Illegal character 'r'
Newline found
Illegal character 'd'
Illegal character 'h'
Illegal character 'd'
Illegal character 'h'
Illegal character 'g'
Newline found
Illegal character 'd'
Illegal character 'f'
Illegal character 'h'
Illegal character 'g'
Newline found
LexToken(QTAG,15,6,18)
Illegal character 'a'
Illegal character 's'
Illegal character 'd'
Illegal character 'a'
Newline found
-============Testing some INVALID inputs===========-
Illegal character 'a'
Illegal character 's'
Illegal character 'd'
Illegal character 'f'
Space(s) and/or tab(s)
Illegal character 'Q'
Illegal character '.'
LexToken(INT,16,8,7)
Space(s) and/or tab(s)
Illegal character 'Q'
Illegal character '.'
Space(s) and/or tab(s)
LexToken(INT,17,8,4)
Space(s) and/or tab(s)

【讨论】:

  • "^ 仅在字符串的开头解释,而不是每一行的开头" - 这让我很困惑!很好地解决了这个问题!
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-02-11
  • 2012-09-23
  • 2018-02-24
  • 2012-09-06
  • 2021-12-23
  • 1970-01-01
相关资源
最近更新 更多