【发布时间】:2019-01-12 14:31:11
【问题描述】:
我正在尝试理解 re.split() 函数与非捕获组来拆分逗号分隔的字符串。
这是我的代码:
pattern = re.compile(r',(?=(?:"[^"]*")*[^"]*$)')
text = 'qarcac,"this is, test1",123566'
results= re.split(pattern, text)
for r in results:
print(r.strip())
当我执行此代码时,结果与预期一致。
split1:qarcac
split2:“这是,test1”
split3:123566
而如果我在源文本中再添加一个双引号字符串,它就不会按预期工作。
text = 'qarcac,"this is, test1","this is, test2", 123566, testdata'
并产生以下输出
split1: qarcac,"这是,test1"
split2:“这是,test2”
split3:123566
谁能解释一下这里发生了什么以及在这两种情况下非捕获组的工作方式有何不同?
【问题讨论】:
-
您应该使用
csv模块来解析CSV 字符串。您使用的正则表达式效率非常低,如果字符串很长,性能可能会大幅下降。 -
感谢 Wiktor,我不打算将其生产化,而是尝试学习,因为我在我的一个学习模块中遇到了这段代码。
-
有效的模式是
,(?=(?:"[^"]*"|[^"])*$)。或,(?=[^"]*(?:"[^"]*"[^"]*)*$)。见Regex to pick commas outside of quotes。 -
见regex101.com/r/dRqJZT/1,你在右边的模式字段中输入的任何正则表达式都有很好的解释。
-
感谢 Wiktor.. 当使用 [^"]* 时,re.split() 如何使用以下正则表达式标记源字符串中逗号的第一次出现..... ,(?= (?:"[^"]*"|[^"])*$)
标签: python regex python-3.x regex-lookarounds