【问题标题】:Python String split by specific pattern with IndicesPython字符串按具有索引的特定模式拆分
【发布时间】:2020-10-31 17:14:31
【问题描述】:

我试图从不同的字符中拆分句子,每个单词都有自己的标签,并用索引存储,名字可以是不同长度的 Mike 或 Steve。内容可以是中文、日文等多种语言。

content = "A:Hello.B:How are you?A:I'm fine."

我想成为什么样的人:

[0]A:Hello.       , 0:7
[1]B:How are you? , 8:21
[2]A:I'm fine.    ,22:33

【问题讨论】:

    标签: python split tokenize python-re


    【解决方案1】:

    您可以使用re.split,如下:

    import re
    s = "A:Hello.B:How are you?A:I'm fine."
    t = re.split(r'[.?]', s)
    print(t)
    

    给了

    ['A:Hello', 'B:How are you', "A:I'm fine", '']
    

    【讨论】:

      【解决方案2】:

      您可以使用re.finditer 执行任务:

      import re
      
      content = "A:Hello.B:How are you?A:I'm fine."
      
      for idx, i in enumerate(re.finditer(r'(.*?[.?])(?=[A-Z]|\Z)', content)):
          print('[{}]{:<20}, {}:{}'.format(idx, i.group(1), i.start(), i.end()-1))
      

      打印:

      [0]A:Hello.            , 0:7
      [1]B:How are you?      , 8:21
      [2]A:I'm fine.         , 22:32
      

      【讨论】:

      • 比我的好多了!
      猜你喜欢
      • 2018-10-23
      • 2013-04-11
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2015-12-18
      • 2019-05-02
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多