【问题标题】:How do I split a string into a list of words?如何将字符串拆分为单词列表?
【发布时间】:2023-01-22 13:47:25
【问题描述】:

如何拆分句子并将每个单词存储在列表中?例如,给定一个像"these are words" 这样的字符串,我如何得到一个像["these", "are", "words"] 这样的列表?

【问题讨论】:

  • 实际上,您将为列表中的每个单词打印完整的单词列表。我认为您打算使用 print(word) 作为最后一行。
  • 请参阅stackoverflow.com/questions/4978787 将字符串拆分为单个字符。

标签: python list split text-segmentation


【解决方案1】:

给定一个字符串 sentence,这会将每个单词存储在一个名为 words 的列表中:

words = sentence.split()

【讨论】:

    【解决方案2】:

    要在任何连续运行的空格上拆分字符串 text

    words = text.split()      
    

    要在自定义分隔符(例如 ",")上拆分字符串 text

    words = text.split(",")   
    

    words 变量将是 list 并包含来自 text 分隔符的单词。

    【讨论】:

      【解决方案3】:

      使用str.split()

      返回一个单词列表在字符串中,使用 sep 作为分隔符 ...如果未指定 sep 或为 None,则应用不同的拆分算法:连续的空格被视为单个分隔符,如果字符串有前导或尾随,结果将在开头或结尾不包含空字符串空格。

      >>> line = "a sentence with a few words"
      >>> line.split()
      ['a', 'sentence', 'with', 'a', 'few', 'words']
      

      【讨论】:

      【解决方案4】:

      根据您打算如何处理句子列表,您可能需要查看 Natural Language Took Kit。它主要处理文本处理和评估。您也可以使用它来解决您的问题:

      import nltk
      words = nltk.word_tokenize(raw_sentence)
      

      这具有拆分标点符号的额外好处。

      例子:

      >>> import nltk
      >>> s = "The fox's foot grazed the sleeping dog, waking it."
      >>> words = nltk.word_tokenize(s)
      >>> words
      ['The', 'fox', "'s", 'foot', 'grazed', 'the', 'sleeping', 'dog', ',', 
      'waking', 'it', '.']
      

      这使您可以过滤掉任何不需要的标点符号,只使用单词。

      请注意,如果您不打算对句子进行任何复杂的操作,那么使用string.split() 的其他解决方案会更好。

      [编辑]

      【讨论】:

      • split() 依赖于空格作为分隔符,因此它无法分隔带连字符的单词——长破折号分隔的短语也无法分隔。如果句子中包含任何没有空格的标点符号,这些标点符号将无法粘贴。对于任何真实世界的文本解析(如此评论),您的 nltk 建议比 split()` 好得多。
      • 可能有用,尽管我不会将其描述为拆分为“单词”。根据任何简单的英语定义,','"'s" 都不是单词。通常,如果您想以标点符号感知的方式将上面的句子拆分为“单词”,您会想要去掉逗号并将 "fox's" 作为单个单词。
      • 截至 2016 年 4 月,Python 2.7+。
      【解决方案5】:

      这个算法怎么样?在空白处拆分文本,然后修剪标点符号。这会小心地去除单词边缘的标点符号,而不会损坏单词中的撇号,例如 we're

      >>> text
      "'Oh, you can't help that,' said the Cat: 'we're all mad here. I'm mad. You're mad.'"
      
      >>> text.split()
      ["'Oh,", 'you', "can't", 'help', "that,'", 'said', 'the', 'Cat:', "'we're", 'all', 'mad', 'here.', "I'm", 'mad.', "You're", "mad.'"]
      
      >>> import string
      >>> [word.strip(string.punctuation) for word in text.split()]
      ['Oh', 'you', "can't", 'help', 'that', 'said', 'the', 'Cat', "we're", 'all', 'mad', 'here', "I'm", 'mad', "You're", 'mad']
      

      【讨论】:

      • 不错,但有些英语单词确实包含尾随标点符号。例如,e.g.Mrs. 中的尾随点,以及所有格 frogs' 中的尾随撇号(如 frogs' legs)是单词的一部分,但会被该算法去除。正确处理缩写可以是大致通过检测以点分隔的首字母缩写加上使用特殊情况字典(如Mr.Mrs.)来实现。区分所有格撇号和单引号要困难得多,因为它需要分析包含该词的句子的语法。
      • @MarkAmery 你是对的。从那以后我还想到,一些标点符号——例如破折号——可以在没有空格的情况下分隔单词。
      【解决方案6】:

      我希望我的 python 函数拆分一个句子(输入)并将每个单词存储在一个列表中

      str().split() 方法就是这样做的,它接受一个字符串,将它拆分成一个列表:

      >>> the_string = "this is a sentence"
      >>> words = the_string.split(" ")
      >>> print(words)
      ['this', 'is', 'a', 'sentence']
      >>> type(words)
      <type 'list'> # or <class 'list'> in Python 3.0
      

      【讨论】:

        【解决方案7】:

        如果你想要一个的所有字符单词/句子在列表中,执行此操作:

        print(list("word"))
        #  ['w', 'o', 'r', 'd']
        
        
        print(list("some sentence"))
        #  ['s', 'o', 'm', 'e', ' ', 's', 'e', 'n', 't', 'e', 'n', 'c', 'e']
        

        【讨论】:

        【解决方案8】:

        shlex 有一个 .split() 功能。它与 str.split() 的不同之处在于它不保留引号并将引用的短语视为单个单词:

        >>> import shlex
        >>> shlex.split("sudo echo 'foo && bar'")
        ['sudo', 'echo', 'foo && bar']
        

        注意:它适用于类 Unix 命令行字符串。它不适用于自然语言处理。

        【讨论】:

        • 谨慎使用,尤其是对于 NLP。它会在单引号字符串上崩溃,例如 "It's good."ValueError: No closing quotation
        【解决方案9】:

        拆分单词而不破坏单词中的撇号 请找出 input_1 和 input_2 摩尔定律

        def split_into_words(line):
            import re
            word_regex_improved = r"(w[w']*w|w)"
            word_matcher = re.compile(word_regex_improved)
            return word_matcher.findall(line)
        
        #Example 1
        
        input_1 = "computational power (see Moore's law) and "
        split_into_words(input_1)
        
        # output 
        ['computational', 'power', 'see', "Moore's", 'law', 'and']
        
        #Example 2
        
        input_2 = """Oh, you can't help that,' said the Cat: 'we're all mad here. I'm mad. You're mad."""
        
        split_into_words(input_2)
        #output
        ['Oh',
         'you',
         "can't",
         'help',
         'that',
         'said',
         'the',
         'Cat',
         "we're",
         'all',
         'mad',
         'here',
         "I'm",
         'mad',
         "You're",
         'mad']
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2011-11-03
          • 2018-03-13
          • 2018-04-06
          • 2011-06-12
          • 2017-08-16
          • 2011-04-22
          相关资源
          最近更新 更多