【问题标题】:Tokenize words in a list of sentences Python标记句子列表中的单词 Python
【发布时间】:2014-02-17 02:58:00
【问题描述】:

我目前有一个文件,其中包含一个看起来像

的列表
example = ['Mary had a little lamb' , 
           'Jack went up the hill' , 
           'Jill followed suit' ,    
           'i woke up suddenly' ,
           'it was a really bad dream...']

“example”是这样的句子列表,我希望输出看起来像:

mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill' ....] 等等。 我需要将句子与标记化的每个单词分开,以便我可以将mod_example(一次使用for循环)句子中的每个单词与参考句子进行比较。

我试过这个:

for sentence in example:
    text3 = sentence.split()
    print text3 

并得到以下输出:

['it', 'was', 'a', 'really', 'bad', 'dream...']

如何获得所有句子的信息? 它不断覆盖。是的,还提到我的方法是否正确? 这应该是一个带有词标记的句子列表。谢谢

【问题讨论】:

  • 您能否更彻底地解释一下“..”的含义,以便我可以将 mod_example 句子中的每个单词(一次使用 for 循环)与参考句子进行比较。
  • " 意味着每个句子仍然是一个单独的实体。所以我希望单词被标记,而不是整个文本。例如:我不想要 ['mary' 'had' 'a' ' little' 'lamb' jack' 'went' 'up' 'the' 'hill'] 等等。它仍然应该是一个列表,每个句子都有标记词..

标签: python-2.7 text nltk


【解决方案1】:

使用列表理解来访问您的句子,然后对其进行 word_tokenize。

 from nltk import word_tokenize
 sentences = ['Mary had a little lamb' , 
        'Jack went up the hill' , 
        'Jill followed suit' ,    
        'i woke up suddenly' ,
        'it was a really bad dream...']

 sentences = [ word_tokenize ( sent ) for sent in sentences ]

 print(sentences)

输出:

 [['Mary', 'had', 'a', 'little', 'lamb'], ['Jack', 'went', 'up', 'the', 'hill'], ['Jill', 'followed', 'suit'], ['i', 'woke', 'up', 'suddenly'], ['it', 'was', 'a', 'really', 'bad', 'dream', '...']]

【讨论】:

    【解决方案2】:

    这也可以通过pytorchtorchtext as 来完成

    from torchtext.data import get_tokenizer
    
    tokenizer = get_tokenizer('basic_english')
    example = ['Mary had a little lamb' , 
                'Jack went up the hill' , 
                'Jill followed suit' ,    
                'i woke up suddenly' ,
                'it was a really bad dream...']
    tokens = []
    for s in example:
        tokens += tokenizer(s)
    # ['mary', 'had', 'a', 'little', 'lamb', 'jack', 'went', 'up', 'the', 'hill', 'jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream', '.', '.', '.']
    

    【讨论】:

      【解决方案3】:

      在 Spacy 中,它会很简单:

      import spacy
      
      example = ['Mary had a little lamb' , 
                 'Jack went up the hill' , 
                 'Jill followed suit' ,    
                 'i woke up suddenly' ,
                 'it was a really bad dream...']
      
      nlp = spacy.load("en_core_web_sm")
      
      result = []
      
      for line in example:
          sent = nlp(line)
          token_result = []
          for token in sent:
              token_result.append(token)
          result.append(token_result)
      
      print(result)
      

      输出将是:

      [[Mary, had, a, little, lamb], [Jack, went, up, the, hill], [Jill, followed, suit], [i, woke, up, suddenly], [it, was, a, really, bad, dream, ...]]
      

      【讨论】:

        【解决方案4】:

        分解列表“示例”

        first_split = []
        
        for i in example:
        
            first_split.append(i.split())
        

        分解first_split列表的元素

        second_split = []
        
        for j in first_split:
        
            for k in j:
        
                second_split.append(k.split())
        

        分解 second_split 列表的元素并将其附加到最终列表中,编码器如何需要输出

        final_list = []
        
        for m in second_split:
        
            for n in m:
        
                if(n not in final_list):
        
                    final_list.append(n)
        
        print(final_list)   
        

        【讨论】:

        • 我希望这是最简单的方法。
        • 试一试。
        【解决方案5】:

        我制作这个脚本是为了让所有人都了解如何标记化,这样他们就可以自己构建自然语言处理的引擎。

        import re
        from contextlib import redirect_stdout
        from io import StringIO
        
        example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
        
        def token_to_sentence(str):
            f = StringIO()
            with redirect_stdout(f):
                regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
                regex_of_sentence = [x for x in regex_of_sentence if x is not '']
                for i in regex_of_sentence:
                    print(i)
                first_step_to_sentence = (f.getvalue()).split('\n')
            g = StringIO()
            with redirect_stdout(g):
                for i in first_step_to_sentence:
                    try:
                        regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
                        print(regex_to_clear_sentence.group(1))
                    except:
                        print(i)
                sentence = (g.getvalue()).split('\n')
            return sentence
        
        def token_to_words(str):
            f = StringIO()
            with redirect_stdout(f):
                for i in str:
                    regex_of_word = re.findall('([\w]{0,})', i)
                    regex_of_word = [x for x in regex_of_word if x is not '']
                    for word in regex_of_word:
                        print(regex_of_word)
                words = (f.getvalue()).split('\n')
        

        我做了一个不同的过程,我从段落重新开始这个过程,让大家更了解文字处理。要处理的段落是:

        example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
        

        将段落标记为句子:

        sentence = token_to_sentence(example)
        

        结果:

        ['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
        

        标记为单词:

        words = token_to_words(sentence)
        

        结果:

        ['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
        

        我将解释这是如何工作的。

        首先,我使用正则表达式搜索所有分隔单词的单词和空格并停止直到找到标点符号,正则表达式是:

        ([\w\s]{0,})[^\w\s]{0,}
        

        所以计算将采用括号中的单词和空格:

        '(Mary had a little lamb),( Jack went up the hill, Jill followed suit),( i woke up suddenly),( it was a really bad dream)...'
        

        结果仍不清楚,包含一些“无”字符。所以我用这个脚本删除了“无”字符:

        [x for x in regex_of_sentence if x is not '']
        

        因此该段落将标记为句子,但不清楚句子的结果是:

        ['Mary had a little lamb', ' Jack went up the hill', ' Jill followed suit', ' i woke up suddenly', ' it was a really bad dream']
        

        如您所见,结果显示了一些以空格开头的句子。所以为了在不开始空格的情况下制作一个清晰的段落,我制作了这个正则表达式:

        \s([\w\s]{0,})
        

        它会做出一个清晰的句子,如:

        ['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
        

        所以,我们必须做两个过程才能取得好的结果。

        你的问题的答案从这里开始……

        为了将句子标记为单词,我进行段落迭代并使用正则表达式来捕获单词,同时它正在使用这个正则表达式进行迭代:

        ([\w]{0,})
        

        然后再次清除空字符:

        [x for x in regex_of_word if x is not '']
        

        所以结果真的很清楚只有单词列表:

        ['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
        

        以后要做好NLP,你需要有自己的词组数据库,如果词组在句子中,搜索一下,做成词组列表后,剩下的词就是一个词了。

        使用这种方法,我可以用我的语言(印度尼西亚语)构建我自己的 NLP,它真的非常缺乏模块。

        编辑:

        我没有看到您想要比较单词的问题。所以你还有一句话要比较....我给你的​​奖金不仅仅是奖金,我给你怎么算。

        mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']
        

        在这种情况下,您必须执行的步骤是: 1. 迭代 mod_example 2. 将第一句话与 mod_example 中的单词进行比较。 3. 计算一下

        所以脚本将是:

        import re
        from contextlib import redirect_stdout
        from io import StringIO
        
        example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
        mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']
        
        def token_to_sentence(str):
            f = StringIO()
            with redirect_stdout(f):
                regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
                regex_of_sentence = [x for x in regex_of_sentence if x is not '']
                for i in regex_of_sentence:
                    print(i)
                first_step_to_sentence = (f.getvalue()).split('\n')
            g = StringIO()
            with redirect_stdout(g):
                for i in first_step_to_sentence:
                    try:
                        regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
                        print(regex_to_clear_sentence.group(1))
                    except:
                        print(i)
                sentence = (g.getvalue()).split('\n')
            return sentence
        
        def token_to_words(str):
            f = StringIO()
            with redirect_stdout(f):
                for i in str:
                    regex_of_word = re.findall('([\w]{0,})', i)
                    regex_of_word = [x for x in regex_of_word if x is not '']
                    for word in regex_of_word:
                        print(regex_of_word)
                words = (f.getvalue()).split('\n')
        
        def convert_to_words(str):
            sentences = token_to_sentence(str)
            for i in sentences:
                word = token_to_words(i)
            return word
        
        def compare_list_of_words__to_another_list_of_words(from_strA, to_strB):
                fromA = list(set(from_strA))
                for word_to_match in fromA:
                    totalB = len(to_strB)
                    number_of_match = (to_strB).count(word_to_match)
                    data = str((((to_strB).count(word_to_match))/totalB)*100)
                    print('words: -- ' + word_to_match + ' --' + '\n'
                    '       number of match    : ' + number_of_match + ' from ' + str(totalB) + '\n'
                    '       percent of match   : ' + data + ' percent')
        
        
        
        #prepare already make, now we will use it. The process start with script below:
        
        if __name__ == '__main__':
            #tokenize paragraph in example to sentence:
            getsentences = token_to_sentence(example)
        
            #tokenize sentence to words (sentences in getsentences)
            getwords = token_to_words(getsentences)
        
            #compare list of word in (getwords) with list of words in mod_example
            compare_list_of_words__to_another_list_of_words(getwords, mod_example)
        

        【讨论】:

          【解决方案6】:

          您可以使用 nltk (as @alvas suggests) 和一个递归函数,它接受任何对象并标记每个 str :

          from nltk.tokenize import word_tokenize
          def tokenize(obj):
              if obj is None:
                  return None
              elif isinstance(obj, str): # basestring in python 2.7
                  return word_tokenize(obj)
              elif isinstance(obj, list):
                  return [tokenize(i) for i in obj]
              else:
                  return obj # Or throw an exception, or parse a dict...
          

          用法:

          data = [["Lorem ipsum dolor. Sit amet?", "Hello World!", None], ["a"], "Hi!", None, ""]
          print(tokenize(data))
          

          输出:

          [[['Lorem', 'ipsum', 'dolor', '.', 'Sit', 'amet', '?'], ['Hello', 'World', '!'], None], [['a']], ['Hi', '!'], None, []]
          

          【讨论】:

            【解决方案7】:

            您可以在 NLTK (http://nltk.org/api/nltk.tokenize.html) 中使用单词标记器进行列表理解,请参阅 http://docs.python.org/2/tutorial/datastructures.html#list-comprehensions

            >>> from nltk.tokenize import word_tokenize
            >>> example = ['Mary had a little lamb' , 
            ...            'Jack went up the hill' , 
            ...            'Jill followed suit' ,    
            ...            'i woke up suddenly' ,
            ...            'it was a really bad dream...']
            >>> tokenized_sents = [word_tokenize(i) for i in example]
            >>> for i in tokenized_sents:
            ...     print i
            ... 
            ['Mary', 'had', 'a', 'little', 'lamb']
            ['Jack', 'went', 'up', 'the', 'hill']
            ['Jill', 'followed', 'suit']
            ['i', 'woke', 'up', 'suddenly']
            ['it', 'was', 'a', 'really', 'bad', 'dream', '...']
            

            【讨论】:

            • 我强烈建议不要使用 NLTK。尽管很受欢迎(因为它是第一个有据可查的 Python NLP 包),但它早已过时了。而且word_tokenize有转换输入的习惯。
            • 同意转换输入,但恕我直言,标记化不应被视为一种转换,而是一种注释。注释是在数据之上添加信息层,而不是替换数据 =)(免责声明:我确实为 NLTK 做出了贡献)
            • 此外,NLTK 中有不止 1 个标记器,NLP 社区广泛使用的原始树库标记器虽然已经过时,但并不是万能的灵丹妙药。自github.com/nltk/nltk/issues/1214 以来,NLTK 中包含/移植/包装了更多的标记器,包括 Moses(来自机器翻译)、Toktok(来自语言建模)、REPP(来自语法工程)和用于多种语言的斯坦福 CoreNLP 标记器( github.com/nltk/nltk/pull/1735#issuecomment-306091826)
            • 如果有人正在寻找速度和可定制的标记器,请查看 SpaCy 标记器spacy.io/docs/usage/customizing-tokenizer。如果在 NLTK 中也有类似的贡献,那就太好了 =)
            【解决方案8】:

            对我来说,很难说你想做什么。

            这个怎么样

            exclude = set(['Mary', 'Jack', 'Jill', 'i', 'it'])
            
            mod_example = []
            for sentence in example:
                words = sentence.split()
                # Optionally sort out some words
                for word in words:
                    if word in exclude:
                        words.remove(word)
                mod_example.append('\'' + '\' \''.join(words) + '\'')
            
            print mod_example
            

            哪些输出

            ["'had' 'a' 'little' 'lamb'", "'went' 'up' 'the' 'hill'", "'followed' 'suit'", 
            "'woke' 'up' 'suddenly'", "'was' 'a' 'really' 'bad' 'dream...'"]
            >>> 
            

            编辑: 基于 OP 提供的进一步信息的另一个建议

            example = ['Area1 Area1 street one, 4454 hikoland' ,
                       'Area2 street 2, 52432 hikoland, area2' ,
                       'Area3 ave three, 0534 hikoland' ]
            
            mod_example = []
            for sentence in example:
                words = sentence.split()
                # Sort out some words
                col1 = words[0]
                col2 = words[1:]
                if col1 in col2:
                    col2.remove(col1)
                elif col1.lower() in col2:
                    col2.remove(col1.lower())
                mod_example.append(col1 + ': ' + ' '.join(col2))
            

            输出

            >>>> print mod_example
            ['Area1: street one, 4454 hikoland', 'Area2: street 2, 52432 hikoland,', 
            'Area3: ave three, 0534 hikoland']
            >>> 
            

            【讨论】:

            • 这还是一个列表吗?这就是我想要的……是的,我也想要每个句子的第一个单词……谢谢……我会检查一下
            • 这会容易得多,@Sword,如果你能说出你正在解决的基本问题是什么。
            • @Sword 只是如果你问精确的操作,没有人能给出解决底层问题的替代方法
            • 哦。我会详细说明。假设我有一个 tsv 文件,它在 1 列中显示区域名称,在第 2 列中显示确切地址(如建筑物名称、街道等)。有许多这样的地址,例如 [jack went up the hill , jill follow suit],其中逗号代表下一行。第 1 列使用自动填充,所以区域名称是正确的,但第 2 列可能有错误。第 2 列(确切地址)中输入的区域名称可能有误。我需要做的是将第一个与第二个进行比较,如果它出现在第 2 列中,则删除重复的区域名称 ..
            • 如果逗号分隔行,用@Sword分隔的列是什么?
            猜你喜欢
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 2014-03-17
            • 2020-05-23
            • 1970-01-01
            • 1970-01-01
            • 2020-10-09
            相关资源
            最近更新 更多