【问题标题】:Convert 4 sentences from text file and append all the words into a new list without repeating the words从文本文件中转换 4 个句子并将所有单词附加到一个新列表中而不重复单词
【发布时间】:2016-11-04 17:28:50
【问题描述】:

我一直在研究从 .txt 文件中读取 4 个句子并将所有单词附加到一个新的空列表中的程序。

我的代码如下:

fname = raw_input("Enter file name: ")
fh = open(fname)
lst = list()
for line in fh:
    line = line.rstrip()
    words = line.split()
    words.sort()
    if words not in lst:
      lst.append(words)
      print lst

我得到了以下结果:

[['但是', 'breaks', 'light', 'soft', 'through', 'what', 'window', '那边']] [['但是', 'breaks', 'light', 'soft', 'through', 'what', 'window', 'yonder'], ['It', 'Juliet', 'and', 'east', 'is', 'is', 'sun', 'the', 'the']] [['but', 'breaks', 'light', 'soft', 'through', 'what', 'window', 'yonder'], ['It', 'Juliet', 'and', 'east', 'is', 'is', 'sun', 'the', 'the'], ['Arise', 'and', '羡慕', 'fair', 'kill', 'moon', 'sun', 'the']] [['But', 'breaks', 'light', 'soft', '通过','什么','窗口','那边'],['它','朱丽叶','和', 'east', 'is', 'is', 'sun', 'the', 'the'], ['Arise', 'and', 'envious', 'fair', 'kill', 'moon', 'sun', 'the'], ['Who', 'already', 'and', '悲伤','是','苍白','生病','与']]

我可以做些什么来获得以下内容:

['Arise', 'But', 'It', 'Juliet', 'Who', 'already', 'and', 'breaks', '东方','羡慕','公平','悲伤','是','杀','光','月亮', '苍白','生病','柔软','阳光','the','透','什么','窗口', '与','那边']

句子是: 但是柔和的光线从那边的窗户打破 它是东方,朱丽叶是太阳 升起美丽的太阳,杀死嫉妒的月亮 谁已经病入膏肓,悲痛欲绝

【问题讨论】:

  • 你能把那4句话的内容显示出来吗?
  • 你能用文字解释一下输出的问题吗?这可能会帮助您了解算法出了什么问题。
  • 看看list.append和list.extend的区别。此外,如果您正在寻找独特的东西,那么您需要设置对象。

标签: python


【解决方案1】:

你想使用一个可以唯一列出元素的集合:

my_string = "But soft what light through yonder window breaks It is the east and Juliet is the sun Arise fair sun and kill the envious moon Who is already sick and pale with grief"    
lst = set(my_string.split(' '))

这会给你你想要的。您可以在字符串、列表等上使用setsets in python 3.5

【讨论】:

    【解决方案2】:

    您正在使用line.split() 正确地将每一行拆分为一个单词列表,但您并没有遍历您刚刚创建的名为words 的新列表。相反,您将列表words 作为对象与lst 的内容进行比较,然后将words 作为对象附加到lst。这会导致lst 成为列表列表,正如您在收到的结果中所显示的那样。

    为了获得您要查找的单词数组,您必须遍历 words 并单独添加每个单词,只要它不在 lst 中即可:

    for word in words:
        if word not in lst:
          lst.append(word)
    

    编辑:发现 another question/answer 涉及相同的问题 - 可能是针对相同的班级作业。

    【讨论】:

      【解决方案3】:

      最简单的方法是使用一个集合,并附加每个单词。

      file_name = raw_input("Enter file name: ")
      with open(file_name, 'r') as fh: 
          all_words = set()
          for line in fh:
              line = line.rstrip()
              words = line.split()
              for word in words:     
                  all_words.add(word)
      print(all_words)
      

      【讨论】:

        【解决方案4】:

        set 可用于删除重复项,split 方法将拆分任何类型的空格 - 包括行尾。所以这个任务可以简化为一个非常简单的单行:

        lst = sorted(set(open(fname).read().split()))
        

        【讨论】:

          【解决方案5】:

          我正在做同样的任务。我使用的代码如下:

          fname = input("Enter file name: ")
          fh = open(fname)
          lst = list()
          for line in fh:
              line = line.rstrip()
              words = line.split()
              for word in words:
                  if word not in lst:
                      lst.append(word)
          lst.sort()
          print(lst)
          

          【讨论】:

            猜你喜欢
            • 1970-01-01
            • 1970-01-01
            • 2020-10-09
            • 2020-05-23
            • 1970-01-01
            • 2016-08-09
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            相关资源
            最近更新 更多