【问题标题】:Finding words that are the reverse of each other in a file在文件中查找彼此相反的单词
【发布时间】:2011-07-03 03:08:09
【问题描述】:

抱歉这个新手问题,刚刚开始。我想要一个简单的程序在文件中查找反向单词,所以写了这个源,但它不起作用。在它进入第二个“for”循环后,它不会回到第一个循环,而是结束程序。有什么线索吗?

def is_reverse(word1, word2):   
   if len(word1) == len(word2):
     if word1 == word2[::-1]:
       return True   
return False

fin = open('List.txt') 
for word1 in fin:
    word1 = word1.strip()
    word1 = word1.lower()
    for word2 in fin:
      word2 = word2.strip()
      word2 = word2.lower()
      print word1 + word2
      if is_reverse(word1, word2) is True:
             print word1 + ' is the opposite of ' + word2 

编辑: 我试图循环一个文件和一个列表,并得到了一个好奇的(对我来说)结果。如果我使用此代码,一切正常:

def is_reverse(word1, word2):
  if len(word1) == len(word2):
      if word1 == word2[::-1]:
        return True
  return False

fin = open('List.txt')
fin2 = ['test1','test2','test3','test4','test5']
for word1 in fin:
    word1 = word1.strip()
    word1 = word1.lower()
    for word2 in fin2:
      word2 = word2.strip()
      word2 = word2.lower()
      print word1 + word2
      if is_reverse(word1, word2) is True:
             print word1 + ' is the opposite of ' + word2

如果我交换 fin 和 fin2 第一个循环只进行一次迭代。谁能解释一下为什么?

【问题讨论】:

  • 粘贴真实代码,这个有语法错误和未定义的变量。
  • 你觉得我的解决方案怎么样?

标签: python file loops iterator for-loop


【解决方案1】:

for word1 in fin 逐行迭代,所以word1 真的是一行,而不是一个词。这是你的本意吗?

for word2 in fin 使用相同的迭代器,所以我认为它会消耗所有输入,而for word1 in fin 只会执行一次。

因此,最简单的更改是拥有两个文件,file1 和 file2,并在每次通过循环时重新打开 file2。

def is_reverse(word1, word2):   
   if len(word1) == len(word2):
     if word1 == word2[::-1]:
       return True   
return False

file1 = open('List.txt') 
for word1 in file1:
    word1 = word1.strip()
    word1 = word1.lower()
    file2 = open('List.txt')
    for word2 in file2:
      word2 = word2.strip()
      word2 = word2.lower()
      print word1 + word2
      if is_reverse(word1, word2):
             print word1 + ' is the opposite of ' + word2 

但可能更好的方法是将文件读入列表一次,然后遍历列表而不是文件,例如

def is_reverse(word1, word2):
    if len(word1) == len(word2):
        if word1 == word2[::-1]:
            return True
    return False

file = open('List.txt')
words = list(file)
for word1 in words:
    word1 = word1.strip()
    word1 = word1.lower()
    for word2 in words:
        word2 = word2.strip()
        word2 = word2.lower()
        print word1 + word2
        if is_reverse(word1, word2):
            print word1 + ' is the opposite of ' + word2 

回答你的另一个问题,关于为什么你可以迭代同一个列表两次但不能迭代同一个文件:

for element in iterable 循环通过调用iterable.__iter__ 向iterable 询问其迭代器。

当 Python 向文件询问其迭代器时,文件会返回自身。这意味着文件上的每个迭代器都共享相同的状态/位置。

>>> file = open('testfile.txt')
>>> it1 = iter(file)
>>> it2 = iter(file)
>>> id(it1)
3078689064L
>>> id(it2)
3078689064L
>>> id(file)
3078689064L

当您向列表询问其迭代器时,您每次都会得到一个不同的迭代器,并带有关于其位置的单独信息。

>>> list = [1,2,3]
>>> it3 = iter(list)
>>> it4 = iter(list)
>>> id(it3)
3078746156L
>>> id(it4)
3078746188L
>>> id(list)
3078731244L

后记

正如 Hugh 所指出的,对每个单词的单词列表进行迭代是非常低效的。

这是一种更快的方法。将List.txt 更改为一个非常大的文件,例如/usr/share/dict/words 在 Linux 系统上看看我的意思。

words = []
wordset = set(())

file = open('List.txt')
for line in file:
    word = line.strip('\n')
    words.append(word)
    wordset.add(word)

for word in words:
    reversed = word[::-1]
    if reversed in wordset:
        print word + ' is the opposite of ' + reversed

【讨论】:

  • 正确。最简单的解决方案是将打开的行更改为fin = list(open('List.txt')) 以将整个文件缓存在内存中。 (并且可能已知源文件每行只包含一个单词 - 在家庭作业中很常见......)
  • 好的,我认为这是我问题的最佳答案。 :)
【解决方案2】:

如果您真的想将列表与自身进行比较,您可以通过测试 set 中的值来避免迭代:

def getWords(fname):
    with open(fname) as inf:
        words = list(w.strip().lower() for w in inf)
    ws = set(words)
    words = list(ws)
    words.sort()
    return words, ws

def wordsInReverse(words, wordset):
    for w in words:
        rw = w[::-1]  # reverse the string
        if rw in wordset:
            yield w,rw

def main():
    words, wordSet = getWords('List.txt')

    for w,rw in wordsInReverse(words, wordSet):
        if rw >= w:  # don't print duplicates
            print('{0} is the opposite of {1}'.format(w, rw))        

if __name__=="__main__":
    main()

并交叉比较两个文件:

from itertools import chain

def main():
    words1, wordSet1 = getWords('List1.txt')
    words2, wordSet2 = getWords('List2.txt')

    for w,rw in chain(wordsInReverse(words1, wordSet2), wordsInReverse(words2, wordSet1)):
        print('{0} is the opposite of {1}'.format(w, rw))        

【讨论】:

    【解决方案3】:

    我的猜测是您在两个循环中都在迭代“fin”(尽管您的示例代码在第一个循环中有一个神秘的变量“x”)。而是尝试在每个循环中对文件使用单独的句柄,如下所示:

    fin1 = open("list.txt")
    for word1 in fin1:
        fin2 = open("list.txt")
        for word2 in fin2:
            ...etc...
    

    【讨论】:

    • 但这是个好主意吗?我的意思是,如果n 是文件中的行数,那么该文件将被打开/读取多少次? n+1 次,对吧?应该不需要多次读取文件。
    • 好的,感谢您的回答! X 变量是我对列表进行的一些测试的结果,因此我编辑了代码。我注意到,即使我使用文件和列表,如果我在开始循环之前将其加载到变量中,循环仍然会停止,但前提是该文件用于 loop2。
    • @nuNce:文件和列表不同的原因是file.__iter__返回self,但list.__iter__在每次调用时返回一个新对象。请参阅我的更新答案。
    【解决方案4】:

    应该不需要读取文件 不止一次。

    –克劳斯·比斯科夫·霍夫曼

    也就是说,对单词进行两次迭代是非常耗时的:如果一个文件包含1000个单词,那么每个单词的反转可能会与1000个单词进行比较,即总共1000000次比较;

    这是一个只有一次迭代的代码,字典会提醒它已经看到的内容

    with open('palindromic.txt') as f:
        ch = f.read()
        li = [ w for w in ch.split() if len(w)>1 ]
    
    dic ={}
    pals = set([])
    
    for line in li:
        word = line.strip().lower()
        if len(word)>1:
            if word not in dic:
                dic[word] = 1
                if word[::-1] in dic and word[::-1]!=word:
                    pals.add(word)
            else:
                dic[word] += 1
    
    
    for w in pals:
        print w,dic[w],'  ',w[::-1],dic[w[::-1]]
    

    [ w for w in ch.split() if len(w)>1 ] 必须改进以从每个单词中删除括号、撇号等

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2011-05-13
      • 2013-04-29
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-09-21
      • 1970-01-01
      相关资源
      最近更新 更多