【问题标题】:Nested loops iterating on a single file在单个文件上迭代的嵌套循环
【发布时间】:2016-03-20 05:57:49
【问题描述】:

我想删除文件中的某些特定行。 我要删除的部分包含在两行之间(也将被删除),名为STARTING_LINECLOSING_LINE。如果文件末尾没有结束行,则操作应该停止。

例子:

...blabla...
[Start] <-- # STARTING_LINE
This is the body that I want to delete
[End] <-- # CLOSING_LINE
...blabla...

我提出了三种不同的方法来实现相同的目标(加上下面 tdelaney 的回答提供的一种),但我想知道哪一种是最好的。请注意,我不是在寻找主观意见:我想知道是否有一些真正的理由让我应该选择一种方法而不是另一种方法。

1。很多if 条件(只有一个for 循环):

def delete_lines(filename):
    with open(filename, 'r+') as my_file:
        text = ''
        found_start = False
        found_end = False

        for line in my_file:
            if not found_start and line.strip() == STARTING_LINE.strip():
                found_start = True
            elif found_start and not found_end:
                if line.strip() == CLOSING_LINE.strip():
                    found_end = True
                continue
            else:
                print(line)
                text += line

        # Go to the top and write the new text
        my_file.seek(0)
        my_file.truncate()
        my_file.write(text)

2。在打开的文件上嵌套for 循环:

def delete_lines(filename):
    with open(filename, 'r+') as my_file:
        text = ''
        for line in my_file:
            if line.strip() == STARTING_LINE.strip():
                # Skip lines until we reach the end of the function
                # Note: the next `for` loop iterates on the following lines, not
                # on the entire my_file (i.e. it is not starting from the first
                # line). This will allow us to avoid manually handling the
                # StopIteration exception.
                found_end = False
                for function_line in my_file:
                    if function_line.strip() == CLOSING_LINE.strip():
                        print("stop")
                        found_end = True
                        break
                if not found_end:
                    print("There is no closing line. Stopping")
                    return False
            else:
                text += line

        # Go to the top and write the new text
        my_file.seek(0)
        my_file.truncate()
        my_file.write(text)

3。 while Truenext()StopIteration 例外)

def delete_lines(filename):
    with open(filename, 'r+') as my_file:
        text = ''
        for line in my_file:
            if line.strip() == STARTING_LINE.strip():
                # Skip lines until we reach the end of the function
                while True:
                    try:
                        line = next(my_file)
                        if line.strip() == CLOSING_LINE.strip():
                            print("stop")
                            break
                    except StopIteration as ex:
                        print("There is no closing line.")
            else:
                text += line

        # Go to the top and write the new text
        my_file.seek(0)
        my_file.truncate()
        my_file.write(text)

4。 itertools(来自 tdelaney 的回答):

def delete_lines_iter(filename):
    with open(filename, 'r+') as wrfile:
        with open(filename, 'r') as rdfile:
            # write everything before startline
            wrfile.writelines(itertools.takewhile(lambda l: l.strip() != STARTING_LINE.strip(), rdfile))
            # drop everything before stopline.. and the stopline itself
            try:
                next(itertools.dropwhile(lambda l: l.strip() != CLOSING_LINE.strip(), rdfile))
            except StopIteration:
                pass
            # include everything after
            wrfile.writelines(rdfile)
        wrfile.truncate()

这四个实现似乎达到了相同的结果。所以...

问题:我应该使用哪一个?哪一个是最 Pythonic 的?哪个效率最高?

有更好的解决方案吗?


编辑:我尝试使用timeit 评估大文件上的方法。为了每次迭代都有相同的文件,我删除了每个代码的编写部分;这意味着评估主要针对读取(和文件打开)任务。

t_if = timeit.Timer("delete_lines_if('test.txt')", "from __main__ import delete_lines_if")
t_for = timeit.Timer("delete_lines_for('test.txt')", "from __main__ import delete_lines_for")
t_while = timeit.Timer("delete_lines_while('test.txt')", "from __main__ import delete_lines_while")
t_iter = timeit.Timer("delete_lines_iter('test.txt')", "from __main__ import delete_lines_iter")

print(t_if.repeat(3, 4000))
print(t_for.repeat(3, 4000))
print(t_while.repeat(3, 4000))
print(t_iter.repeat(3, 4000))

结果:

# Using IF statements:
[13.85873354100022, 13.858520206999856, 13.851908310999988]
# Using nested FOR:
[13.22578497800032, 13.178281234999758, 13.155530822999935]
# Using while:
[13.254994718000034, 13.193942980999964, 13.20395484699975]
# Using itertools:
[10.547019549000197, 10.506679693000024, 10.512742852999963]

【问题讨论】:

  • 就我个人而言,我最喜欢 while 循环,根据您的 timeit,它是最有效的。但是,对于 SO,这个问题似乎有点过于基于意见,应该关闭。一般提示:任何看起来最易读的(如果效率差别不大的话)通常是最好的。
  • @RNar 谢谢。实际上,我正在寻找一种方法优于其他方法的具体原因,例如关于(例如)Python 最佳实践、效率、异常处理、Python 版本等。因为The Zen of Python 声明 “应该成为一种——最好只有一种——明显的方式”,我想知道在这种情况下应该选择哪一种。

标签: python performance if-statement for-loop while-loop


【解决方案1】:

您可以使用itertools 来制作它。我会对时间比较感兴趣。

import itertools
def delete_lines(filename):
    with open(filename, 'r+') as wrfile:
        with open(filename, 'r') as rdfile:
            # write everything before startline
            wrfile.writelines(itertools.takewhile(lambda l: l.strip() != STARTING_LINE.strip(), rdfile))
            # drop everything before stopline.. and the stopline itself
            next(itertools.dropwhile(lambda l: l.strip() != CLOSING_LINE.strip(), rdfile))
            # include everything after 
            wrfile.writelines(rdfile)
        wrfile.truncate()

【讨论】:

  • 谢谢,这个方法我没想过,很有意思。据我所见,它是性能方面最好的(我在上面的问题中添加了它的评估)。但是,它对我来说似乎也不太可读。在这种情况下,似乎在可读性和效率之间进行了权衡。
猜你喜欢
  • 2015-09-25
  • 2014-10-22
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-12-23
相关资源
最近更新 更多