【问题标题】:python script to concatenate all the files in the directory into one filepython脚本将目录中的所有文件连接到一个文件中
【发布时间】:2013-07-19 15:08:05
【问题描述】:

我编写了以下脚本来将目录中的所有文件连接到一个文件中。

这可以优化吗,就

而言
  1. 惯用的python

  2. 时间

这里是sn-p:

import time, glob

outfilename = 'all_' + str((int(time.time()))) + ".txt"

filenames = glob.glob('*.txt')

with open(outfilename, 'wb') as outfile:
    for fname in filenames:
        with open(fname, 'r') as readfile:
            infile = readfile.read()
            for line in infile:
                outfile.write(line)
            outfile.write("\n\n")

【问题讨论】:

标签: python file copy


【解决方案1】:

使用shutil.copyfileobj复制数据:

import shutil

with open(outfilename, 'wb') as outfile:
    for filename in glob.glob('*.txt'):
        if filename == outfilename:
            # don't want to copy the output into the output
            continue
        with open(filename, 'rb') as readfile:
            shutil.copyfileobj(readfile, outfile)

shutil 以块的形式从readfile 对象中读取,直接将它们写入outfile 文件对象。不要使用readline() 或迭代缓冲区,因为您不需要查找行尾的开销。

阅读和写作使用相同的模式;这在使用 Python 3 时尤其重要;我在这里都使用了二进制模式。

【讨论】:

  • 为什么使用相同的读写模式很重要?
  • @JuanDavid:因为shutil 将使用.read() 调用一个,.write() 调用另一个文件对象,将读取的数据从一个传递到另一个。如果一个以二进制模式打开,另一个以文本模式打开,则您正在传递不兼容的数据(二进制数据到文本文件,或文本数据到二进制文件)。
  • 这里的代码不适用于 CSV 文件,该死。但它确实给了我一些很好的灵感,让我知道如何用 CSV 完成这个任务。我对 Python 比较陌生。
  • @bretts:文件的内容无关紧要;也许您的 CSV 文件缺少最后一个换行符,或者使用了不同的分隔符格式?
【解决方案2】:

您可以直接遍历文件对象的行,而无需将整个内容读入内存:

with open(fname, 'r') as readfile:
    for line in readfile:
        outfile.write(line)

【讨论】:

    【解决方案3】:

    不需要使用那么多变量。

    with open(outfilename, 'w') as outfile:
        for fname in filenames:
            with open(fname, 'r') as readfile:
                outfile.write(readfile.read() + "\n\n")
    

    【讨论】:

      【解决方案4】:

      我很想了解更多关于性能的信息,我使用了 Martijn Pieters 和 Stephen Miller 的答案。

      我尝试了使用shutil 和没有shutil 的二进制和文本模式。我尝试合并 270 个文件。

      文本模式 -

      def using_shutil_text(outfilename):
          with open(outfilename, 'w') as outfile:
              for filename in glob.glob('*.txt'):
                  if filename == outfilename:
                      # don't want to copy the output into the output
                      continue
                  with open(filename, 'r') as readfile:
                      shutil.copyfileobj(readfile, outfile)
      
      def without_shutil_text(outfilename):
          with open(outfilename, 'w') as outfile:
              for filename in glob.glob('*.txt'):
                  if filename == outfilename:
                      # don't want to copy the output into the output
                      continue
                  with open(filename, 'r') as readfile:
                      outfile.write(readfile.read())
      

      二进制模式 -

      def using_shutil_text(outfilename):
          with open(outfilename, 'wb') as outfile:
              for filename in glob.glob('*.txt'):
                  if filename == outfilename:
                      # don't want to copy the output into the output
                      continue
                  with open(filename, 'rb') as readfile:
                      shutil.copyfileobj(readfile, outfile)
      
      def without_shutil_text(outfilename):
          with open(outfilename, 'wb') as outfile:
              for filename in glob.glob('*.txt'):
                  if filename == outfilename:
                      # don't want to copy the output into the output
                      continue
                  with open(filename, 'rb') as readfile:
                      outfile.write(readfile.read())
      

      二进制模式的运行时间 -

      Shutil - 20.161773920059204
      Normal - 17.327500820159912
      

      文本模式的运行时间 -

      Shutil - 20.47757601737976
      Normal - 13.718038082122803
      

      看起来在两种模式下,shutil 执行相同,而文本模式比二进制更快。

      操作系统:Mac OS 10.14 Mojave。 Macbook Air 2017。

      【讨论】:

        【解决方案5】:

        使用 Python 2.7,我做了一些“基准”测试

        outfile.write(infile.read())
        

        对

        shutil.copyfileobj(readfile, outfile)
        

        我迭代了 20 多个 .txt 文件,大小从 63 MB 到 313 MB 不等,联合文件大小约为 2.6 GB。在这两种方法中,普通读取模式的性能都优于二进制读取模式,shutil.copyfileobj 通常比 outfile.write 快。

        将最差组合(outfile.write,二进制模式)与最佳组合(shutil.copyfileobj,普通读取模式)进行比较,差异非常显着:

        outfile.write, binary mode: 43 seconds, on average.
        
        shutil.copyfileobj, normal mode: 27 seconds, on average.
        

        正常读取模式下输出文件的最终大小为 2620 MB,而二进制读取模式下为 2578 MB。

        【讨论】:

        • 有趣。那是什么平台?
        • 我大致在两个平台上工作:Linux Fedora 16、不同的节点或带有 Intel Core(TM)2 Quad CPU Q9550、2.83 GHz 的 Windows 7 Enterprise SP1。我认为是后者。
        【解决方案6】:

        fileinput 模块提供了一种自然的方式来迭代多个文件

        for line in fileinput.input(glob.glob("*.txt")):
            outfile.write(line)
        

        【讨论】:

        • 如果它不局限于一次读取一行会更好。
        • @Marcin,这是正确的。我曾经认为这是一个很酷的解决方案 - 直到我看到 Martijn Pieter 的 shutil.copyfileobj humdinger。
        猜你喜欢
        • 2013-06-03
        • 1970-01-01
        • 2015-05-03
        • 1970-01-01
        • 2021-02-19
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多