【问题标题】:reading tar file contents without untarring it, in python script在 python 脚本中读取 tar 文件内容而不解压缩它
【发布时间】:2010-01-07 05:58:08
【问题描述】:

我有一个 tar 文件,里面有很多文件。 我需要编写一个 python 脚本,它将读取文件的内容并给出总字符数,包括字母总数、空格、换行符等所有内容,而无需解压缩 tar 文件。

【问题讨论】:

  • 如何在不提取到其他地方的情况下计算字符/字母/空格/所有内容?
  • 问的正是这个问题。

标签: python tar


【解决方案1】:

你可以使用getmembers()

>>> import  tarfile
>>> tar = tarfile.open("test.tar")
>>> tar.getmembers()

之后,您可以使用extractfile() 将成员提取为文件对象。只是一个例子

import tarfile,os
import sys
os.chdir("/tmp/foo")
tar = tarfile.open("test.tar")
for member in tar.getmembers():
    f=tar.extractfile(member)
    content=f.read()
    print "%s has %d newlines" %(member, content.count("\n"))
    print "%s has %d spaces" % (member,content.count(" "))
    print "%s has %d characters" % (member, len(content))
    sys.exit()
tar.close()

上面例子中的文件对象f,可以使用read()readlines()等。

【讨论】:

  • “for member in tar.getmembers()”可以更改为“for member in tar”,它可以是生成器或迭代器(我不确定是哪个)。但它一次只有一个成员。
  • 我刚刚遇到了类似的问题,但 tarfile 模块似乎吃掉了我的内存,即使我使用了 'r|' 选项。
  • 啊。我解决了。假设您将按照 Huggie 的提示编写代码,则必须不时“清理”成员列表。所以给定上面的代码示例,那就是tar.members = []。更多信息在这里:bit.ly/JKXrg6
  • tar.getmembers()放入for member in tar.getmembers()循环时会被多次调用吗?
  • 执行“f=tar.extractfile(member)”后,还需要关闭f吗?
【解决方案2】:

您需要使用 tarfile 模块。具体来说,您使用 TarFile 类的实例来访问文件,然后使用 TarFile.getnames() 访问名称

 |  getnames(self)
 |      Return the members of the archive as a list of their names. It has
 |      the same order as the list returned by getmembers().

如果您想阅读内容,则使用此方法

 |  extractfile(self, member)
 |      Extract a member from the archive as a file object. `member' may be
 |      a filename or a TarInfo object. If `member' is a regular file, a
 |      file-like object is returned. If `member' is a link, a file-like
 |      object is constructed from the link's target. If `member' is none of
 |      the above, None is returned.
 |      The file-like object is read-only and provides the following
 |      methods: read(), readline(), readlines(), seek() and tell()

【讨论】:

    【解决方案3】:

    之前,这篇文章展示了一个“dict(zip(()”) 将成员名称和成员列表放在一起的示例,这很愚蠢并且会导致对存档的过度读取,为了实现同样的效果,我们可以使用字典理解:

    index = {i.name: i for i in my_tarfile.getmembers()}
    

    更多关于如何使用 tarfile 的信息

    提取一个 tarfile 成员

    #!/usr/bin/env python3
    import tarfile
    
    my_tarfile = tarfile.open('/path/to/mytarfile.tar')
    
    print(my_tarfile.extractfile('./path/to/file.png').read())
    

    索引 tar 文件

    #!/usr/bin/env python3
    import tarfile
    import pprint
    
    my_tarfile = tarfile.open('/path/to/mytarfile.tar')
    
    index = my_tarfile.getnames()  # a list of strings, each members name
    # or
    # index = {i.name: i for i in my_tarfile.getmembers()}
    
    pprint.pprint(index)
    

    索引、读取、动态额外的 tar 文件

    #!/usr/bin/env python3
    
    import tarfile
    import base64
    import textwrap
    import random
    
    # note, indexing a tar file requires reading it completely once
    # if we want to do anything after indexing it, it must be a file
    # that can be seeked (not a stream), so here we open a file we
    # can seek
    my_tarfile = tarfile.open('/path/to/mytar.tar')
    
    
    # tarfile.getmembers is similar to os.stat kind of, it will
    # give you the member names (i.name) as well as TarInfo attributes:
    #
    # chksum,devmajor,devminor,gid,gname,linkname,linkpath,
    # mode,mtime,name,offset,offset_data,path,pax_headers,
    # size,sparse,tarfile,type,uid,uname
    #
    # here we use a dictionary comprehension to index all TarInfo
    # members by the member name
    index = {i.name: i for i in my_tarfile.getmembers()}
    
    print(index.keys())
    
    # pick your member
    # note: if you can pick your member before indexing the tar file,
    # you don't need to index it to read that file, you can directly
    # my_tarfile.extractfile(name)
    # or my_tarfile.getmember(name)
    
    # pick your filename from the index dynamically
    my_file_name = random.choice(index.keys())
    
    my_file_tarinfo = index[my_file_name]
    my_file_size = my_file_tarinfo.size
    my_file_buf = my_tarfile.extractfile( 
        my_file_name
        # or my_file_tarinfo
    )
    
    print('file_name: {}'.format(my_file_name))
    print('file_size: {}'.format(my_file_size))
    print('----- BEGIN FILE BASE64 -----'
    print(
        textwrap.fill(
            base64.b64encode(
                my_file_buf.read()
            ).decode(),
            72
        )
    )
    print('----- END FILE BASE64 -----'
    

    包含重复成员的 tar 文件

    如果我们有一个奇怪创建的 tar,在这个例子中,通过将同一个文件的多个版本附加到同一个 tar 存档中,我们可以仔细处理,我已经注释了哪些成员包含哪些文本,假设我们想要第四个(索引 3)成员,“capturetheflag\n”

    tar -tf mybadtar.tar 
    mymember.txt  # "version 1\n"
    mymember.txt  # "version 1\n"
    mymember.txt  # "version 2\n"
    mymember.txt  # "capturetheflag\n"
    mymember.txt  # "version 3\n"
    
    #!/usr/bin/env python3
    
    import tarfile
    my_tarfile = tarfile.open('mybadtar.tar')
    
    # >>> my_tarfile.getnames()
    # ['mymember.txt', 'mymember.txt', 'mymember.txt', 'mymember.txt', 'mymember.txt']
    
    # if we use extracfile on a name, we get the last entry, I'm not sure how python is smart enough to do this, it must read the entire tar file and buffer every valid member and return the last one
    
    # >>> my_tarfile.extractfile('mymember.txt').read()
    # b'version 3\n'
    
    # >>> my_tarfile.extractfile(my_tarfile.getmembers()[3]).read()
    # b'capturetheflag\n'
    

    或者,我们可以遍历 tar 文件 #!/usr/bin/env python3

    import tarfile
    my_tarfile = tarfile.open('mybadtar.tar')
    # note, if we do anything to the tarfile object that will 
    # cause a full read, the tarfile.next() method will return none,
    # so call next in a loop as the first thing you do if you want to
    # iterate
    
    while True:
        my_member = my_tarfile.next()
        if not my_member:
            break
        print((my_member.offset, mytarfile.extractfile(my_member).read,))
    
    # (0, b'version 1\n')
    # (1024, b'version 1\n')
    # (2048, b'version 2\n')
    # (3072, b'capturetheflag\n')
    # (4096, b'version 3\n')
    
    
        
    

    【讨论】:

    • 这让我向后寻求是不允许的例外
    • @KIC 在我上面的示例中,我们必须读取文件两次,一次是为了索引它(列出它包含的所有文件),第二次是按名称提取我们想要的文件,这是一个焦油结构的结果/特征。如果您需要从 tar 中提取文件,您只能读取一次(例如从流中),那么您必须提前知道文件名。如果您知道成员名称,并且只想提取一个成员,则可以在第一遍直接使用myArchive.extractfile('my/member/name.png') 提取它
    【解决方案4】:

    你可以使用 tarfile.list() 例如:

    filename = "abc.tar.bz2"
    with open( filename , mode='r:bz2') as f1:
        print(f1.list())
    

    获得这些数据后。您可以操作或将此输出写入文件并执行您的任何要求。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-10-08
      • 2014-06-20
      • 2012-09-29
      相关资源
      最近更新 更多