【问题标题】:How to iterate through every directory specified and run commands on files (Python)如何遍历指定的每个目录并在文件上运行命令(Python)
【发布时间】:2016-09-06 19:52:24
【问题描述】:

我一直在编写一个脚本,该脚本将检查目录中的每个子目录并使用正则表达式匹配文件,然后根据文件的类型使用不同的命令。

所以我完成的是基于正则表达式匹配的不同命令的使用。现在它会检查 .zip 文件、.rar 文件或 .r00 文件,并为每个匹配项使用不同的命令。但是我需要帮助遍历每个目录并首先检查那里是否有 .mkv 文件,然后它应该只是传递该目录并跳转到下一个,但如果有匹配,它应该运行命令,然后当它完成继续下一个目录。

import os
import re

rx = '(.*zip$)|(.*rar$)|(.*r00$)'
path = "/mnt/externa/folder"

for root, dirs, files in os.walk(path):

    for file in files:
        res = re.match(rx, file)
        if res:
            if res.group(1):
                print("Unzipping ",file, "...")
                os.system("unzip " + root + "/" + file + " -d " + root)
            elif res.group(2):
                os.system("unrar e " + root + "/" + file + " " + root)
            if res.group(3):
                print("Unraring ",file, "...")
                os.system("unrar e " + root + "/" + file + " " + root)

编辑:

这是我现在拥有的代码:

import os
import re
from subprocess import check_call
from os.path import join

rx = '(.*zip$)|(.*rar$)|(.*r00$)'
path = "/mnt/externa/Torrents/completed/test"

for root, dirs, files in os.walk(path):
    if not any(f.endswith(".mkv") for f in files):
        found_r = False
        for file in files:
            pth = join(root, file)
            try:
                 if file.endswith(".zip"):
                    print("Unzipping ",file, "...")
                    check_call(["unzip", pth, "-d", root])
                    found_zip = True
                 elif not found_r and file.endswith((".rar",".r00")):
                     check_call(["unrar","e","-o-", pth, root,])
                     found_r = True
                     break
            except ValueError:
                print ("Oops! That did not work")

这个脚本工作正常,但有时当文件夹中有 Subs 时我似乎遇到了问题,这是我在运行脚本时收到的错误消息:

$ python unrarscript.py

UNRAR 5.30 beta 2 freeware      Copyright (c) 1993-2015    Alexander Roshal


Extracting from /mnt/externa/Torrents/completed/test/The.Conjuring.2013.1080p.BluRay.x264-ALLiANCE/Subs/the.conjuring.2013.1080p.bluray.x264-alliance.subs.rar

No files to extract
Traceback (most recent call last):
  File "unrarscript.py", line 19, in <module>
    check_call(["unrar","e","-o-", pth, root])
  File "/usr/lib/python2.7/subprocess.py", line 541, in     check_call
    raise CalledProcessError(retcode, cmd)
subprocess.CalledProcessError: Command '['unrar', 'e', '-o-', '/mnt/externa/Torrents/completed/test/The.Conjuring.2013.1080p.BluRay.x264-ALLiANCE/Subs/the.conjuring.2013.1080p.bluray.x264-alliance.subs.rar', '/mnt/externa/Torrents/completed/test/The.Conjuring.2013.1080p.BluRay.x264-ALLiANCE/Subs']' returned non-zero exit status 10

我无法真正理解代码有什么问题,所以我希望你们中的一些人愿意帮助我。

【问题讨论】:

  • Python 对缩进很敏感,因此您的代码不会像您发布的那样工作。我已经为你固定了间距。

标签: python loops


【解决方案1】:

在继续之前,只需使用 any 查看是否有任何文件以 .mkv 结尾,您也可以像做同样的事情一样简化为 if/else最后两场比赛。同样使用subprocess.check_call 会是更好的方法:

import os
import re
from subprocess import check_call
from os.path import join

rx = '(.*zip$)|(.*rar$)|(.*r00$)'
path = "/mnt/externa/folder"


for root, dirs, files in os.walk(path):
    if not any(f.endswith(".mkv") for f in files):
        for file in files:
            res = re.match(rx, file)
            if res:
                # use os.path.join 
                pth = join(root, file)
                # it can only be res.group(1) or  one of the other two so we only need if/else. 
                if res.group(1): 
                    print("Unzipping ",file, "...")
                    check_call(["unzip" , pth, "-d", root])
                else:
                    check_call(["unrar","e", pth,  root])

您也可以忘记 rex,只使用 if/elif 和 str.endswith:

for root, dirs, files in os.walk(path):
    if not any(f.endswith(".mkv") for f in files):
        for file in files:
            pth = join(root, file)
            if file.endswith("zip"):
                print("Unzipping ",file, "...")
                check_call(["unzip" , pth, "-d", root])
            elif file.endswith((".rar",".r00")):
                check_call(["unrar","e", pth,  root])

如果您真的关心不重复步骤和速度,您可以在迭代时进行过滤,您可以在检查 .mkv 并使用 for/else 逻辑时通过切片扩展收集:

good = {"rar", "zip", "r00"}
for root, dirs, files in os.walk(path):
    if not any(f.endswith(".mkv") for f in files):
        tmp = {"rar": [], "zip": []}
        for file in files:
            ext = file[-4:]
            if ext == ".mkv":
                break
            elif ext in good:
                tmp[ext].append(join(root, file))
        else:
            for p in tmp.get(".zip", []):
                print("Unzipping ", p, "...")
                check_call(["unzip", p, "-d", root])
            for p in tmp.get(".rar", []):
                check_call(["unrar", "e", p, root])

这将在.mkv 的任何匹配上短路,否则只会迭代.rar 或.r00 的任何匹配,但除非你真的关心效率,否则我会使用第二个逻辑。

为避免覆盖,您可以使用计数器将每个解压缩/解压缩到新的子目录,以帮助创建新的目录名称:

from itertools import count


for root, dirs, files in os.walk(path):
        if not any(f.endswith(".mkv") for f in files):
            counter = count()
            for file in files:
                pth = join(root, file)
                if file.endswith("zip"):
                    p = join(root, "sub_{}".format(next(counter)))
                    os.mkdir(p)
                    print("Unzipping ",file, "...")
                    check_call(["unzip" , pth, "-d", p])
                elif file.endswith((".rar",".r00")):
                    p = join(root, "sub_{}".format(next(counter)))
                    os.mkdir(p)
                    check_call(["unrar","e", pth,  p])

每个都将被解压到根目录下的一个新目录中,即root_path/sub_1等。

您可能会更好地为您的问题添加一个示例,但如果真正的问题是您只想要 .rar 或 .r00 之一,那么您可以在找到任何匹配的 .rar 或 .r00 时设置一个标志,并且只有在未设置标志时才解包:

for root, dirs, files in os.walk(path):
    if not any(f.endswith(".mkv") for f in files):
        found_r = False
        for file in files:
            pth = join(root, file)
            if file.endswith("zip"):
                print("Unzipping ",file, "...")
                check_call(["unzip", pth, "-d", root])
                found_zip = True
            elif not found_r and file.endswith((".rar",".r00"))
                check_call(["unrar","e", pth,  root])
                found_r = True     

如果也只有一个 zip,您可以设置两个标志并离开设置两个标志的循环。

【讨论】:

  • 我试过这个,它似乎工作得很好,但是当同时存在一个 .rar 文件和几个 .r00 文件时会出现问题,脚本将成功解压缩 .r00 文件,然后在完成后开始提取 .rar 文件,但问题是它们包含相同的内容,因此它只想替换刚刚解压缩的文件。有没有办法跳过这个?
  • @nillenilsson,是的,为每个目录使用单独的目录
  • 不是你看不懂,目录下的文件结构如下: file1.r00 file.r01 file.r02 ... file.r99 file.rar
  • 我明白,为每个目录创建一个唯一的目录并指定解压到该目录
  • 我运行了新脚本,但它仍然不正确,因为我会这样得到它: 1. 脚本运行并提取 .r00/.r01.. 文件,我得到一个 .mkv正确的文件夹,一切都很棒! 2. 但是现在脚本在文件夹中也找到了一个 .rar 并开始解包,但这次是在一个新文件夹中,这根本不是我想要的。这些文件是部分的。 .mkv 被打包成 rar 文件,我不需要提取 .r01 和 .rar 文件,因为它们包含相同的文件。因此,我只需要提取 .r01 或 .rar 之一,非常感谢您的帮助!
【解决方案2】:

下面的例子可以直接工作!正如@Padraic 所建议的,我将 os.system 替换为更合适的子进程。

将所有文件连接到一个字符串中并在字符串中查找 *.mkv 怎么样?

import os
import re
from subprocess import check_call
from os.path import join

rx = '(.*zip$)|(.*rar$)|(.*r00$)'
path = "/mnt/externa/folder"
regex_mkv = re.compile('.*\.mkv\,')
for root, dirs, files in os.walk(path):

    string_files = ','.join(files)+', '
    if regex_mkv.match(string_files): continue

    for file in files:
        res = re.match(rx, file)
        if res:
            # use os.path.join 
            pth = join(root, file)
            # it can only be res.group(1) or  one of the other two so we only need if/else. 
            if res.group(1): 
                print("Unzipping ",file, "...")
                check_call(["unzip" , pth, "-d", root])
            else:
                check_call(["unrar","e", pth,  root])

【讨论】:

  • 对不起,我不明白,我知道这会让我找到 .mkv 文件,但我如何解压缩 .zip 和 rar 文件?
  • 也许我真的不明白你想要什么...我建议的 sn-p 将包含至少一个以“.mkv”结尾的文件的目录。这不是你想要的?
  • @Padraic Cunningham 尽管更简单,但使用 any 的解决方案需要列表中的每个元素都有一个“if”。比较 780 个名称的列表,您的方法比使用正则表达式慢约 3 倍。无论如何,用子进程替换 os.system 非常有用!我将基于此编辑我的评论。
  • 您是否先为加入计时?你的正则表达式也是错误的,当最后一个文件以 .mkv 结尾时会发生什么?它看起来不像foo.mkv,,因此您需要添加更多逻辑来捕捉它。
  • @PadraicCunningham 是的,我考虑过加入时间。我更正了代码以解决您发现的错误...谢谢
【解决方案3】:

re 对于这样的事情来说太过分了。有一个用于提取文件扩展名的库函数os.path.splitext。在下面的示例中,我们构建了一个扩展名到文件名的映射,我们使用它来在恒定时间内检查 .mkv 文件的存在,并将每个文件名映射到适当的命令。

请注意,您可以使用zipfile(标准库)和第三方包are available for .rar files 解压缩文件。

import os

for root, dirs, files in os.walk(path):
    ext_map = {}
    for fn in files:
        ext_map.setdefault(os.path.splitext(fn)[1], []).append(fn)
    if '.mkv' not in ext_map:
        for ext, fnames in ext_map.iteritems():
            for fn in fnames:
                if ext == ".zip":
                    os.system("unzip %s -d %s" % (fn, root))
                elif ext == ".rar" or ext == ".r00":
                    os.system("unrar %s %s" % (fn, root))

【讨论】:

  • 构建字典是 O(n) 所以你不会减少搜索时间。 dict 唯一有意义的方法是,如果您使用 dict.get 在最后一个循环中查找扩展名。
  • 您必须迭代一次来检查 .mkv 和两次解压档案。由于字典是在检查 .mkv 时构建的,无论如何您都必须这样做,因此不会增加任何复杂性。但是,为每次迭代进行正则表达式匹配可能会使其成为二次方。
【解决方案4】:
import os
import re

regex = re.complile(r'(.*zip$)|(.*rar$)|(.*r00$)')
path = "/mnt/externa/folder"
for root, dirs, files in os.walk(path):
    for file in files:
        res = regex.match(file)
        if res:
           if res.group(1):
              print("Unzipping ",file, "...")
              os.system("unzip " + root + "/" + file + " -d " + root)
           elif res.group(2):
              os.system("unrar e " + root + "/" + file + " " + root)
           else:
              print("Unraring ",file, "...")
              os.system("unrar e " + root + "/" + file + " " + root)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2018-01-04
    • 2011-12-29
    • 1970-01-01
    • 2010-11-03
    • 2019-01-16
    • 1970-01-01
    • 1970-01-01
    • 2019-09-15
    相关资源
    最近更新 更多