【问题标题】:python glob or listdir to create then save files from one directory to anotherpython glob 或 listdir 创建然后将文件从一个目录保存到另一个
【发布时间】:2020-07-06 06:29:53
【问题描述】:

我正在将文档从 pdf 转换为文本。 pdfs 当前位于一个文件夹中,然后在 txt 转换后保存到另一个文件夹中。我有很多这样的文档,并且更喜欢迭代子文件夹并保存到 txt 文件夹中具有相同名称的子文件夹,但在添加该层时遇到问题。

我知道我可以使用 glob 递归迭代并为文件列表等执行此操作。但不清楚如何将文件从该文件夹保存到新文件夹。这不是完全必要的,但会更加方便和高效。

有什么好办法吗?

import os
import io
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
from pdfminer.pdfpage import PDFPage


def convert(fname, pages=None):
    if not pages:
        pagenums = set()
    else:
        pagenums = set(pages)

    output = io.StringIO()
    manager = PDFResourceManager()
    converter = TextConverter(manager, output, laparams=LAParams())
    interpreter = PDFPageInterpreter(manager, converter)

    infile = open(fname, 'rb')
    for page in PDFPage.get_pages(infile, pagenums):
        interpreter.process_page(page)
    infile.close()
    converter.close()
    text = output.getvalue()
    output.close
    return text 
    print(text)



def convertMultiple(pdfDir, txtDir):
    if pdfDir == "": pdfDir = os.getcwd() + "\\" #if no pdfDir passed in 
    for pdf in os.listdir(pdfDir): #iterate through pdfs in pdf directory
        fileExtension = pdf.split(".")[-1]
        if fileExtension == "pdf":
            pdfFilename = pdfDir + pdf 
            text = convert(pdfFilename)  
            textFilename = txtDir + pdf.split(".")[0] + ".txt"
            textFile = open(textFilename, "w")  
            textFile.write(text)  


pdfDir = r"C:/Users/Documents/pdf/"
txtDir = r"C:/Users/Documents/txt/"
convertMultiple(pdfDir, txtDir)   

【问题讨论】:

  • 那么如果在子目录中找到了一个pdf文件(例如C:/Users/Documents/pdf/Setec Astronomy/employees.pdf),你希望将文本文件保存到C:/Users/Documents/txt/Setec Astronomy/employees.pdf还是C:/Users/Documents/txt/employees.pdf?
  • @GordonAitchJay 第一个 C:/Users/Documents/txt/Setec Astronomy/employees.txt - 程序负责 txt 部分

标签: python directory glob listdir


【解决方案1】:

正如您所建议的,glob 在这里工作得很好。它甚至可以只过滤.pdf 文件。

测试完后取消注释这 3 行。

import os, glob

def convert_multiple(pdf_dir, txt_dir):
    if pdf_dir == "": pdf_dir = os.getcwd() # If no pdf_dir passed in 
    for filepath in glob.iglob(f"{pdf_dir}/**/*.pdf", recursive=True):
        text = convert(filepath)
        root, _ = os.path.splitext(filepath) # Remove extension
        txt_filepath = os.path.join(txt_dir, os.path.relpath(root, pdf_dir)) + ".txt"
        txt_filepath = os.path.normpath(txt_filepath) # Not really necessary
        print(txt_filepath)
#        os.makedirs(os.path.dirname(txt_filepath), exist_ok=True)
#        with open(txt_filepath, "wt") as f:
#            f.write(text)


pdf_dir = r"C:/Users/Documents/pdf/"
txt_dir = r"C:/Users/Documents/txt/"
convert_multiple(pdf_dir, txt_dir)   

要确定新.txt 文件的文件路径,请使用os.path 模块中的函数。

os.path.relpath(filepath, pdf_dir) 返回文件的文件路径,包括与pdf_dir 相关的任何子目录。

假设filepath 是:

C:/Users/Documents/pdf/Setec Astronomy/employees.pdf

而pdf_dir 是

C:/Users/Documents/pdf/

它将返回Setec Astronomy/employees.pdf,然后可以将其与txt_dir一起传递给os.path.join(),从而为我们提供包含额外子目录的完整文件路径。

您可以使用txt_filepath = filepath.replace(filepath, pdf_dir),但您必须确保所有相应的斜线都在同一个方向,并且没有多余/缺失的前导/尾随斜线。

在打开新的.txt 文件之前,需要创建所有子目录。调用os.path.dirname() 以获取文件目录的文件路径,并调用os.makedirs() 并将其exist_ok 参数设置为True,以在目录已存在时抑制FileExistsError 异常。

在打开.txt 文件时使用with 语句以避免显式调用.close(),尤其是在出现任何异常的情况下。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-01-22
    • 2016-04-04
    • 2020-09-02
    • 1970-01-01
    • 2012-06-26
    • 2017-07-20
    • 2012-02-15
    • 1970-01-01
    相关资源
    最近更新 更多