【问题标题】:Why do I receive this error in Python PDFMiner: TypeError: can only concatenate str (not "bytes") to str为什么我在 Python PDFMiner 中收到此错误:TypeError: can only concatenate str (not "bytes") to str
【发布时间】:2020-12-24 18:00:55
【问题描述】:

我是 python 新手,尝试使用 PDFminer 将 pdf 转换为 txt 文件,每次TypeError: can only concatenate str (not "bytes") to str*- 时都会出现此错误

我很困惑,因为错误消息似乎表明错误是由于pdfminer 包中的文件引起的?我知道这里还有其他关于此错误消息的问题,但我无法根据它们找出我的问题 - 可能主要是因为我不知道他们的代码在做什么而且我是初学者,但也可能是因为它看起来像我的问题是由于与PDFminer 相关的文件。

我正在运行这段代码:

from pdfminer.layout import LAParams
from pdfminer.converter import TextConverter
from io import StringIO
from pdfminer.pdfpage import PDFPage

def get_pdf_file_content(path_to_pdf):
    resource_manager = PDFResourceManager(caching=True)
    out_text = StringIO
    laParams = LAParams()
    text_converter = TextConverter(resource_manager, out_text, laparams= laParams)
    fp = open(path_to_pdf, 'rb')
    interpreter = PDFPageInterpreter(resource_manager, text_converter)
    for page in PDFPage.get_pages(fp, pagenos=set(), maxpages=0, password="", caching= True, check_extractable= True):
        interpreter.process_page(page)

    text = out_text.getvalue()

    fp.close()
    text_converter.close()
    out_text.close()

    return text

path_to_pdf = "C:\\files\\raw\\AZO - CALLSTREET REPORT  AutoZone, Inc.(AZO), Q1 2002 Earnings Call, 5-December-2001 10 00 AM ET - 05-Dec-01.pdf"
print(get_pdf_file_content(path_to_pdf))

我收到此错误消息:

  File "<stdin>", line 1, in <module>
  File "<stdin>", line 8, in get_pdf_file_content
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfpage.py", line 122, in get_pages
    doc = PDFDocument(parser, password=password, caching=caching)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 575, in __init__
    self._initialize_password(password)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 599, in _initialize_password
    handler = factory(docid, param, password)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 300, in __init__
    self.init()
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 307, in init
    self.init_key()
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 320, in init_key
    self.key = self.authenticate(self.password)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 368, in authenticate
    key = self.authenticate_user_password(password)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 374, in authenticate_user_password
    key = self.compute_encryption_key(password)
  File "C:\text_analysis\project\lib\site-packages\pdfminer\pdfdocument.py", line 351, in compute_encryption_key
    password = (password + self.PASSWORD_PADDING)[:32]  # 1
TypeError: can only concatenate str (not "bytes") to str```

【问题讨论】:

    标签: python python-3.x pdf pdfminer


    【解决方案1】:

    你有两个选择:

    1) 您可以将密码设置为字节,从而以

    结尾
    for page in PDFPage.get_pages(fp, pagenos=set(), maxpages=0, password=b"", caching= True, check_extractable= True):
            interpreter.process_page(page)
    

    (注意定义密码的引号前的 b)

    2) 你可以摆脱那个论点

    密码参数不是强制性的(它有一个默认值),如果你不是特别需要它,你可以去掉它。你最终会得到:

    for page in PDFPage.get_pages(fp, pagenos=set(), maxpages=0, caching= True, check_extractable= True):
            interpreter.process_page(page)
    

    【讨论】:

      【解决方案2】:

      我之前遇到过这个问题。我将密码设置为字节,并将数据作为字节传递给解析器,它可以为我将多个 PDF 转换为多个 txt 文件。这是我的代码:

          def main():
      
              for path in Path(PDFS_FOLDER).glob("*.pdf"):
                  with path.open("rb") as file:
                       parser = PDFParser(file)
                       document = PDFDocument(parser, b"")
                       if not document.is_extractable:
                          continue
      
                       manager = PDFResourceManager()
                       params = LAParams()
      
                       device = PDFPageAggregator(manager, laparams=params)
                       interpreter = PDFPageInterpreter(manager, device)
              
                       password =b""
                       text = ""
      
                       for page in PDFPage.create_pages(document):
                             interpreter.process_page(page)
                             for obj in device.get_result():
                                 if isinstance(obj, LTTextBox) or isinstance(obj, LTTextLine):
                          text += obj.get_text()
                   with open(TEXTS_FOLDER + "{}.txt".format(path.stem), "w") as file:
                       file.write(text)
               return 0
      
      
           if __name__ == "__main__":
               import sys
               sys.exit(main())
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2021-08-31
        • 2020-08-29
        • 1970-01-01
        • 1970-01-01
        • 2020-11-21
        • 2022-11-01
        • 2019-02-02
        • 2020-09-18
        相关资源
        最近更新 更多