【发布时间】:2019-01-18 02:17:48
【问题描述】:
我使用的是 32 位的 Windows 7。当我解析俄语文本 PDF 时,我收到带有 ??? 的结果文件而不是俄语字符。 开发人员通过此修复解决了此问题
我知道了?在 Windows 上带有结果的字符。我怎样才能避免它?如果 PDF 的编码是 UTF-8,你应该在你的终端上设置 chcp 65001 在启动 Python 进程之前。
chcp 65001
我在 windows cmd 中更改了这个但没有结果。
我的代码
import tabula
tabula.convert_into(r"C:\Code\Active\kartoteka\misc\ExampleExtract.pdf", r"C:\Code\Active\kartoteka\misc\output.csv", output_format="csv",pages = "all",java_options="-Dfile.encoding=utl-8")
错误日志:
?? 10, 2018 11:15:18 PM org.apache.pdfbox.pdmodel.font.PDCIDFontType2Font getawtFont
INFO: Can't read the embedded font Times-Roman
??? 10, 2018 11:15:18 PM org.apache.pdfbox.pdmodel.font.PDCIDFontType2Font getawtFont
INFO: Using font Times New Roman instead
??? 10, 2018 11:15:19 PM org.apache.pdfbox.pdmodel.font.PDCIDFontType2Font getawtFont
INFO: Can't read the embedded font Times-Roman
??? 10, 2018 11:15:19 PM org.apache.pdfbox.pdmodel.font.PDCIDFontType2Font getawtFont
INFO: Using font Times New Roman instead
我生成的文件仍然在 ????? 中显示所有俄语字符 你如何解决这个问题?
【问题讨论】:
-
它是正确的 java_options 还是错字?应该是
java_options="-Dfile.encoding=UTF8"。另见:stackoverflow.com/questions/6031877/…
标签: python-3.x tabula