【发布时间】:2022-01-23 21:13:33
【问题描述】:
我们有许多 pdf 文件,它们都已解锁,它们有文本、图片等。每次我们必须在 adobe 上打开文件并手动执行时,我想也许有更好的方法来处理 PowerShell,如果不是,是的,我们有要做超过 1000 个文件,还有更多的文件,但谢谢你的回答 佩吉
【问题讨论】:
标签: windows powershell
我们有许多 pdf 文件,它们都已解锁,它们有文本、图片等。每次我们必须在 adobe 上打开文件并手动执行时,我想也许有更好的方法来处理 PowerShell,如果不是,是的,我们有要做超过 1000 个文件,还有更多的文件,但谢谢你的回答 佩吉
【问题讨论】:
标签: windows powershell
在进一步研究之后,我发现了一个命令行工具,您可以将它与 PowerShell 结合使用。它被称为 tesseract。对于 Windows 和 Linux,请下载 prebuilt binaries。对于 MacOS,您需要使用 MacPorts 或 Homebrew。
你会想做这样的事情:
# Using Get-ChildItem's -Include parameter to filter file types
# requires the target path to end in an asterisk. Using just an
# asterisk as the path makes it target the current directory.
foreach ($pdf in (Get-ChildItem * -Include *.pdf))
{
# An array isn't needed, it's just good for arranging arguments
tesseract @(
#INPUT:
$pdf
#OUTPUT:
"$($pdf.Directory)\{OCR} $($pdf.Name)"
#LANGUAGE:
'-l','eng'
)
# The directory is included in the output path so that you can
# change Get-ChildItem's target without adjusting the argument
}
或者,没有绒毛:
foreach ($pdf in (Get-ChildItem * -Include *.pdf))
{
tesseract $pdf "$($pdf.Directory)\{OCR} $($pdf.Name)" -l eng
}
当然,我还没有实际测试过 tesseract,但我确实阅读了其他问答页面以得出适当的命令。如果有任何问题,请告诉我。
【讨论】:
你的问题有点不清楚。有一种使用 PowerShell 对图像进行 OCR 的方法,例如使用 this function,您可以使用 this function 将 pdf 转换为图像(它确实需要 imagemagick,可以使用 here,如果你不这样做,还有便携式选项想要安装任何东西)。这将有效地让您搜索尚未经过 OCR 处理的 PDF 文件。
但是,就使用 PowerShell 直接编辑 PDF 文件以将其转换为 OCR 的 PDF 而言,虽然 PowerShell 功能可能会帮助您自动化该过程,但您首先需要找到一个可以执行此类操作的程序命令行。 PDF 文件也必须全部解锁,这样才能编辑它们(尽管有一些方法可以绕过 PDF 锁来解锁它们)。
不幸的是,我真的不知道有什么程序可以做到这一点。也许可以使用一些高级的Ghostscript 参数,但我还没有研究过。这肯定不容易!
【讨论】: