这是一个有趣的问题。如果您愿意在 .NET 中的 Windows 上工作,您可以使用 dotImage 来完成此操作(免责声明,我为 Atalasoft 工作并编写了大部分 OCR 引擎代码)。让我们将问题分解为多个部分 - 首先是遍历所有 PDF:
string[] candidatePDFs = Directory.GetFiles(sourceDirectory, "*.pdf");
PdfDecoder decoder = new PdfDecoder();
foreach (string path in candidatePDFs) {
using (FileStream stm = new FileStream(path, FileMode.Open)) {
if (decoder.IsValidFormat(stm)) {
ProcessPdf(path, stm);
}
}
}
这将获取以 .pdf 结尾的所有文件的列表,如果该文件是有效的 pdf,则调用一个例程来处理它:
public void ProcessPdf(string path, Stream stm)
{
using (Document doc = new Document(stm)) {
int i=0;
foreach (Page p in doc.Pages) {
if (p.SingleImageOnly) {
ProcessWithOcr(path, stm, i);
}
else {
ProcessWithTextExtract(path, stm, i);
}
i++;
}
}
}
这会将文件作为 Document 对象打开,并询问每个页面是否仅为图像。如果是这样,它将 OCR 页面,否则它将提取文本:
public void ProcessWithOcr(string path, Stream pdfStm, int page)
{
using (Stream textStream = GetTextStream(path, page)) {
PdfDecoder decoder = new PdfDecoder();
using (AtalaImage image = decoder.Read(pdfStm, page)) {
ImageCollection coll = new ImageCollection();
coll.Add(image);
ImageCollectionImageSource source = new ImageCollectionImageSource(coll);
OcrEngine engine = GetOcrEngine();
engine.Initialize();
engine.Translate(source, "text/plain", textStream);
engine.Shutdown();
}
}
}
它的作用是将 PDF 页面光栅化为图像,并将其放入适合 engine.Translate 的形式。这并不一定需要以这种方式完成 - 可以通过调用识别从 AtalaImage 从引擎中获取 OcrPage 对象,但随后将由客户端代码循环遍历结构并写出文本。
您会注意到我省略了 GetOcrEngine() - 我们提供 4 个 OCR 引擎供客户使用:Tesseract、GlyphReader、RecoStar 和 Iris。您将选择最适合您需求的那一款。
最后,您需要代码从已经有完美文本的页面中提取文本:
public void ProcessWithTextExtract(string path, Stream pdfStream, int page)
{
using (Stream textStream = GetTextStream(path, page)) {
StreamWriter writer = new StreamWriter(textStream);
using (PdfTextDocument doc = new PdfTextDocument(pdfStream)) {
PdfTextPage page = doc.GetPage(i);
writer.Write(page.GetText(0, page.CharCount));
}
}
}
这会从给定页面中提取文本并将其写入输出流。
最后,你需要GetTextStream():
public Stream GetTextStream(string sourcePath, int pageNo)
{
string dir = Path.GetDirectoryName(sourcePath);
string fname = Path.GetFileNameWithoutExtension(sourcePath);
string finalPath = Path.Combine(dir, String.Format("{0}p{1}.txt", fname, pageNo));
return new FileStream(finalPath, FileMode.Create);
}
这会是 100% 的解决方案吗?不,当然不是。您可以想象 PDF 页面包含单个图像并在其周围绘制一个框 - 这显然会使仅图像测试失败,但不会返回任何有用的文本。可能更好的方法是仅使用提取的文本,如果没有返回任何内容,请尝试使用 OCR 引擎。从一种方法更改为另一种方法是编写不同的谓词。