【问题标题】:Extract text from pdf file using javascript [duplicate]使用javascript从pdf文件中提取文本[重复]
【发布时间】:2013-06-29 18:21:07
【问题描述】:

我想在客户端仅使用 Javascript 而不使用服务器从 pdf 文件中提取文本。我已经在以下链接中找到了 javascript 代码:extract text from pdf in Javascript

然后在

http://hublog.hubmed.org/archives/001948.html

进入:

https://github.com/hubgit/hubgit.github.com/tree/master/2011/11/pdftotext

1) 我想知道从以前的文件中提取这些文件所需的文件是什么。 2) 我不确切知道如何在应用程序中调整这些代码,而不是在网络中。

欢迎任何答案。谢谢。

【问题讨论】:

标签: javascript pdf text-extraction pdf.js


【解决方案1】:

这是一个很好的例子,说明如何使用 pdf.js 提取文本: http://git.macropus.org/2011/11/pdftotext/example/

当然,你必须为你的目的删除很多代码,但它应该这样做

【讨论】:

  • 未来 Google 员工注意事项:官方 pdf.js 项目自上述链接发布以来似乎已多次易手,但它目前位于 Mozilla 的 GitHub 页面 - github.com/mozilla/pdf.js
  • @Allanon 你知道提取文本并保持其语义的任何方法吗?该示例仅抓取所有文本,而不考虑换行符、段落、标题等。
  • @Jun711 你是怎么得到换行符的?我实现了吗?
【解决方案2】:

我制作了一种更简单的方法,不需要使用相同的库(使用最新版本)using pdf.js 在 iframe 之间发布消息。

以下示例将仅从 PDF 的第一页中提取所有文本:

/**
 * Retrieves the text of a specif page within a PDF Document obtained through pdf.js 
 * 
 * @param {Integer} pageNum Specifies the number of the page 
 * @param {PDFDocument} PDFDocumentInstance The PDF document obtained 
 **/
function getPageText(pageNum, PDFDocumentInstance) {
    // Return a Promise that is solved once the text of the page is retrieven
    return new Promise(function (resolve, reject) {
        PDFDocumentInstance.getPage(pageNum).then(function (pdfPage) {
            // The main trick to obtain the text of the PDF page, use the getTextContent method
            pdfPage.getTextContent().then(function (textContent) {
                var textItems = textContent.items;
                var finalString = "";

                // Concatenate the string of the item to the final string
                for (var i = 0; i < textItems.length; i++) {
                    var item = textItems[i];

                    finalString += item.str + " ";
                }

                // Solve promise with the text retrieven from the page
                resolve(finalString);
            });
        });
    });
}

/**
 * Extract the test from the PDF
 */

var PDF_URL  = '/path/to/example.pdf';
PDFJS.getDocument(PDF_URL).then(function (PDFDocumentInstance) {

    var totalPages = PDFDocumentInstance.pdfInfo.numPages;
    var pageNumber = 1;

    // Extract the text
    getPageText(pageNumber , PDFDocumentInstance).then(function(textPage){
        // Show the text of the page in the console
        console.log(textPage);
    });

}, function (reason) {
    // PDF loading error
    console.error(reason);
});

Read the article about this solution here。正如@xarxziux 所提到的,自发布第一个解决方案以来,该库已经发生了变化(它不应再与最新版本的 pdf.js 一起使用)。这应该适用于大多数情况。

【讨论】:

  • 此方法未提供正确格式的数据。我们找不到换行符、段落的位置。
  • @RishabhGarg 请记住,PDF 不知道文本的格式甚至顺序。你很幸运,你能得到文本。导出的格式甚至可能不一致。这就是为什么原始演示用一个空格替换所有空格的原因。这至少有点保持格式一致。
  • PDFDocumentInstance.pdfInfo.numPages 现在应该是 PDFDocumentInstance.numPages
  • @Sancarn 你是对的。为了获得更好的结果,请改用 OCR(光学字符识别)。
  • @CarlosDelgado 我认为综合方法对个人来说是最好的。根据我的经验,OCR 往往会使许多字符不正确。结合这两种技术可能会产生更好的结果。但是不确定是否有任何图书馆。
猜你喜欢
  • 1970-01-01
  • 2011-01-18
  • 2014-09-11
  • 2012-12-30
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-04-30
  • 1970-01-01
相关资源
最近更新 更多