【问题标题】:how to know if a field is on a particular page?如何知道某个字段是否在特定页面上?
【发布时间】:2014-02-27 16:26:42
【问题描述】:

PDFbox 内容流是按页面完成的,但字段来自目录的表单,目录来自 pdf 文档本身。所以我不确定哪些字段在哪些页面上,以及它导致将文本写入不正确的位置/页面。

即。我正在处理每页的字段,但不确定哪些字段在哪些页面上。

有没有办法判断哪个字段在哪个页面上?或者,有没有办法只获取当前页面上的字段?

谢谢!

标记

代码sn-p:

PDDocument pdfDoc = PDDocument.load(file);
PDDocumentCatalog docCatalog = pdfDoc.getDocumentCatalog();
PDAcroForm acroForm = docCatalog.getAcroForm();

// Get field names
List<PDField> fieldList = acroForm.getFields();
List<PDPage> pages = pdfDoc.getDocumentCatalog().getAllPages();
for (PDPage page : pages) {
  PDPageContentStream contentStream = new PDPageContentStream(pdfDoc, page, true, true, true);
  processFields(acroForm, fieldList, contentStream, page);
  contentStream.close();
}

【问题讨论】:

    标签: java pdfbox


    【解决方案1】:

    PDFbox 内容流是按页面完成的,但字段来自目录的表单,目录来自 pdf 文档本身。所以我不确定哪些字段在哪些页面上

    这样做的原因是 PDF 包含定义表单的全局对象结构。此结构中的表单域可能在 0、1 或更多实际 PDF 页面上有 0、1 或更多可视化。此外,在只有 1 个可视化的情况下,允许合并字段对象和可视化对象。

    PDFBox 1.8.x

    不幸的是,PDFBox 在其PDAcroForm 和PDField 对象中仅表示此对象结构,并且不提供对相关页面的轻松访问。但是,通过访问底层结构,您可以建立连接。

    下面的代码应该清楚地说明如何做到这一点:

    @SuppressWarnings("unchecked")
    public void printFormFields(PDDocument pdfDoc) throws IOException {
        PDDocumentCatalog docCatalog = pdfDoc.getDocumentCatalog();
    
        List<PDPage> pages = docCatalog.getAllPages();
        Map<COSDictionary, Integer> pageNrByAnnotDict = new HashMap<COSDictionary, Integer>();
        for (int i = 0; i < pages.size(); i++) {
            PDPage page = pages.get(i);
            for (PDAnnotation annotation : page.getAnnotations())
                pageNrByAnnotDict.put(annotation.getDictionary(), i + 1);
        }
    
        PDAcroForm acroForm = docCatalog.getAcroForm();
    
        for (PDField field : (List<PDField>)acroForm.getFields()) {
            COSDictionary fieldDict = field.getDictionary();
    
            List<Integer> annotationPages = new ArrayList<Integer>();
            List<COSObjectable> kids = field.getKids();
            if (kids != null) {
                for (COSObjectable kid : kids) {
                    COSBase kidObject = kid.getCOSObject();
                    if (kidObject instanceof COSDictionary)
                        annotationPages.add(pageNrByAnnotDict.get(kidObject));
                }
            }
    
            Integer mergedPage = pageNrByAnnotDict.get(fieldDict);
    
            if (mergedPage == null)
                if (annotationPages.isEmpty())
                    System.out.printf("i Field '%s' not referenced (invisible).\n", field.getFullyQualifiedName());
                else
                    System.out.printf("a Field '%s' referenced by separate annotation on %s.\n", field.getFullyQualifiedName(), annotationPages);
            else
                if (annotationPages.isEmpty())
                    System.out.printf("m Field '%s' referenced as merged on %s.\n", field.getFullyQualifiedName(), mergedPage);
                else
                    System.out.printf("x Field '%s' referenced as merged on %s and by separate annotation on %s. (Not allowed!)\n", field.getFullyQualifiedName(), mergedPage, annotationPages);
        }
    }
    

    注意,PDFBoxPDAcroForm表单域处理有两个缺点:

    1. PDF 规范允许定义表单的全局对象结构为深树,即实际字段不必是根的直接子级,而是可以通过内部树节点进行组织。 PDFBox 会忽略这一点,并希望这些字段是根的直接子级。

    2. 有些 PDF,最重要的是较旧的 PDF,不包含字段树,而仅通过可视化小部件注释从页面中引用字段对象。 PDFBox 在其PDAcroForm.getFields 列表中没有看到这些字段。

    PS: @mikhailvs in his answer 正确显示您可以使用PDField.getWidget().getPage() 从字段小部件中检索页面对象,并使用catalog.getAllPages().indexOf 确定其页码。虽然速度很快,但这个getPage() 方法有一个缺点:它从小部件注释字典的可选 条目中检索页面引用。因此,如果您处理的 PDF 是由填充该条目的软件创建的,那么一切都很好,但如果 PDF 创建者没有填充该条目,那么您得到的只是一个null 页面。

    PDFBox 2.0.x

    在 2.0.x 中,一些访问相关元素的方法已经改变,但整体情况并未改变,为了安全地检索小部件的页面,您仍然必须遍历页面并找到引用注释的页面。

    安全的方法:

    int determineSafe(PDDocument document, PDAnnotationWidget widget) throws IOException
    {
        COSDictionary widgetObject = widget.getCOSObject();
        PDPageTree pages = document.getPages();
        for (int i = 0; i < pages.getCount(); i++)
        {
            for (PDAnnotation annotation : pages.get(i).getAnnotations())
            {
                COSDictionary annotationObject = annotation.getCOSObject();
                if (annotationObject.equals(widgetObject))
                    return i;
            }
        }
        return -1;
    }
    

    快速方法

    int determineFast(PDDocument document, PDAnnotationWidget widget)
    {
        PDPage page = widget.getPage();
        return page != null ? document.getPages().indexOf(page) : -1;
    }
    

    用法:

    PDAcroForm acroForm = document.getDocumentCatalog().getAcroForm();
    if (acroForm != null)
    {
        for (PDField field : acroForm.getFieldTree())
        {
            System.out.println(field.getFullyQualifiedName());
            for (PDAnnotationWidget widget : field.getWidgets())
            {
                System.out.print(widget.getAnnotationName() != null ? widget.getAnnotationName() : "(NN)");
                System.out.printf(" - fast: %s", determineFast(document, widget));
                System.out.printf(" - safe: %s\n", determineSafe(document, widget));
            }
        }
    }
    

    (DetermineWidgetPage.java)

    (与 1.8.x 代码相比,这里的安全方法只是搜索单个字段的页面。如果在您的代码中您必须确定许多小部件的页面,您应该创建一个查找 Map,如1.8.x 的情况。)

    示例文档

    快速方法失败的文档:aFieldTwice.pdf

    快速方法适用的文档:test_duplicate_field2.pdf

    【讨论】:

      【解决方案2】:

      授予这个答案可能对 OP 没有帮助(一年后),但如果其他人遇到它,这里是解决方案:

      PDDocumentCatalog catalog = doc.getDocumentCatalog();
      
      int pageNumber = catalog.getAllPages().indexOf(yourField.getWidget().getPage());
      

      【讨论】:

      • 如果一个字段在多个页面上有多个widget,你会得到哪些widget的页面?
      • @mkl 这是个好问题。文档说它将获得“作为该字段一部分的单个关联小部件”。不完全清楚您所指的情况会发生什么
      • “作为该字段一部分的单个关联小部件”听起来像是涵盖了将小部件对象合并到字段对象中的情况。仅具有单个小部件的表单字段允许此合并。
      • 是的......我目前在一个项目中正在努力解决这个问题,我遇到了一个 pdf,其中小部件没有与之关联的页面(或其他东西,.getPage() 返回空)
      • 好的,我已经查看了来源。 A getWidget 返回合并到字段字典中的小部件或 Kids 数组中的第一个小部件,如果 Kids为空,则返回 null > 阵列。 B getPage 返回在 P 条目中引用的页面。此条目通常是可选的。因此,null 是每隔一段时间就会发生一次的结果。
      【解决方案3】:

      本示例使用 Lucee (cfml) https://lucee.org/

      非常感谢 mkl,因为他的上述回答非常宝贵,如果没有他的帮助,我无法构建此功能。

      调用函数:pageForSignature(doc, fieldName),它将返回字段名所在的页面编号。如果未找到 fieldName,则返回 -1。

        <cfscript>
        try{
      
        /*
        java is used by using CreateObject()
        */
      
        variables.File = CreateObject("java", "java.io.File");
      
        //references lucee bundle directory - typically on tomcat: /usr/local/tomcat/lucee-server/bundles
        variables.PDDocument = CreateObject("java", "org.apache.pdfbox.pdmodel.PDDocument", "org.apache.pdfbox.app", "2.0.18")
      
        function determineSafe(doc, widget){
      
          var i = '';
          var widgetObject = widget.getCOSObject();
          var pages = doc.getPages();
          var annotation = '';
          var annotationObject = '';
      
          for (i = 0; i < pages.getCount(); i=i+1){
      
          for (annotation in pages.get(i).getAnnotations()){
              if(annotation.getSubtype() eq 'widget'){
                  annotationObject = annotation.getCOSObject();
                  if (annotationObject.equals(widgetObject)){
                      return i;
                  }
              }
          }
      
          }
          return -1;
        }
      
        function pageForSignature(doc, fieldName){
          try{
          var acroForm = doc.getDocumentCatalog().getAcroForm();
          var field = '';
          var widget = '';
          var annotation = '';
          var pageNo = '';
      
          for(field in acroForm.getFields()){
      
          if(field.getPartialName() == fieldName){
      
              for(widget in field.getWidgets()){
      
                 for(annotation in widget.getPage().getAnnotations()){
      
                   if(annotation.getSubtype() == 'widget'){
      
                      pageNo = determineSafe(doc, widget);
                      doc.close();
                      return pageNo;
                   }
                 }
      
              }
          }
        }
      return -1;  
      }catch(e){
          doc.close();
      writeDump(label="catch error",var='#e#');
        }
      } 
      
      doc = PDDocument.init().load(File.init('/**********/myfile.pdf'));
      
      //returns no,  page numbers start at 0
      pageNo = pageForSignature(doc, 'twtzceuxvx');
      
      writeDump(label="pageForSignature(doc, fieldName)", var="#pageNo#");
      </cfscript
      

      【讨论】:

        【解决方案4】:

        单个或多个小部件的通用解决方案(单个页面的重复限定名称)..

        List<PDAnnotationWidget>  widget=field.getWidgets();
        PDDocumentCatalog catalog = doc.getDocumentCatalog();
        for(int i=0;i<widget.size();i++) {
        int pageNumber = 1+ catalog.getPages().indexOf(field.getWidgets().get(i).getPage());
        

        /* 字段坐标也可以到这里单个或多个都可以工作..*/

        //PDRectangle r= widget.get(i).getRectangle();

        }
        

        【讨论】:

        • getPage 返回值的条目是可选的。如果提供来自野外的 PDF,您很可能会获得至少与 PDPage 实例一样频繁的 null。
        • 您使用的是哪个版本..在pdfbox 2.x版本中我没有找到PDDocumentCatalog类的getAllPages方法。
        • 我没有得到任何空值我的 pdf 有 2 页,1 页有一个同名的单选按钮 3 个小部件(字段),在第 2 页复选框有 2 个同名字段..您能否发送您的 pdf 或您获得 null 的场景...对于仅获取页面 no 您可以使用 int pageNumber = 1+ catalog.getPages().indexOf(field.getWidgets().get(0).getPage( ));
        • 我的pdf没有得到任何空值 - 如前所述,该值是可选的。如果您的 PDF 是由添加此值的 PDF 制作者创建的,则您的代码可以工作并且相当快。但其他生产商可能不会增加价值。您可以将您的代码用作快速路线,如果该路线失败,请使用类似于我旧答案中的代码的代码。
        • “您使用的是哪个版本..在 pdfbox 2.x 版本中我没有找到任何方法 getAllPages of PDDocumentCatalog” - 如果您参考我上面的回答,请知道我是在 2014 年 3 月写的,所以可能是 PDFBox 的 1.8.x 版本。在 2.0.x 中,检索所有页面的首选方式是通过 PDDocument.getPages(),参见 Migration to PDFBox 2.0.0。
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2013-01-12
        • 1970-01-01
        • 2015-11-18
        • 1970-01-01
        • 1970-01-01
        • 2020-04-14
        • 1970-01-01
        相关资源
        最近更新 更多