PDFbox 内容流是按页面完成的,但字段来自目录的表单,目录来自 pdf 文档本身。所以我不确定哪些字段在哪些页面上
这样做的原因是 PDF 包含定义表单的全局对象结构。此结构中的表单域可能在 0、1 或更多实际 PDF 页面上有 0、1 或更多可视化。此外,在只有 1 个可视化的情况下,允许合并字段对象和可视化对象。
PDFBox 1.8.x
不幸的是,PDFBox 在其PDAcroForm 和PDField 对象中仅表示此对象结构,并且不提供对相关页面的轻松访问。但是,通过访问底层结构,您可以建立连接。
下面的代码应该清楚地说明如何做到这一点:
@SuppressWarnings("unchecked")
public void printFormFields(PDDocument pdfDoc) throws IOException {
PDDocumentCatalog docCatalog = pdfDoc.getDocumentCatalog();
List<PDPage> pages = docCatalog.getAllPages();
Map<COSDictionary, Integer> pageNrByAnnotDict = new HashMap<COSDictionary, Integer>();
for (int i = 0; i < pages.size(); i++) {
PDPage page = pages.get(i);
for (PDAnnotation annotation : page.getAnnotations())
pageNrByAnnotDict.put(annotation.getDictionary(), i + 1);
}
PDAcroForm acroForm = docCatalog.getAcroForm();
for (PDField field : (List<PDField>)acroForm.getFields()) {
COSDictionary fieldDict = field.getDictionary();
List<Integer> annotationPages = new ArrayList<Integer>();
List<COSObjectable> kids = field.getKids();
if (kids != null) {
for (COSObjectable kid : kids) {
COSBase kidObject = kid.getCOSObject();
if (kidObject instanceof COSDictionary)
annotationPages.add(pageNrByAnnotDict.get(kidObject));
}
}
Integer mergedPage = pageNrByAnnotDict.get(fieldDict);
if (mergedPage == null)
if (annotationPages.isEmpty())
System.out.printf("i Field '%s' not referenced (invisible).\n", field.getFullyQualifiedName());
else
System.out.printf("a Field '%s' referenced by separate annotation on %s.\n", field.getFullyQualifiedName(), annotationPages);
else
if (annotationPages.isEmpty())
System.out.printf("m Field '%s' referenced as merged on %s.\n", field.getFullyQualifiedName(), mergedPage);
else
System.out.printf("x Field '%s' referenced as merged on %s and by separate annotation on %s. (Not allowed!)\n", field.getFullyQualifiedName(), mergedPage, annotationPages);
}
}
注意,PDFBoxPDAcroForm表单域处理有两个缺点:
-
PDF 规范允许定义表单的全局对象结构为深树,即实际字段不必是根的直接子级,而是可以通过内部树节点进行组织。 PDFBox 会忽略这一点,并希望这些字段是根的直接子级。
-
有些 PDF,最重要的是较旧的 PDF,不包含字段树,而仅通过可视化小部件注释从页面中引用字段对象。 PDFBox 在其PDAcroForm.getFields 列表中没有看到这些字段。
PS: @mikhailvs in his answer 正确显示您可以使用PDField.getWidget().getPage() 从字段小部件中检索页面对象,并使用catalog.getAllPages().indexOf 确定其页码。虽然速度很快,但这个getPage() 方法有一个缺点:它从小部件注释字典的可选 条目中检索页面引用。因此,如果您处理的 PDF 是由填充该条目的软件创建的,那么一切都很好,但如果 PDF 创建者没有填充该条目,那么您得到的只是一个null 页面。
PDFBox 2.0.x
在 2.0.x 中,一些访问相关元素的方法已经改变,但整体情况并未改变,为了安全地检索小部件的页面,您仍然必须遍历页面并找到引用注释的页面。
安全的方法:
int determineSafe(PDDocument document, PDAnnotationWidget widget) throws IOException
{
COSDictionary widgetObject = widget.getCOSObject();
PDPageTree pages = document.getPages();
for (int i = 0; i < pages.getCount(); i++)
{
for (PDAnnotation annotation : pages.get(i).getAnnotations())
{
COSDictionary annotationObject = annotation.getCOSObject();
if (annotationObject.equals(widgetObject))
return i;
}
}
return -1;
}
快速方法
int determineFast(PDDocument document, PDAnnotationWidget widget)
{
PDPage page = widget.getPage();
return page != null ? document.getPages().indexOf(page) : -1;
}
用法:
PDAcroForm acroForm = document.getDocumentCatalog().getAcroForm();
if (acroForm != null)
{
for (PDField field : acroForm.getFieldTree())
{
System.out.println(field.getFullyQualifiedName());
for (PDAnnotationWidget widget : field.getWidgets())
{
System.out.print(widget.getAnnotationName() != null ? widget.getAnnotationName() : "(NN)");
System.out.printf(" - fast: %s", determineFast(document, widget));
System.out.printf(" - safe: %s\n", determineSafe(document, widget));
}
}
}
(DetermineWidgetPage.java)
(与 1.8.x 代码相比,这里的安全方法只是搜索单个字段的页面。如果在您的代码中您必须确定许多小部件的页面,您应该创建一个查找 Map,如1.8.x 的情况。)
示例文档
快速方法失败的文档:aFieldTwice.pdf
快速方法适用的文档:test_duplicate_field2.pdf