【问题标题】:Using Java PDFBox library to write Russian PDF使用 Java PDFBox 库编写俄语 PDF
【发布时间】:2009-11-11 08:09:58
【问题描述】:

我正在使用一个名为 PDFBox 的 Java 库尝试将文本写入 PDF。它非常适合英文文本,但是当我尝试在 PDF 中写俄文文本时,字母看起来很奇怪。似乎问题出在使用的字体上,但我不太确定,所以我希望有人能指导我完成这个。这是重要的代码行:

PDTrueTypeFont font = PDTrueTypeFont.loadTTF( pdfFile, new File( "fonts/VREMACCI.TTF" ) );  // Windows Russian font imported to write the Russian text.
font.setEncoding( new WinAnsiEncoding() );  // Define the Encoding used in writing.
// Some code here to open the PDF & define a new page.
contentStream.drawString( "отделом компьютерной" ); // Write the Russian text.

WinAnsiEncoding 源代码为:Click here

--------------------- 2009 年 11 月 18 日编辑

经过一番调查,我现在确定这是一个编码问题,这可以通过使用名为 DictionaryEncoding 的有用 PDFBox 类定义我自己的编码来解决。

我不知道如何使用它,但这是我迄今为止尝试过的:

COSDictionary cosDic = new COSDictionary();
cosDic.setString( COSName.getPDFName("Ercyrillic"), "0420 " ); // Russian letter.
font.setEncoding( new DictionaryEncoding( cosDic ) );

这不起作用,因为我似乎以错误的方式填写字典,当我使用它编写 PDF 页面时,它显示为空白。

DictionaryEncoding 源码为:Click here

【问题讨论】:

标签: java pdf encoding


【解决方案1】:

长话短说——为了从 TrueType 字体以 PDF 格式输出 unicode,输出必须包含大量详细且看似多余的信息。它归结为 - 在 TrueType 字体中,字形存储为字形 ID。这些字形 id 与特定的 unicode 字符相关联(和 IIRC,内部的 unicode 字形可能指代几个代码点 - 比如 é 指代 e 和一个尖锐的重音 - 我的记忆很模糊)。除了说存在从字符串中的 UTF16BE 值到 TrueType 字体中的字形 id 的映射以及从 UTF16BE 值到 Unicode 的映射(即使它是标识)之外,PDF 并没有真正的 unicode 支持。

  • 子类型 Type0 的字体字典
    • 一个 DescendantFonts 数组,其条目如下所述
    • 将 UTF16BE 值映射到 unicode 的 ToUnicode 条目
    • 编码设置为 Identity-H

我在自己的工具上进行的一项单元测试的输出如下所示:

13 0 obj
<< 
   /BaseFont /DejaVuSansCondensed 
   /DescendantFonts [ 4 0 R  ]   
   /ToUnicode 14 0 R 
   /Type /Font 
   /Subtype /Type0 
   /Encoding /Identity-H 
>> endobj

14 0 obj
<< /Length 346 >> stream
/CIDInit /ProcSet findresource begin 12 dict begin begincmap /CIDSystemInfo <<
/Registry (Adobe) /Ordering (UCS) /Supplement 0 >> def /CMapName /Adobe-Identity-UCS
def /CMapType 2 def 1 begincodespacerange <0000> <FFFF> endcodespacerange 1
beginbfrange <0000> <FFFF> <0000> endbfrange endcmap CMapName currentdict /CMap
defineresource pop end end

endstream % 注意流的格式错误

  • 子类型 CIDFontTYpe2 的字体字典
    • CIDSsytemInfo
    • 一个字体描述符
    • DW 和 W
    • 从字符 ID 映射到字形 ID 的 CIDToGIDMap

这是来自同一测试的一个 - 这是 DescendantFonts 数组中的对象:

4 0 obj
<< 
   /Subtype /CIDFontType2 
   /Type /Font 
   /BaseFont /DejaVuSansCondensed 
   /CIDSystemInfo 8 0 R 
   /FontDescriptor 9 0 R 
   /DW 1000 
   /W 10 0 R 
   /CIDToGIDMap 11 0 R 
>>

8 0 obj
<< 
   /Registry (Adobe)
   /Ordering (UCS)
   /Supplement 0 
>>
endobj

我为什么要告诉你这个?它与 PDFBox 有什么关系?就是这样:坦率地说,PDF 中的 Unicode 输出是一件让人头疼的事情。 Acrobat 是在 Unicode 出现之前开发的,从一开始就没有 Unicode 的 CJK 编码很痛苦(我知道 - 我当时在 Acrobat 上工作)。后来添加了对 Unicode 的支持,但它真的感觉像是被蒙上了一层阴影。有人会希望您只说 /Encoding /Unicode 并拥有以 thorn 和 y-dieresis 字符开头的字符串,然后就可以了。没有这样的运气。如果你没有把每一个细节都放进去(实际上,Acrobat,嵌入了一个 PostScript 程序来翻译成 Unicode?WTH?),你​​会在 Acrobat 中得到一个空白页。我发誓,这不是我编造的。

此时,我为一家单独的公司编写 PDF 生成工具(现在是 .NET,所以它对您没有帮助),并且我将隐藏所有废话作为设计目标。所有文本都是 unicode - 如果您只使用与 WinAnsi 相同的字符代码,那就是您在后台得到的。使用其他任何东西,你会得到所有其他的东西。如果 PDFBox 为您工作,我会感到惊讶 - 这是一个严重的麻烦。

【讨论】:

    【解决方案2】:

    解决方案非常简单。

    1) 您必须找到与要显示的字符兼容的字体。
    2) 将字体的.ttf文件下载到本地。
    3) 从您的应用程序加载字体

    例如,如果您想使用希腊字符,您必须这样做:

    content = new PDPageContentStream(document, page);
    pdfFont = PDType0Font.load( document, new File( "arialuni.ttf" ) )
    content.setFont(pdfFont, fontSize);
    

    【讨论】:

      【解决方案3】:

      尝试使用这种结构:

      PDFont font = PDType0Font.load( pdfFile, new File( "fonts/VREMACCI.TTF" ) );  // Windows Russian font imported to write the Russian text.
      // Some code here to open the PDF & define a new page.
      contentStream.beginText();
      contentStream.setFont(font, 12);
      contentStream.showText( "отделом компьютерной" ); // Write the Russian text.
      contentStream.endText();
      

      【讨论】:

      • 当时我问了这个问题 PDFBox 只支持写拉丁语,所以不支持写俄语。现在创建 PDFBox 的人已经解决了这个问题,它现在支持俄语和其他语言,感谢他们,感谢您分享解决方案:)
      【解决方案4】:

      也许需要编写俄语编码类,我想它应该看起来像WinAnsiEncoding 之一。
      现在,我不知道该放什么!

      或者,如果您还没有这样做,也许您应该将源文件编码为 UTF-8 并使用默认编码。
      我看到了一些与从现有 PDF 文件中提取俄语文本相关的消息(当然使用 PDFBox),但我不知道输出是否相关。
      您也可以写信给 PDFBox 邮件列表。

      【讨论】:

      • 好吧,使用 PDFBox 提取俄语文本可以正常工作,问题在于在 PDF 中编写俄语文本。
      • 对于编写俄语编码,有 DictionaryEncoding 类,我认为可以让我定义自己的编码......但对我来说似乎是一个迷宫:kickjava.com/src/org/pdfbox/encoding/…
      【解决方案5】:

      测试这是否是一个编码问题应该很容易做到(只需切换到 UTF16 编码)。

      我假设您已尝试使用带有 VREMACCI 字体的编辑器或其他工具,并确认它以您期望的方式显示?

      您可能想尝试在 iText 中做同样的事情,只是为了了解问题是否与 PdfBox 库本身有关...如果您的主要目标是生成 PDF 文件,那么 iText 可能是一个更好的解决方案.

      编辑 - 对 cme​​ts 的长回答:

      好的-对于编码问题的来回感到抱歉...您的核心问题(您可能已经知道)是写入内容流的字节编码与用于查看的编码不同向上字形。现在我会尽力提供帮助:

      我查看了 PdfBox 中的字典编码类,它看起来很不直观……有问题的“字典”是 PDF 字典。所以你基本上需要做的是创建一个 Pdf 字典对象(我认为 PdfBox 将其称为 COSObject 的一种),然后向其中添加条目。

      字体的编码在 PDF 中定义为字典(参见上述规范的第 266 页)。该字典包含一个基本编码名称,以及一个可选的差异数组。从技术上讲,差异数组不应该与真字体一起使用(尽管我已经看到它在某些情况下使用过 - 但不要使用它)。

      然后,您将为编码的 cmap 指定一个条目。此 cmap 将是您字体的编码。

      我的建议是采用现有的 PDF 来满足您的需求,然后获取该字体的字典结构转储,以便您查看它的外观。

      这绝对不适合胆小的人。我可以提供一些帮助 - 如果您需要字典转储,请给我一个带有示例 PDF 的超链接,我将通过我在 iText 开发中使用的一些算法运行它(我是 iText 文本提取子的维护者-系统)。

      编辑 - 2009 年 11 月 17 日

      好的 - 这是来自 russian.pdf 文件的字典转储(子字典以缩进方式列出,并按照它们在包含字典中出现的顺序):

      (/CropBox=[0, 0, 595, 842], /Parent=Dictionary of type: /Pages, /Type=/Page, /Contents=[209 0 R, 210 0 R, 211 0 R, 214 0 R, 215 0 R, 216 0 R, 222 0 R, 223 0 R], /Resources=Dictionary, /MediaBox=[0, 0, 595, 842], /StructParents=0, /Rotate=0)
          Subdictionary /Parent = (/Type=/Pages, /Count=6, /Kids=[195 0 R, 1 0 R, 3 0 R, 5 0 R, 7 0 R, 9 0 R])
          Subdictionary /Resources = (/ExtGState=Dictionary, /ProcSet=[/PDF, /Text], /ColorSpace=Dictionary, /Font=Dictionary, /Properties=Dictionary)
              Subdictionary /ExtGState = (/GS0=Dictionary of type: /ExtGState)
                  Subdictionary /GS0 = (/OPM=1, /op=false, /Type=/ExtGState, /SA=false, /OP=false, /SM=0.02)
              Subdictionary /ColorSpace = (/CS0=[/ICCBased, 228 0 R])
              Subdictionary /Font = (/C2_1=Dictionary of type: /Font, /C2_2=Dictionary of type: /Font, /C2_3=Dictionary of type: /Font, /C2_4=Dictionary of type: /Font, /TT2=Dictionary of type: /Font, /TT1=Dictionary of type: /Font, /TT0=Dictionary of type: /Font, /C2_0=Dictionary of type: /Font, /TT3=Dictionary of type: /Font)
                  Subdictionary /C2_1 = (/DescendantFonts=[243 0 R], /BaseFont=/LDMIEC+TimesNewRomanPS-BoldMT, /Type=/Font, /Subtype=/Type0, /Encoding=/Identity-H, /ToUnicode=Stream)
                  Subdictionary /C2_2 = (/DescendantFonts=[233 0 R], /BaseFont=/LDMIBO+TimesNewRomanPSMT, /Type=/Font, /Subtype=/Type0, /Encoding=/Identity-H, /ToUnicode=Stream)
                  Subdictionary /C2_3 = (/DescendantFonts=[224 0 R], /BaseFont=/LDMIHD+TimesNewRomanPS-ItalicMT, /Type=/Font, /Subtype=/Type0, /Encoding=/Identity-H, /ToUnicode=Stream)
                  Subdictionary /C2_4 = (/DescendantFonts=[229 0 R], /BaseFont=/LDMIDA+Tahoma, /Type=/Font, /Subtype=/Type0, /Encoding=/Identity-H, /ToUnicode=Stream)
                  Subdictionary /TT2 = (/LastChar=58, /BaseFont=/LDMIFC+TimesNewRomanPS-BoldMT, /Type=/Font, /Subtype=/TrueType, /Encoding=/WinAnsiEncoding, /Widths=[250, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 250, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 333], /FontDescriptor=Dictionary of type: /FontDescriptor, /FirstChar=32)
                      Subdictionary /FontDescriptor = (/Type=/FontDescriptor, /StemV=136, /Descent=-216, /FontWeight=700, /FontBBox=[-558, -307, 2000, 1026], /CapHeight=656, /FontFile2=Stream, /FontStretch=/Normal, /Flags=34, /XHeight=0, /FontFamily=Times New Roman, /FontName=/LDMIFC+TimesNewRomanPS-BoldMT, /Ascent=891, /ItalicAngle=0)
                  Subdictionary /TT1 = (/LastChar=187, /BaseFont=/LDMICP+TimesNewRomanPSMT, /Type=/Font, /Subtype=/TrueType, /Encoding=/WinAnsiEncoding, /Widths=[250, 0, 0, 0, 0, 833, 778, 0, 333, 333, 0, 0, 250, 333, 250, 278, 500, 500, 500, 500, 500, 500, 500, 500, 500, 500, 278, 278, 0, 564, 0, 444, 0, 722, 667, 667, 722, 611, 556, 0, 722, 333, 389, 0, 611, 889, 722, 722, 556, 0, 667, 556, 611, 0, 722, 944, 0, 722, 0, 333, 0, 333, 0, 500, 0, 444, 500, 444, 500, 444, 333, 500, 500, 278, 0, 500, 278, 778, 500, 500, 500, 0, 333, 389, 278, 500, 500, 722, 0, 500, 444, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 500, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 500, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 500], /FontDescriptor=Dictionary of type: /FontDescriptor, /FirstChar=32)
                      Subdictionary /FontDescriptor = (/Type=/FontDescriptor, /StemV=82, /Descent=-216, /FontWeight=400, /FontBBox=[-568, -307, 2000, 1007], /CapHeight=656, /FontFile2=Stream, /FontStretch=/Normal, /Flags=34, /XHeight=0, /FontFamily=Times New Roman, /FontName=/LDMICP+TimesNewRomanPSMT, /Ascent=891, /ItalicAngle=0)
                  Subdictionary /TT0 = (/LastChar=55, /BaseFont=/LDMIBN+TimesNewRomanPS-BoldItalicMT, /Type=/Font, /Subtype=/TrueType, /Encoding=/WinAnsiEncoding, /Widths=[250, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 250, 0, 500, 500, 500, 0, 0, 0, 0, 500], /FontDescriptor=Dictionary of type: /FontDescriptor, /FirstChar=32)
                      Subdictionary /FontDescriptor = (/Type=/FontDescriptor, /StemV=116.867004, /Descent=-216, /FontWeight=700, /FontBBox=[-547, -307, 1206, 1032], /CapHeight=656, /FontFile2=Stream, /FontStretch=/Normal, /Flags=98, /XHeight=468, /FontFamily=Times New Roman, /FontName=/LDMIBN+TimesNewRomanPS-BoldItalicMT, /Ascent=891, /ItalicAngle=-15)
                  Subdictionary /C2_0 = (/DescendantFonts=[238 0 R], /BaseFont=/LDMHPN+TimesNewRomanPS-BoldItalicMT, /Type=/Font, /Subtype=/Type0, /Encoding=/Identity-H, /ToUnicode=Stream)
                  Subdictionary /TT3 = (/LastChar=169, /BaseFont=/LDMIEB+Tahoma, /Type=/Font, /Subtype=/TrueType, /Encoding=/WinAnsiEncoding, /Widths=[313, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 546, 0, 546, 0, 0, 546, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 929], /FontDescriptor=Dictionary of type: /FontDescriptor, /FirstChar=32)
                      Subdictionary /FontDescriptor = (/Type=/FontDescriptor, /StemV=92, /Descent=-206, /FontWeight=400, /FontBBox=[-600, -208, 1338, 1034], /CapHeight=734, /FontFile2=Stream, /FontStretch=/Normal, /Flags=32, /XHeight=546, /FontFamily=Tahoma, /FontName=/LDMIEB+Tahoma, /Ascent=1000, /ItalicAngle=0)
              Subdictionary /Properties = (/MC0=Dictionary of type: /OCMD)
                  Subdictionary /MC0 = (/Type=/OCMD, /OCGs=Dictionary of type: /OCG)
                      Subdictionary /OCGs = (/Usage=Dictionary, /Type=/OCG, /Name=HeaderFooter)
                          Subdictionary /Usage = (/CreatorInfo=Dictionary, /PageElement=Dictionary)
                              Subdictionary /CreatorInfo = (/Creator=Acrobat PDFMaker 6.0 äëÿ Word)
                              Subdictionary /PageElement = (/SubType=/HF)
      

      这里有很多活动部件。你可能想把一个只有 3 或 4 个字符的测试文档放在一起……这里使用了很多 type-1 字体(除了 TT 字体),所以很难说您的特定问题涉及什么。

      (你确定你至少不想用 iText 试试这个吗?;-) 我并不是说它会起作用,只是它可能值得一试)。

      作为参考,上面的字典转储是使用 com.lowagie.text.pdf.parser.PdfContentReaderTool 类获得的

      【讨论】:

      • PDFBox 中没有支持 UTF8 或 UTF16 的类,但我认为是的,这是编码问题。我知道 iText 是一个很棒的库,但我已经开始使用 PDFBox 了,到目前为止它还不错,所以我想坚持使用 PDFBox。
      • 呃。如果您使用 PDFBox 来解析您创建的内容,您是否能够恢复文本?如果是这样,那么它可能不是编码的限制,预先...也许这只是 PDFBox 如何将字节元组映射到字形的问题?
      • 恢复它是什么意思?我可以写一些其他的外语,比如法语、德语……但是像俄语这样的其他语言似乎是个问题。这是一个编码问题,我敢肯定。并且创建了 DictionaryEncoding 类以允许扩展其他不受支持的编码,但我仍然不知道如何使用它。
      • 好吧,如果你使用 PDFBox 解析文本,你得到的是你输入的文本,还是被咀嚼了?换句话说,将文本 A 写入 PDF,然后从 PDF 中读取文本 A,然后查看 A?=A。如果这是一个编码问题,它不太可能是对称的,所以你很可能会得到 A!=B 出来。如果您确实得到 A=A,那么问题可能不是编码,您正在处理字符代码->字形转换问题。强烈建议您使用 iText 进行尝试,这样您至少有一个您应该获得的内容流的基线。
      • 好吧,我试过了,它得到的文本与我输入的文本相同,因此 A=A 返回 true。另一方面,我不明白编码和字形转换之间的区别......我认为它们是同一回事。当我与图书馆管理员交谈时,他是这么说的:“问题是要添加的字符串和字体编码之间的映射。AFAIK WinAnsiEncoding,它被用作真实类型字体的默认值,不包含俄罗斯字母. 所以最后你必须找到另一种映射方式。你应该能够使用 DictionaryEncoding 定义自己的映射。"
      【解决方案6】:

      试试这个:

      Phrase leftTitle = new Phrase("САНКТ-ПЕТЕРБУРГ", FontFactory.getFont("Tahoma", "Cp1251", true, 25));

      这至少适用于最新的 (5.0.1) iText

      【讨论】:

        猜你喜欢
        • 2012-12-11
        • 1970-01-01
        • 2022-08-02
        • 1970-01-01
        • 2011-10-15
        • 1970-01-01
        • 2013-07-15
        • 2012-10-27
        • 2015-03-23
        相关资源
        最近更新 更多