【问题标题】:Tesseract not able to recognize characters even for a high quality Image即使是高质量图像,Tesseract 也无法识别字符
【发布时间】:2014-07-24 08:44:12
【问题描述】:

我正在使用leptonica进行清理和图像处理,然后将其传递给tesseract进行OCR。但是即使图像质量很高,它也无法识别字符。图像规格如下。

1 bpp, uncompressed, 1280 * 960 , 300dpi horizontal and vertical resolution

以下是我使用 leptonica 依次进行的图像处理操作

pixConvertTo8
pixBackgroundNormSimple
pixOtsuAdaptiveThreshold
pixContrastTRC {Regarding this - I am passing high values like 1.0 or even 5.0 but image doesnt really change}
pixFindSkew
pixRotate { rotate by angle found by pixFindSkew}
pixRotate90 {do this 4 times to read image in all 4 orientations}
pixClipRectangle {crop image}
Finally tesseract command

我在输出中得到垃圾字符。示例输入图像如下。

我得到的输出如下

Final K-1
II]
s h d | K-1 ,.,
(F°o.~?n‘i&1) 5/>.©12 mm E2‘;
Deparlrnenl of tho Treasury , ,
I 1 I l I
‘mama, Ravenuo SGMW For cnlundm your 201), ‘ " °F°$ "'100fTIO
or lax yum boqmnnnq 7 _ 20\Q_
‘ 7660
and ondmg _  W vv I go
Beneï¬ciary's Share of Income, Deductions,
cl'editS, etc. F 800 buck 01 loam nnd lnstruoflons»
___lnformatI0n About mo Estate or Trust
‘ Ordmary d|v|dm
i 12113
 _
‘; Quahfmd dlVIdG
\ 8132
3 1
Net shun-term
A Estate's at trust's omgiuym ldonnlmnluon numbol
56-0987654
B Estate's u trust‘: namo
ESTATE OF MARTHA SMITH
0 Fiduc§ary's name, address, clly, smlu‘ and /IP codo
N01 long~lerm c
\ 24043 
u 
‘ 28% vale gann
Ti
Unreptumd 5
Omar porfloho 4
nonbuslness lfll
/\..4........ L. ._.._ ,.

我应该怎么做才能提高准确性。

第 2 部分:

我尝试关注this link。并创建了一个 eng.user-words.traineddata 文件和 bazaar.train 文件并尝试使用“bazaar”作为附加参数运行。但我得到“read_params_file:无法打开市场错误”。 有什么建议么?

【问题讨论】:

  • 嗨@nnm,您是否获得了有关使用 tesseract 处理这些税务文件的进一步帮助?我们需要在我们的一个项目中实现相同的要求。

标签: image-processing tesseract leptonica


【解决方案1】:

对于第一部分,

我不知道您在此处发布的图像是否是您尝试扫描的实际图像,但当我尝试时,我得到了这个:-


财政部国税局

对于 cnlundm 你的 V019, 1 ‘ '"l0T°5' |nC0m0

or tax yam boqlnnlnq , 2o12_ ‘ 7660 and ondlng I go 2: ‘ 普通 dlvndm " "T ' x 12113

1;合格 dwnda ' 8132 Netshun-term:

M 不长~terrn c

i 24043 Ab ‘ 2896 ralagann

受益人的收入份额、扣除额、Cfedits 等 5 800 back oi 形成 nnd 指令

| Partl 关于州或信托的信息

A Estate 或 IvLsl 的 omuoym Idonnlncnluon numhu

56-0987654

8 房地产‘:信托’:namo

玛莎·史密斯庄园

M: Unreptumd 5

017161 portioho : nonbuslness Inl

C Fiduc§ary 的姓名、地址、城市、smlul an-(V1/If’Eooo


这不是很好,但它似乎比你得到的要好一些。我在 Windows 上使用 Tesseract v3。 我的基本命令是:

-    tesseract.exe  nnm.tif  nnm

对于第二部分,

您的bazaar 文件应位于configs 文件夹中

 .....\Tesseract-OCR\tessdata\configs\bazaar

并且有一些要求以特定格式保存它,例如UTF8,行尾只有LF而不是CR + LF,文件格式似乎相当挑剔。

你可以从http://code.metager.de/source/raw/google/tesseract-ocr/tessdata/configs/bazaar得到一份副本

我制作了一个数字配置文件,用于扫描一些我只对数字感兴趣的图像并且效果很好:

-   tesseract.exe  scanfile.jpg  scanfile  digits 

Tesseract 的文档很差,而且在 PC 上运行不佳。

【讨论】:

    【解决方案2】:

    对于第一部分,

    我认为您应该考虑 Capture2Text 完成的预处理。它同时使用 Leptonica 和 Tesseract 对图像进行 OCR。

    我不确定第 2 部分。

    【讨论】:

      猜你喜欢
      • 2014-12-21
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-03-26
      • 2020-07-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多