【问题标题】:Pytesseract does not detect me numbersPytesseract 没有检测到我的数字
【发布时间】:2021-01-24 23:28:54
【问题描述】:

我正在制作一个简单的程序来使用 python 和 pytesseract 检测图像中的数字,但情况是它总是返回我♀,我正在分析这样的图像:

my image

我读取数字的代码如下:

import pytesseract
from pytesseract import (
    Output,
    TesseractError,
    TesseractNotFoundError,
    TSVNotSupported,
    get_tesseract_version,
    image_to_boxes,
    image_to_data,
    image_to_osd,
    image_to_pdf_or_hocr,
    image_to_string,
    run_and_get_output
)

def analizar_resultado(path): 
    image = cv2.imread(path, 1)
    
    text = pytesseract.image_to_string(image, config = 'digits')
    print('texto detectado:', text)

但我无法让它为我工作,我已经尝试了更多质量更好的此类图像和其他图像,但我无法获得任何数字,我该如何解决这个问题?非常感谢

【问题讨论】:

  • 想要改进 Tesseract 文本识别? Goggle for tesseract 提高识别率
  • 我只想检测数字,但你对谷歌的 tesseract 是什么意思?谢谢
  • 搜索 Tesseract 提高识别率
  • 你现在还有什么我可以尝试的吗?其他ocr或类似的东西?谢谢
  • 在 StackOverflow 上请求库是题外话。

标签: python artificial-intelligence text-processing python-tesseract text-recognition


【解决方案1】:

我有一个三步解决方案


    1. 分别获取每个数字
    1. 应用阈值
    1. 读取输出

第 1 部分:分别获取每个数字

  • 您可以通过使用索引变量来获取每个数字。例如:

    • s_idx = 0  # start index
      e_idx = int(w/5) - 10  # end index
      
  • 先获取图片的高度和宽度,然后对于每个数字,增加索引

    • for _ in range(0, 6):
          gry_crp = gry[0:h, s_idx:e_idx]
          s_idx = e_idx
          e_idx = s_idx + int(w/5) - 20
      
  • 结果

    • 0 0 9 9 7 6
  • 第 2 部分:应用阈值

    • 0 0 9 9 7 6
  • 第 3 部分:阅读

    • 0.9976
      

很遗憾,由于伪影,第二个零无法识别为数字。

如果看不懂图片,换个psmconfigurations试试

代码:


import cv2
from pytesseract import image_to_string

img = cv2.imread("A3QRw.png")
gry = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
(h, w) = gry.shape[:2]
s_idx = 0  # start index
e_idx = int(w/5) - 10  # end index

result = []

for i, _ in enumerate(range(0, 6)):
    gry_crp = gry[0:h, s_idx:e_idx]
    (h_crp, w_crp) = gry_crp.shape[:2]
    gry_crp = cv2.resize(gry_crp, (w_crp*3, h_crp*3))
    thr = cv2.threshold(gry_crp, 0, 255,
                        cv2.THRESH_BINARY + cv2.THRESH_OTSU)[1]
    txt = image_to_string(thr, config="--psm 6 digits")
    result.append(txt[0])
    s_idx = e_idx
    e_idx = s_idx + int(w/5) - 20
    cv2.imshow("thr", thr)
    cv2.waitKey(0)

print("".join([digit for digit in result]))

【讨论】:

  • 但是为什么 e_index 首先是-10,然后是-20?
  • 将每个单独的图像对齐到中心。
  • 有 6 张图片,所以我认为将 width 分成 5 个单独的标签会给我 6 张图片。
  • 为什么在调整大小时将 w*3 和 h*3 相乘?为什么 3
  • 调整图像大小有利于获得准确的结果。 3 只是获得所需结果的数字。您可以使用任何其他正数
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2021-03-12
  • 1970-01-01
  • 2021-03-07
  • 2021-12-12
  • 2021-07-12
  • 1970-01-01
  • 2021-08-06
相关资源
最近更新 更多