【问题标题】:Pytesseract random bug when reading textPytesseract 阅读文本时出现随机错误
【发布时间】:2019-09-20 06:33:22
【问题描述】:

我正在为视频游戏创建一个机器人,我必须阅读屏幕上显示的一些信息。鉴于信息总是在同一个位置,我可以截图并将图片裁剪到正确的位置。

90% 的情况下,识别是完美的,但有时它会返回看起来完全随机的东西(参见下面的示例)。

我试过把图片转成黑白没有成功,并尝试改变pytesseract的配置(config = ("-l fra --oem 1 --psm 6"))

def readScreenPart(x,y,w,h):
    monitor = {"top": y, "left": x, "width": w, "height": h}
    output = "monitor.png"
    with mss.mss() as sct:
        sct_img = sct.grab(monitor)        
        mss.tools.to_png(sct_img.rgb, sct_img.size, output=output)

    img = cv2.imread("monitor.png")
    img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    cv2.imwrite("result.png", img)
    config = ("-l fra --oem 1 --psm 6")

    return pytesseract.image_to_string(img,config=config)

例子:这张图片产生了一个bug,它返回字符串“IRPMV/LEIILK”

另一张图片

现在我不知道问题出在哪里,因为它不仅仅是一个错误的字符,而是一个完全随机的结果..

感谢您的帮助

【问题讨论】:

  • 另一个例子,这个返回 "plooWWLEIlÿ 3" (i.imgur.com/IhWeT0F.png)
  • Pytesseract 使用程序tesseract,该程序旨在识别白纸上带有黑色文本的扫描文档。在tesseract 页面上,您甚至可以找到有关如何在白色背景上创建更好的黑色文本以使用tesseract 获得更好结果的信息。您的代码在深灰色背景上创建浅灰色文本,因此可能不足以正确识别文本。
  • tesseract documentation: Improving the quality of the output。在反转图像中,您可以阅读:虽然 tesseract 3.05 版(及更早版本)可以毫无问题地处理反转图像(深色背景和浅色文本),但 4.x 版本在浅色背景上使用深色文本。
  • 您可以使用img = 255 - img 反转您的图像。
  • 我用你的例子运行代码,我得到了正确的结果。甚至我也不必转换为灰度。 PyTesseract 0.2.7 / Tesseract 4.0.0-beta.1 / Python 3.7.4 / Linux Mint 19.2。

标签: python image opencv ocr python-tesseract


【解决方案1】:

预处理是将图像投入 Pytesseract 之前的重要步骤。通常,您希望所需文本为黑色,背景为白色。目前,您的前景文本是绿色而不是白色。这是修复格式的简单过程

  • 将图像转换为灰度
  • Otsu 获取二值图像的阈值
  • 反转图像

原图

大津的门槛

反转图像

Pytesseract 的输出

122 活力

其他图片

200 活力

在反转图像之前,最好执行morphological operations 以平滑/过滤文本。但是对于您的图像,文本不需要额外的平滑

import cv2
import pytesseract

pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe"

image = cv2.imread('3.png',0)
thresh = cv2.threshold(image, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[1]
result = 255 - thresh

data = pytesseract.image_to_string(result, lang='eng',config='--psm 6')
print(data)

cv2.imshow('thresh', thresh)
cv2.imshow('result', result)
cv2.waitKey()

【讨论】:

  • 效果很好,感谢您的帮助 :)
【解决方案2】:

正如评论所说,这与您的文本和背景颜色有关。对于深色背景上的浅色文本,Tesseract 基本上是无用的,这是我在将其提供给 tesseract 之前应用于任何文本图像的几行:

# convert color image to grayscale
grayscale_image = cv2.cvtColor(your_image, cv2.COLOR_BGR2GRAY)

# Otsu Tresholding method find perfect treshold, return an image with only black and white pixels
_, binary_image = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU)

# we just don't know if the text is in black and background in white or vice-versa
# so we count how many black pixels and white pixels there are
count_white = numpy.sum(binary > 0)
count_black = numpy.sum(binary == 0)

# if there are more black pixels than whites, then it's the background that is black so we invert the image's color
if count_black > count_white:
    binary_image = 255 - binary_image

black_text_white_background_image = binary_image

现在,无论原始图像是哪种颜色,您都可以肯定在白色背景上有黑色文本,而且如果字符的高度为 35 像素,Tesseract 也是(奇怪地)最有效的,较大的字符不会显着降低准确度,但只需缩短几个像素就会使 tesseract 无用!

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-08-27
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多