【问题标题】:OCR with python and tesseract使用 python 和 tesseract 进行 OCR
【发布时间】:2017-09-02 09:36:22
【问题描述】:

OCr 示例代码

from PIL import Image
from pytesser import *

image_file = 'E:\Downloads\menu.jpg'
im = Image.open(image_file)
text = image_to_string(im)
text = image_file_to_string(image_file)
text = image_file_to_string(image_file, graceful_errors=True)
print ("=====output=======\n")
print (text)

错误

  File "C:\Users\XXX\Anaconda3\lib\site-packages\pytesser\__init__.py", line 64
    except errors.Tesser_General_Exception, value:
                                          ^
SyntaxError: invalid syntax

.I was following this tutorial on python and OCR using tesseract

我正在使用 python 3,我已经下载了 tesseract 库并添加到 anaconda 库中。但是在第一次运行它时,我收到了错误,显示打印缺少括号,所以我改变了它,现在我发现这个错误任何人都可以帮助我这会很棒。 我还在这里添加了 tesseract 包装器的源代码 """使用 Google 的 Tesseract 引擎在 Python 中进行 OCR

http://code.google.com/p/pytesser/
by Michael J.T. O'Kelly
V 0.0.1, 3/10/07"""

PIL import Image
import subprocess

import util
import errors

tesseract_exe_name = 'C:\Users\SACHIN\Anaconda3\Lib\site-packages\pytesser\\tesseract' # Name of executable to be called at command line
scratch_image_name = "temp.bmp" # This file must be .bmp or other Tesseract-compatible format
scratch_text_name_root = "temp" # Leave out the .txt extension
cleanup_scratch_flag = True  # Temporary files cleaned up after OCR operation

def call_tesseract(input_filename, output_filename):
    """Calls external tesseract.exe on input file (restrictions on types),
    outputting output_filename+'txt'"""
    args = [tesseract_exe_name, input_filename, output_filename]
    proc = subprocess.Popen(args)
    retcode = proc.wait()
    if retcode!=0:
        errors.check_for_errors()

def image_to_string(im, cleanup = cleanup_scratch_flag):
    """Converts im to file, applies tesseract, and fetches resulting text.
    If cleanup=True, delete scratch files after operation."""
    try:
        util.image_to_scratch(im, scratch_image_name)
        call_tesseract(scratch_image_name, scratch_text_name_root)
        text = util.retrieve_text(scratch_text_name_root)
    finally:
        if cleanup:
            util.perform_cleanup(scratch_image_name, scratch_text_name_root)
    return text

def image_file_to_string(filename, cleanup = cleanup_scratch_flag, graceful_errors=True):
    """Applies tesseract to filename; or, if image is incompatible and graceful_errors=True,
    converts to compatible format and then applies tesseract.  Fetches resulting text.
    If cleanup=True, delete scratch files after operation."""
    try:
        try:
            call_tesseract(filename, scratch_text_name_root)
            text = util.retrieve_text(scratch_text_name_root)
        except errors.Tesser_General_Exception:
            if graceful_errors:
                im = Image.open(filename)
                text = image_to_string(im, cleanup)
            else:
                raise
    finally:
        if cleanup:
            util.perform_cleanup(scratch_image_name, scratch_text_name_root)
    return text


if __name__=='__main__':
    im = Image.open('phototest.tif')
    text = image_to_string(im)
    print (text)
    try:
        text = image_file_to_string('fnord.tif', graceful_errors=False)
    except errors.Tesser_General_Exception, value:
        print "fnord.tif is incompatible filetype.  Try graceful_errors=True"
        print value
    text = image_file_to_string('fnord.tif', graceful_errors=True)
    print ("fnord.tif contents:", text)
    text = image_file_to_string('fonts_test.png', graceful_errors=True)
    print (text)

【问题讨论】:

  • @user10089632 和第二个?。我按照教程中的说明做了

标签: python python-3.x ocr tesseract


【解决方案1】:

而不是使用:

except errors.Tesser_General_Exception, value:

替换为

except errors.Tesser_General_Exception as value:

它对我有用。这就是从Python2 升级到Python3。

【讨论】:

  • 兄弟这个问题是我在 python 3 中使用了 python2 库,不管有什么代码会出现问题,我实际上不得不调试整个库。即使在那之后,一些问题仍然存在所以我决定使用python2.7,一切都落到了它的位置
猜你喜欢
  • 2015-12-21
  • 1970-01-01
  • 1970-01-01
  • 2013-05-11
  • 1970-01-01
  • 2020-01-25
  • 1970-01-01
  • 1970-01-01
  • 2023-02-09
相关资源
最近更新 更多