【问题标题】:Perl Image::OCR::Tesseract module on WindowsWindows 上的 Perl Image::OCR::Tesseract 模块
【发布时间】:2020-12-16 16:04:14
【问题描述】:

有人知道在 Windows 上安装“Image::OCR::Tesseract”模块的优雅方法吗?由于名为“LEOCHARRE::CLI”的仅限 *NIX 的模块依赖项,该模块无法通过 CPAN 在 Windows 上安装。这个模块似乎不需要运行“Image::OCR::Tesseract”本身。

我已经设法使模块工作,首先手动安装 makefile.pl 中列出的依赖模块(“LEOCHARRE::CLI”除外),然后将模块文件移动到“C”下的正确目录结构:\Perl\site\lib\Image\OCR"。让它工作的最后一部分是更改从命令行调用 ImageMagick 和 Tesseract 可执行文件的代码部分,以便在模块调用可执行文件时在程序名称周围加上引号。

这行得通,但我真的觉得通过在 Windows 上运行的存储库在生产系统上安装 PPM 或 CPAN 会感觉更好。

【问题讨论】:

    标签: windows perl tesseract


    【解决方案1】:

    没关系,我知道了,但我无法确定更好的解决方案。

    要让安装程序通过传统的“perl makefile.pl、make、make test、make install”例程在 Windows 上运行,需要编辑 Makefile.pl 脚本,包括缺少的 Windows 安装模块(Devel::AssertOS ::MSWin32),并修补 AssertEXE.pm 以使用“File::Which”而不是 Windows 缺少的内置 shell“which”命令。所有这一切仍然需要修补“Image::OCR::Tesseract”以在从命令行执行“convert”和“tesseract”时在程序名称周围加上引号。

    考虑到使安装程序在 Windows 上运行所涉及的步骤数量,以及该模块不会为模块创建要链接到的二进制组件这一事实,我会说安装和使 Tesseract 模块正常工作的最佳选择在 windows 上将首先安装以下二进制包:

    ImageMagick Link

    正方体 http://code.google.com/p/tesseract-ocr/downloads/list

    接下来,找到您的 Perl 模块目录 - 在我的系统上它是“C:\Perl\site\lib”。创建一个文件夹“图像”,如果你没有。接下来,打开 Image 文件夹并创建一个名为“OCR”的文件夹。打开 OCR 文件夹。此时,您的路径应该类似于“C:\Perl\site\lib\Image\OCR”。创建一个名为“Tesseract.pm”的新文本文件,并复制以下内容...

    package Image::OCR::Tesseract;
    use strict;
    use Carp;
    use Cwd;
    use String::ShellQuote 'shell_quote';
    use Exporter;
    use vars qw(@EXPORT_OK @ISA $VERSION $DEBUG $WHICH_TESSERACT $WHICH_CONVERT %EXPORT_TAGS @TRASH);
    @ISA = qw(Exporter);
    @EXPORT_OK = qw(get_ocr get_hocr _tesseract convert_8bpp_tif tesseract);
    $VERSION = sprintf "%d.%02d", q$Revision: 1.24 $ =~ /(\d+)/g;
    %EXPORT_TAGS = ( all => \@EXPORT_OK );
    
    
    BEGIN {
       use File::Which 'which';
       $WHICH_TESSERACT = which('tesseract');
       $WHICH_CONVERT   = which('convert');
       
       if($^O=~m/MSWin/) {
          $WHICH_TESSERACT='"'.$WHICH_TESSERACT.'"';
          $WHICH_CONVERT='"'.$WHICH_CONVERT.'"';
       }
       $WHICH_TESSERACT or die("Is tesseract installed? Cannot find bin path to tesseract.");
       $WHICH_CONVERT or die("Is convert installed? Cannot find bin path to convert.");
    }
    
    END {
       scalar @TRASH or return;
       if ( $DEBUG ){
          print STDERR "Debug on, these are trash files:\n".join("\n",@TRASH) ;
       }
       else {
          unlink @TRASH;
       }
    }
    
    sub DEBUG { Carp::cluck("Image::OCR::Tesseract::DEBUG() deprecated") }
    
    sub get_hocr {
       my ($abs_image,$abs_tmp_dir,$lang)= @_;
       -f $abs_image or croak("$abs_image is not a file on disk");
       my $hocr="hocr";
       if(defined $abs_tmp_dir){
    
          -d $abs_tmp_dir or die("tmp dir arg $abs_tmp_dir not a dir on disk.");
    
          $abs_image=~/([^\/]+)$/ or die("cant match filename in path arg '$abs_image'");
          my $abs_copy = "$abs_tmp_dir/$1";
    
          # TODO, what if source and dest are same, i want it to die
          require File::Copy;
          File::Copy::copy($abs_image, $abs_copy) 
             or die("cant make copy of $abs_image to $abs_copy, $!");
    
          # change the image to get ocr from to be the copy
          $abs_image = $abs_copy;
          # since it's a copy. erase that on exit
          push @TRASH, $abs_image;      
       }
    
       my $tmp_tif = convert_8bpp_tif($abs_image);
       
       push @TRASH, $tmp_tif; # for later delete
    
       _tesseract($tmp_tif,$lang,$hocr) || '';
    }
    
    sub get_ocr {
       my ($abs_image,$abs_tmp_dir,$lang)= @_;
       -f $abs_image or croak("$abs_image is not a file on disk");
       if(defined $abs_tmp_dir){
    
          -d $abs_tmp_dir or die("tmp dir arg $abs_tmp_dir not a dir on disk.");
    
          $abs_image=~/([^\/]+)$/ or die("cant match filename in path arg '$abs_image'");
          my $abs_copy = "$abs_tmp_dir/$1";
    
          # TODO, what if source and dest are same, i want it to die
          require File::Copy;
          File::Copy::copy($abs_image, $abs_copy) 
             or die("cant make copy of $abs_image to $abs_copy, $!");
    
          # change the image to get ocr from to be the copy
          $abs_image = $abs_copy;
          # since it's a copy. erase that on exit
          push @TRASH, $abs_image;      
       }
    
       my $tmp_tif = convert_8bpp_tif($abs_image);
       
       push @TRASH, $tmp_tif; # for later delete
    
       _tesseract($tmp_tif,$lang) || '';
    }
    
    sub convert_8bpp_tif {
       my ($abs_img,$abs_out) = (shift,shift);
       defined $abs_img or die('missing image arg');
    
       $abs_out ||= $abs_img.'.tmp.'.time().(int rand(9000)).'.tif';
       
       my @arg = ( $WHICH_CONVERT, $abs_img, '-compress','none','+matte', $abs_out );
       
       #die (join(" ", @arg));
       
       system(@arg) == 0 or die("convert $abs_img error.. $?");
    
       $DEBUG and warn("made $abs_out 8bpp tiff.");
       $abs_out;
    }
    
    
    
    # people expect tesseract to automatically convert
    
    *tesseract = \&_tesseract;
    sub _tesseract {
        my ($abs_image,$lang,$hocr) = @_;
       defined $abs_image or croak('missing image path arg');
       
       $abs_image=~/\.tif+$/i or warn("Are you sure '$abs_image' is a tif image? This operation may fail.");
       
       #my @arg = (
       #   $WHICH_TESSERACT, shell_quote($abs_image), shell_quote($abs_image), 
       #   (defined $lang and ('-l', $lang) ), '2>/dev/null'
       #); 
    
       my $cmd = 
          ( sprintf '%s %s %s', 
             $WHICH_TESSERACT, 
             shell_quote($abs_image), 
             shell_quote($abs_image) 
          ) .
          ( defined $lang ? " -l $lang" : '' ) .
          ( defined $hocr ? " hocr" : '' ) .
          "  2>/dev/null";
       $DEBUG and warn "command: $cmd";
    
        system($cmd); # hard to check ==0 
    
        my $txt = $abs_image.($hocr?".html":".txt");
       unless( -f $txt ){      
            Carp::cluck("no text output for image '$abs_image'. (No text file '$txt' found on disk)");
          return;
       }
    
        $DEBUG and warn "Found text file '$txt'";
       
       my $content = (_slurp($txt) || '');   
       $DEBUG and warn("content length of text in '$txt' from image '$abs_image' is ". length $content );
       push @TRASH, $txt;
    
       $content;
    }
    
    sub _slurp {
       my $abs = shift;
       open(FILE,'<', $abs) or die("can't open file for reading '$abs', $!");
       local $/;
       my $txt = <FILE>;
       close FILE;
       $txt;
    }  
    
    1;
    
    
    __END__
    
    #sub _force_imgtype {
    #   my $img = shift;
    #   my $type = shift;
    #   my $delete_original = shift;
    #   $delete_original ||=0;
    #   
    #
    #   if($img=~/\.$type$/i){
    #      return $img;
    #   }
    #
    #   my $img_out= $img;
    #   $img_out=~s/\.\w{1,5}$/\.$type/ or die("cant get file ext for $img");
    #
    #
    #
    #}
    

    保存并关闭。如果您在安装 ImageMagick 和 Tesseract 二进制文件之前打开了一个命令行会话,请关闭命令行会话并打开一个新会话。使用以下脚本测试模块:

    use Image::OCR::Tesseract;
    my $image = 'SomeImageFileThatContainsText.jpg';
    
    my $text = Image::OCR::Tesseract::get_ocr($image);
    
    print "Text...\n";
    print $text."\n";
    
    print "Normal Exit\n";
    
    exit;
    

    就是这样。混乱,我知道,但是模块安装程序确实需要更新以支持 Windows(和其他)系统这一事实没有好的方法,即使实际的模块代码几乎无需修改即可运行。确实,如果将 Tesseract 和 ImageMagick 安装到没有空格的路径,则“Image::OCR::Tesseract”模块代码不需要任何更改,但是这个小调整可以让支持的可执行文件安装在任何地方,包括默认位置。

    【讨论】:

    • 或许您应该在rt.cpan.org/Dist/Display.html?Name=Image-OCR-Tesseract 上分享您的发现。
    • 我会在接下来的几天内尝试直接联系作者,一旦我解决了一些小问题。他对 CPAN 的反应似乎不是很好,但我今天早些时候发现了他的主页。
    • 进一步更新:我在发布此问题后不久尝试联系模块作者,但没有任何回复。可悲的是,我认为我们必须将此 Perl 模块视为已被废弃的类别。
    • Rick:你能推荐一个好的和活跃的 perl 模块吗?我正在开始一个项目,我必须从图像中读取文本。
    • 即使在当前的“废弃”状态 Image::OCR::Tesseract 似乎仍然是最好的选择(对于 Perl)。实际上,这种类型的模块并不难编写,因为命令行 Tesseract 可执行文件正在为您完成大部分工作。一组新代码需要做的就是将图像转换为 TIFF,将正确的参数传递给 Tesseract,然后收集并返回输出。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2015-04-21
    • 2013-07-17
    • 2015-10-30
    • 2015-09-10
    • 2011-11-29
    • 2018-11-28
    • 1970-01-01
    相关资源
    最近更新 更多