【问题标题】:Receiving HTTP Error 404: Not Found, during image to text conversion在图像到文本的转换过程中接收 HTTP 错误 404:未找到
【发布时间】:2020-12-11 14:50:56
【问题描述】:

我收到 HTTPError: HTTP Error 404: Not Found error,同时对来自 excel 电子表格的图像 URL 实施 easyOCR。这个想法是遍历 excel 文件中的每个 URL,并创建一个带有提取文本的输出 excel 文件。

输入文件是这样的,input data

除了 iterrows(),我还使用了 range(len()) 并遇到了同样的问题。请就解决此问题提出建议。提前致谢!

def modelling(url):
  reader = easyocr.Reader(['en'])
  bounds = reader.readtext(url, detail = 0)
  return bounds

new_col = []

for index, row in df.iterrows():
  img = row['url']
  col = modelling(img)
  new_col.append(col)

df['new_col'] = new_col
filename = "Output_Data.xlsx"
df.to_excel(filename)
print("Data frame is written to Excel file successfully.")

错误日志如下所示。

<ipython-input-10-013fb5c72447> in modelling(url)
      1 def modelling(url):
      2   reader = easyocr.Reader(['en'])
----> 3   bounds = reader.readtext(url, detail = 0)
      4   return bounds

/usr/local/lib/python3.6/dist-packages/easyocr/easyocr.py in readtext(self, image, decoder, beamWidth, batch_size, workers, allowlist, blocklist, detail, paragraph, min_size, contrast_ths, adjust_contrast, filter_ths, text_threshold, low_text, link_threshold, canvas_size, mag_ratio, slope_ths, ycenter_ths, height_ths, width_ths, add_margin)
    345         image: file path or numpy-array or a byte stream object
    346         '''
--> 347         img, img_cv_grey = reformat_input(image)
    348 
    349         horizontal_list, free_list = self.detect(img, min_size, text_threshold,\

/usr/local/lib/python3.6/dist-packages/easyocr/utils.py in reformat_input(image)
    645     if type(image) == str:
    646         if image.startswith('http://') or image.startswith('https://'):
--> 647             tmp, _ = urlretrieve(image , reporthook=printProgressBar(prefix = 'Progress:', suffix = 'Complete', length = 50))
    648             img_cv_grey = cv2.imread(tmp, cv2.IMREAD_GRAYSCALE)
    649             os.remove(tmp)

/usr/lib/python3.6/urllib/request.py in urlretrieve(url, filename, reporthook, data)
    246     url_type, path = splittype(url)
    247 
--> 248     with contextlib.closing(urlopen(url, data)) as fp:
    249         headers = fp.info()
    250 

/usr/lib/python3.6/urllib/request.py in urlopen(url, data, timeout, cafile, capath, cadefault, context)
    221     else:
    222         opener = _opener
--> 223     return opener.open(url, data, timeout)
    224 
    225 def install_opener(opener):

/usr/lib/python3.6/urllib/request.py in open(self, fullurl, data, timeout)
    530         for processor in self.process_response.get(protocol, []):
    531             meth = getattr(processor, meth_name)
--> 532             response = meth(req, response)
    533 
    534         return response

/usr/lib/python3.6/urllib/request.py in http_response(self, request, response)
    640         if not (200 <= code < 300):
    641             response = self.parent.error(
--> 642                 'http', request, response, code, msg, hdrs)
    643 
    644         return response

/usr/lib/python3.6/urllib/request.py in error(self, proto, *args)
    568         if http_err:
    569             args = (dict, 'default', 'http_error_default') + orig_args
--> 570             return self._call_chain(*args)
    571 
    572 # XXX probably also want an abstract factory that knows when it makes

/usr/lib/python3.6/urllib/request.py in _call_chain(self, chain, kind, meth_name, *args)
    502         for handler in handlers:
    503             func = getattr(handler, meth_name)
--> 504             result = func(*args)
    505             if result is not None:
    506                 return result

/usr/lib/python3.6/urllib/request.py in http_error_default(self, req, fp, code, msg, hdrs)
    648 class HTTPDefaultErrorHandler(BaseHandler):
    649     def http_error_default(self, req, fp, code, msg, hdrs):
--> 650         raise HTTPError(req.full_url, code, msg, hdrs, fp)
    651 
    652 class HTTPRedirectHandler(BaseHandler):

HTTPError: HTTP Error 404: Not Found

【问题讨论】:

    标签: python excel machine-learning ocr http-error


    【解决方案1】:

    使用tryexcept 处理错误并防止它们破坏您的程序。做这样的事情

    try:
        reader = easyocr.Reader(['en'])
        bounds = reader.readtext(url, detail = 0)
        return bounds
    except:
        #error raised
        #do something else
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2022-06-10
      • 2021-02-22
      • 1970-01-01
      • 2018-08-08
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多