【问题标题】:What is the pythonic way to catch errors and keep going in this loop?捕获错误并继续循环的pythonic方法是什么?
【发布时间】:2012-11-13 07:53:24
【问题描述】:

我有两个可以正常工作的函数,但是当我将它们嵌套在一起运行时似乎崩溃了。

def scrape_all_pages(alphabet):
    pages = get_all_urls(alphabet)
    for page in pages:
        scrape_table(page)

我正在尝试系统地抓取一些搜索结果。所以get_all_pages() 为字母表中的每个字母创建一个 URL 列表。有时有数千页,但效果很好。然后,对于每一页,scrape_table 只抓取我感兴趣的表格。这也很好。我可以运行整个程序并且它运行良好,但我在 Scraperwiki 工作,如果我将它设置为运行并离开它总是会给我一个“列表索引超出范围”错误。这绝对是 scraperwiki 中的一个问题,但我想通过添加一些 try/except 子句并在遇到错误时记录错误来找到解决问题的方法。比如:

def scrape_all_pages(alphabet):
    try:
        pages = get_all_urls(alphabet)
    except:
        ## LOG THE ERROR IF THAT FAILS.
    try:
        for page in pages:
            scrape_table(page)
    except:
        ## LOG THE ERROR IF THAT FAILS

不过,我无法弄清楚如何一般地记录错误。此外,上面的代码看起来很笨拙,根据我的经验,当某些东西看起来很笨拙时,Python 有更好的方法。有没有更好的办法?

【问题讨论】:

    标签: python error-handling scraperwiki


    【解决方案1】:

    最好这样写:

        try:
            pages = get_all_urls(alphabet)
        except IndexError:
            ## LOG THE ERROR IF THAT FAILS.
        for page in pages:
            try:
                scrape_table(page)
            except IndexError:
                continue ## this will bring you to the next item in for
            ## LOG THE ERROR IF THAT FAILS
    

    【讨论】:

      【解决方案2】:

      将日志信息包装在上下文管理器周围,尽管您可以轻松更改详细信息以满足您的要求:

      import traceback
      
      # This is a context manager
      class LogError(object):
          def __init__(self, logfile, message):
              self.logfile = logfile
              self.message = message
          def __enter__(self):
              return self
          def __exit__(self, type, value, tb):
              if type is None or not issubclass(type, Exception):
                  # Allow KeyboardInterrupt and other non-standard exception to pass through
                  return
      
              self.logfile.write("%s: %r\n" % (self.message, value))
              traceback.print_exception(type, value, tb, file=self.logfile)
              return True # "swallow" the traceback
      
      # This is a helper class to maintain an open file object and
      # a way to provide extra information to the context manager.
      class ExceptionLogger(object):
          def __init__(self, filename):
              self.logfile = open(filename, "wa")
          def __call__(self, message):
              # override function() call so that I can specify a message
              return LogError(self.logfile, message)
      

      关键部分是__exit__可以返回'True',在这种情况下异常被忽略,程序继续进行。代码也需要小心一点,因为可能会引发 KeyboardInterrupt (control-C)、SystemExit 或其他非标准异常,并且您确实希望程序在哪里停止。

      您可以像这样在代码中使用上述内容:

      elog = ExceptionLogger("/dev/tty")
      
      with elog("Can I divide by 0?"):
          1/0
      
      for i in range(-4, 4):
          with elog("Divisor is %d" % (i,)):
              print "5/%d = %d" % (i, 5/i)
      

      那个 sn-p 给了我输出:

      Can I divide by 0?: ZeroDivisionError('integer division or modulo by zero',)
      Traceback (most recent call last):
        File "exception_logger.py", line 24, in <module>
          1/0
      ZeroDivisionError: integer division or modulo by zero
      5/-4 = -2
      5/-3 = -2
      5/-2 = -3
      5/-1 = -5
      Divisor is 0: ZeroDivisionError('integer division or modulo by zero',)
      Traceback (most recent call last):
        File "exception_logger.py", line 28, in <module>
          print "5/%d = %d" % (i, 5/i)
      ZeroDivisionError: integer division or modulo by zero
      5/1 = 5
      5/2 = 2
      5/3 = 1
      

      我认为也很容易看出如何修改代码以处理仅记录 IndexError 异常,甚至传递基本异常类型以进行捕获。

      【讨论】:

        【解决方案3】:

        这是一个好方法,但是。你不应该只使用except 子句,你必须指定你试图捕获的异常的类型。您也可以捕获错误并继续循环。

        def scrape_all_pages(alphabet):
            try:
                pages = get_all_urls(alphabet)
            except IndexError: #IndexError is an example
                ## LOG THE ERROR IF THAT FAILS.
        
            for page in pages:
                try:
                    scrape_table(page)
                except IndexError: # IndexError is an example
                    ## LOG THE ERROR IF THAT FAILS and continue this loop
        

        【讨论】:

        • 我实际上得到了类似except Exception,e: print str(e) 的东西——仍然不是那么优雅,但它帮助我将错误归零。
        【解决方案4】:

        也许记录每次迭代的错误,以便一次迭代中的错误不会破坏您的循环:

        for page in pages:
            try:
                scrape_table(page)
            except:
                #open error log file for append:
                f=open("errors.txt","a")
                #write error to file:
                f.write("Error occured\n") # some message specific to this iteration (page) should be added here...
                #close error log file:
                f.close()
        

        【讨论】:

          【解决方案5】:

          您可以指定要捕获的某种类型的异常和保存异常实例的变量:

          def scrape_all_pages(alphabet):
              try:
                  pages = get_all_urls(alphabet)
                  for page in pages:
                      scrape_table(page)
              except OutOfRangeError as error:
                  # Will only catch OutOfRangeError
                  print error
              except Exception as error:
                  # Will only catch any other exception
                  print error
          

          捕获 Exception 类型将捕获所有错误,因为它们应该都是从 Exception 继承的。

          这是我所知道的捕捉错误的唯一方法。

          【讨论】:

          • 您还可以通过上下文管理器捕获错误。我在这里的其他地方发布了一个示例。
          猜你喜欢
          • 2011-10-20
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2020-07-01
          • 1970-01-01
          • 2019-07-19
          • 1970-01-01
          • 2011-02-07
          相关资源
          最近更新 更多