【问题标题】:Check response using urllib2使用 urllib2 检查响应
【发布时间】:2023-04-07 19:46:01
【问题描述】:

我正在尝试通过使用 opencorporates api 增加页面计数器来访问页面。但问题是有时会有无用的数据。例如,在下面的权限代码 = ae_az 的网址中,我得到的网页只显示了这个:

{"api_version":"0.2","results":{"companies":[],"page":1,"per_page":26,"total_pages":0,"total_count":0}}

这在技术上是空的。如何检查此类数据并跳过此以进入下一个司法管辖区?

这是我的代码

import urllib2
import json,os

f = open('codes','r')
for line in f.readlines():
   id = line.strip('\n')
   url = 'http://api.opencorporates.com/v0.2/companies/search?q=&jurisdiction_code={0}&per_page=26&current_status=Active&page={1}?api_token=ab123cd45' 
   i = 0
   directory = id
   os.makedirs(directory)
   while True:
      i += 1
      req = urllib2.Request(url.format(id, i))
      print url.format(id,i)
      try:
        response = urllib2.urlopen(url.format(id, i))
      except urllib2.HTTPError, e:
        break
      content = response.read()
      fo = str(i) + '.json'    
      OUTFILE = os.path.join(directory, fo)
      with open(OUTFILE, 'w') as f:
        f.write(content)

【问题讨论】:

    标签: python json api opencore


    【解决方案1】:

    解释您返回的响应(您已经知道它是 json)并检查您想要的数据是否存在。

    ...
    content = response.read()
    data = json.loads(content)
    if not data.get('results', {}).get('companies'):
        break
    ...
    

    这是您使用 Requests 编写的代码并在此处使用答案。它远没有应有的强大或干净,但展示了您可能想要采取的路径。速率限制是一个猜测,似乎不起作用。请记住输入您的实际 API 密钥。

    import json
    import os
    from time import sleep
    import requests
    
    url = 'http://api.opencorporates.com/v0.2/companies/search'
    token = 'ab123cd45'
    rate = 20  # seconds to wait after rate limited
    
    with open('codes') as f:
        codes = [l.strip('\n') for l in f]
    
    
    def get_page(code, page, **kwargs):
        params = {
            # 'api_token': token,
            'jurisdiction_code': code,
            'page': page,
        }
        params.update(kwargs)
    
        while True:
            r = requests.get(url, params=params)
    
            try:
                data = r.json()
            except ValueError:
                return None
    
            if 'error' in data:
                print data['error']['message']
                sleep(rate)
                continue
    
            return data['results']
    
    
    def dump_page(code, page, data):
        with open(os.path.join(code, str(page) + '.json'), 'w') as f:
            json.dump(data, f)
    
    
    for code in codes:
        try:
            os.makedirs(code)
        except os.error:
            pass
    
        data = get_page(code, 1)
        if data is None:
            continue
    
        dump_page(code, 1, data['companies'])
    
        for page in xrange(1, int(data.get('total_pages', 1))):
            data = get_page(code, page)
            if data is None:
                break
    
            dump_page(code, page, data['companies'])
    

    【讨论】:

      【解决方案2】:

      我认为实际上这个例子不是“技术上是空的”。它包含数据,因此在技术上不是空的。数据只是不包括对您有用的任何字段。 :-)

      如果您希望您的代码跳过包含无趣数据的响应,那么只需在写入任何数据之前检查 JSON 是否具有必要的字段:

      content = response.read()
      try:
          json_content = json.loads(content)
          if json_content['results']['total_count'] > 0:
              fo = str(i) + '.json'    
              OUTFILE = os.path.join(directory, fo)
              with open(OUTFILE, 'w') as f:
                  f.write(content)
      except KeyError:
          break
      except ValueError:
          break
      

      等等。您可能想要报告 ValueError 或 KeyError,但这取决于您。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2012-08-18
        • 1970-01-01
        • 1970-01-01
        • 2012-10-24
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2014-06-18
        相关资源
        最近更新 更多