免责声明:我不是 HTTP 或 Requests 库方面的专家,也没有使用过 neuromorpho.org,所以请把这个和一粒盐一起吃。
您可以查询第一个请求的页面数,然后循环浏览各个页面。在循环中,您必须将请求的页面作为参数包含在 HTTP GET 方法中,例如?page=42&...,像这样:
url = 'http://neuromorpho.org/api/neuron/select'
params = {
'page' : 0,
'q' : 'species:rat',
'fq' : [
'brain_region:hippocampus,CA1',
'experiment_condition:Control',
'cell_type:Pyramidal,principal cell' ] }
totalPages = requests.get(url, params).json()['page']['totalPages']
df_dict = {
'NeuronID' : list(),
'Archive' : list(),
'Strain' : list(),
'Cell' : list(),
'Region' : list() }
for pageNum in range(totalPages):
params['page'] = pageNum
response = requests.get(url, params)
print('Querying page {} -> status code: {}'.format(
pageNum, response.status_code))
if (response.status_code == 200): #only parse successful requests
data = response.json()
for row in data['_embedded']['neuronResources']:
df_dict['NeuronID'].append(str(row['neuron_id']))
df_dict['Archive'].append(str(row['archive']))
df_dict['Strain'].append(str(row['strain']))
df_dict['Cell'].append(str(row['cell_type']))
df_dict['Region'].append(str(row['brain_region']))
rat_df = pd.DataFrame(df_dict)
print(rat_df)
您可以在控制台输出中看到生成的DataFrame 以及请求的页码如何变化:
Querying page 0 -> status code: 200
Querying page 1 -> status code: 200
Querying page 2 -> status code: 200
Querying page 3 -> status code: 200
Querying page 4 -> status code: 200
Querying page 5 -> status code: 200
Querying page 6 -> status code: 200
Querying page 7 -> status code: 200
Querying page 8 -> status code: 200
Querying page 9 -> status code: 200
Querying page 10 -> status code: 200
Querying page 11 -> status code: 200
Querying page 12 -> status code: 200
Querying page 13 -> status code: 200
Querying page 14 -> status code: 200
Querying page 15 -> status code: 200
Querying page 16 -> status code: 200
Querying page 17 -> status code: 200
Querying page 18 -> status code: 200
Querying page 19 -> status code: 200
Querying page 20 -> status code: 200
Querying page 21 -> status code: 200
Querying page 22 -> status code: 200
NeuronID Archive Strain Cell Region
0 100 Turner Fischer 344 ['pyramidal', 'principal cell'] ['hippocampus', 'CA1']
1 101 Turner Fischer 344 ['pyramidal', 'principal cell'] ['hippocampus', 'CA1']
2 1016 Ascoli Sprague-Dawley ['pyramidal', 'principal cell'] ['hippocampus']
3 1019 Ascoli Sprague-Dawley ['pyramidal', 'principal cell'] ['hippocampus']
4 102 Turner Fischer 344 ['pyramidal', 'principal cell'] ['hippocampus', 'CA1']
... ... ... ... ... ...
1110 99614 Guizzetti Sprague-Dawley ['principal cell', 'pyramidal'] ['hippocampus', 'CA1', 'left']
1111 99615 Guizzetti Sprague-Dawley ['principal cell', 'pyramidal'] ['hippocampus', 'CA1', 'left']
1112 99616 Guizzetti Sprague-Dawley ['principal cell', 'pyramidal'] ['hippocampus', 'CA1', 'left']
1113 99617 Guizzetti Sprague-Dawley ['principal cell', 'pyramidal'] ['hippocampus', 'CA1', 'left']
1114 99618 Guizzetti Sprague-Dawley ['principal cell', 'pyramidal'] ['hippocampus', 'CA1', 'left']
[1115 rows x 5 columns]
更新 #1:
我通过添加代码的修改版本来更改我发布的代码以解析循环中的响应。我认为 neuromorpho.org API 中存在一个小错误,因为它在最后一页(第 22 号)以 size: 50 响应,而它仅包含 15 个(索引 0-14)对象JSON 响应。您可以通过遍历 JSON 对象并忽略报告的大小来规避该问题。
更新 #2:
意识到 GET 参数不必在 URL 中进行编码,但 Requests 在将它们作为 dict 传递时为我们执行此操作(更新了代码)。
我希望这会有所帮助!