【发布时间】:2019-05-11 13:43:56
【问题描述】:
我试图从这个网站获取世界人口:https://www.worldometers.info/world-population/ 但我只能得到html代码,而不是实际数字的数据。
我已经尝试找到我尝试从中获取数据的对象的子对象。我也尝试列出整个对象,但似乎没有任何效果。
'''只是导入东西'''
import urllib.request
import requests
from bs4 import BeautifulSoup
'''从网站获取html到文本'''
r = requests.get('https://www.worldometers.info/world-population/')
soup = BeautifulSoup(r.text,'html.parser')
'''这里它只找到下面列出的一个对象'''
current_population = soup.find('div',{'class':'maincounter-number'}).find_all('span', recursive=False)
print(current_population)
这是存储信息的对象:
(span class="rts-counter" rel="current_population">retrieving data... </span>
在“检查模式”中你可以看到:
(span class="rts-counter" rel="current_population">(span class="rts-nr-sign"></span>(span class="rts-nr-int rts-nr-10e9">7</span>(span class="rts-nr-thsep">,</span>(span class="rts-nr-int rts-nr-10e6">703</span>(span class="rts-nr-thsep">,</span>(span class="rts-nr-int rts-nr-10e3">227</span><span class="rts-nr-thsep">,</span>(span class="rts-nr-int rts-nr-10e0">630</span></span>
我总是只得到第一个,但想从'inspect-mode'中得到第二个。
Here 是检查模式的图片。
【问题讨论】:
-
本网站所依赖的 API 已获得许可。一定会有公共 api 有这些数据。
标签: python html web-scraping