【发布时间】:2018-04-17 22:18:06
【问题描述】:
我正在使用 BeautifulSoup 尝试从该 URL 获取所有 2000 家公司的整个表格:
https://www.forbes.com/global2000/list/#tab:overall.
这是我写的代码:
from bs4 import BeautifulSoup
import urllib.request
html_content = urllib.request.urlopen('https://www.forbes.com/global2000/list/#header:position')
soup = BeautifulSoup(html_content, 'lxml')
table = soup.find_all('table')[0]
new_table = pd.DataFrame(columns=range(0,7), index = [0])
row_marker = 0
for row in table.find_all('tr'):
column_marker = 0
columns = row.find_all('td')
for column in columns:
new_table.iat[row_marker,column_marker] = column.get_text()
column_marker += 1
new_table
在结果中,我只得到列的名称,而不是表本身。
我怎样才能得到整张桌子。
【问题讨论】:
-
如果页面使用 javascript 填充表格,BeautifulSoup 不会运行它。也许看看Selenium。
标签: python parsing web-scraping beautifulsoup