【问题标题】:Can't get a table from a web page无法从网页获取表格
【发布时间】:2018-04-17 22:18:06
【问题描述】:

我正在使用 BeautifulSoup 尝试从该 URL 获取所有 2000 家公司的整个表格:

https://www.forbes.com/global2000/list/#tab:overall.

这是我写的代码:

from bs4 import BeautifulSoup
import urllib.request

html_content = urllib.request.urlopen('https://www.forbes.com/global2000/list/#header:position')

soup = BeautifulSoup(html_content, 'lxml')
table = soup.find_all('table')[0]
new_table = pd.DataFrame(columns=range(0,7), index = [0])

row_marker = 0
for row in table.find_all('tr'):
   column_marker = 0
   columns = row.find_all('td')

   for column in columns:
      new_table.iat[row_marker,column_marker] = column.get_text()
      column_marker += 1
new_table

在结果中,我只得到列的名称,而不是表本身。

我怎样才能得到整张桌子。

【问题讨论】:

  • 如果页面使用 javascript 填充表格,BeautifulSoup 不会运行它。也许看看Selenium。

标签: python parsing web-scraping beautifulsoup


【解决方案1】:

内容是通过javascript生成的,所以你必须selenium来模仿浏览器和滚动动作,然后用beautiful soup解析页面源,或者在某些情况下,像这样,你可以通过查询它们的 ajax API 来访问这些值:

import requests
import json

headers = {'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:50.0) Gecko/20100101 Firefox/50.0'}

target = 'https://www.forbes.com/ajax/list/data?year=2017&uri=global2000&type=organization'

with requests.Session() as s:
    s.headers = headers
    data = json.loads(s.get(target).text)

print([x['name'] for x in data[:5]])

输出(前 5 项):

['3M', '3i Group', '77 Bank', 'AAC Technologies Holdings', 'ABB']

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2021-12-05
    • 2021-08-29
    • 1970-01-01
    • 2015-06-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多