【问题标题】:How To Scrape Table Into Dataframe From The Webpage如何将表格从网页中抓取到数据框中
【发布时间】:2021-01-13 01:06:08
【问题描述】:

我正在尝试用一页将表格刮入数据框。

import pandas as pd
import requests
from bs4 import BeautifulSoup

res = requests.get("https://www.viewbase.com/funding")
soup = BeautifulSoup(res.content,'lxml')

table1 = soup.find_all('tr')

【问题讨论】:

  • df = pd.read_html(str(soup.find('table'))) 试试这样的。
  • 谢谢。已尝试但只能转义表格标题而不是数据。
  • 似乎 bs4 无法获取使用 selenium 的值。

标签: python dataframe web-scraping


【解决方案1】:

该表是通过 JS 脚本填充的,因此 BS4 不会看到它。但是,您可以在headless 模式下使用selenium 并获取您需要的内容。

这是如何做到这一点的:

import time

import pandas as pd

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)

driver.get("https://www.viewbase.com/funding")
time.sleep(5)
headers = driver.find_elements_by_xpath('//*[@class="tablesorter-headerRow"][2]/th/div')
table = driver.find_element_by_xpath('//*[@id="inverse_swap"]')

columns = [i.text for i in headers]
data = [r.split() for r in table.text.split('\n')]

df = pd.DataFrame(data, columns=columns)
df.to_csv("data.csv", index=False)

输出:

【讨论】:

  • 感谢您的回答。但是在运行代码时会出现一些错误。我正在努力:)
  • 您需要安装selenium 和pip install -U selenium 并从这里获取webdriver - chromedriver.chromium.org/downloads
  • 对堆栈溢出的人才如此之多感到惊讶。我赞成您的回答,但我的声誉太低而无法显示。非常感谢。
【解决方案2】:

你是否也想要标题

driver.get("https://www.viewbase.com/funding")
tbl = driver.find_element_by_tag_name('table')
headers = tbl.find_elements_by_xpath('//thead/tr[2]/th')
headers = [item.text.strip() for item in headers]
trs = tbl.find_elements_by_xpath('tbody/tr')
lst=[]

for tr in trs:
    tds = tr.find_elements_by_tag_name('td')
    tds = [item.text.strip() for item in tds]
    lst.append([item for item in tds if item])

#print(lst)
print(headers)
df = pd.DataFrame(lst)
print(df)

目前的输出

['', 'Binance', 'FTX', 'Okex', 'Bybit', 'Binance', 'Huobi', 'Okex', 'Bitmex', 'Bybit', '']
        0         1         2         3  ...         6         7        8        9
0     BTC   0.0100%   0.0020%  -0.0062%  ...   0.0034%  -0.0047%  0.0100%  0.0100%
1     ETH   0.0100%   0.0010%  -0.0034%  ...   0.0100%   0.0111%  0.0193%  0.0100%
2     XRP   0.0100%   0.0013%   0.0003%  ...  -0.0096%  -0.0089%  0.0100%  0.0100%
3     EOS   0.0100%   0.0017%  -0.0098%  ...   0.0100%   0.0186%        -  0.0100%
4     BCH   0.0100%   0.0006%   0.0227%  ...   0.0100%   0.0135%  0.0110%        -
5     LTC   0.0100%  -0.0012%   0.0180%  ...   0.0100%   0.0006%  0.0100%        -
6    LINK   0.0031%  -0.0004%  -0.0272%  ...  -0.0321%  -0.0221%        -        -
7     BSV         -  -0.0020%  -0.0280%  ...   0.0105%  -0.0468%        -        -
8     BNB  -0.1475%  -0.0046%         -  ...         -         -        -        -

[59 行 x 10 列]

导入

from selenium import webdriver
import pandas as pd

【讨论】:

  • 谢谢。我想要标题。但是当我运行代码时,我遇到了 "NameError: name 'driver' is not defined" 的错误。即使我运行“driver = webdriver.Chrome(options=options)”,它也不起作用。
  • 你必须 pip install selenium 并获得一个 chromedriver
  • 您可以使用 driver = webdriver.Chrome(ChromeDriverManager().install(),options=options),从 webdriver_manager.chrome 导入 ChromeDriverManager,然后 pip install webdriver-manager。
  • 而不是获取 chromedriver。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2023-01-15
  • 1970-01-01
  • 1970-01-01
  • 2014-05-21
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多