【发布时间】:2021-07-08 13:28:16
【问题描述】:
我会从 Transfermarkt 玩家资料页面中的两个 html 表中抓取数据。 这是一个页面示例:https://www.transfermarkt.com/cristiano-ronaldo/profil/spieler/8198
第一个是“事实和数据”表,第二个是“统计”表。我想从搜索页面开始抓取并获取网址。 一旦我从搜索页面的每一页获取了 url,就开始为每个播放器链接抓取统计信息。
如何从该链接中抓取 html 表格的数据?
这是我的完整代码
import requests
from bs4 import BeautifulSoup
import pandas as pd
import time
url_page="https://www.transfermarkt.com/detailsuche/spielerdetail/suche/27403221"
response = requests.get(url=url_page,
headers={'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/35.0.1916.47 Safari/537.36'})
response.elapsed.seconds
soup = BeautifulSoup(response.content, "html.parser")
for link in soup.find_all('table',class_='items'):
for link_pag in link.find_all(class_='spielprofil_tooltip'):
#add page loop
url_page="https://www.transfermarkt.com"+link_pag.attrs["href"]
response_pagina = requests.get(url=url_page,
headers={'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/35.0.1916.47 Safari/537.36'})
soup_pagina = BeautifulSoup(response_pagina.content, "html.parser")
time.sleep(3)
for n_player in soup_pagina('h1', itemprop="name"):
name = n_player.text
for value_player in soup_pagina('span', class_="waehrung"):
price = value_player.text
data_table = soup_pagina.find('table', class_='auflistung')
for data in data_table.find_all('tbody'):
rows = data.find_all('tr')
for row in rows:
try:
date_of_birth = row.find('td', [1]).text
except:
date_of_birth = ""
place_of_birth = row.find('td', [2]).text
age = row.find('td', [3]).text
height = row.find('td', [4]).text
citizenship = row.find('td', [5]).text
position = row.find('td', [6]).text
foot = row.find('td', [7]).text
agent = row.find('td', [8]).text
club = row.find('td', [9]).text
joined = row.find('td', [10]).text
contract_expired = row.find('td', [11]).text
contract_extension = row.find('td', [12]).text
stats_table = soup_pagina.find('table', class_='items')
for stats in stats_table.find_all('tfoot'):
rows_s = stats.find_all('td'):
for row_s in rows_s:
total = row.find('td', [3]).text
goal = row.find('td', [4]).text
assist = row.find('td', [5]).text
goal_per_min = row.find('td', [6]).text
total_min = row.find('td', [7]).text
data_stats = {
'name': name,
'price': price,
'data_of_birth': data_of_birth,
'place_of_birth': place_of_birth,
'age': age,
'height': height,
'citizenship': citizenship,
'position': position,
'foot': foot,
'agent': agent,
'club': club,
'joined': joined,
'contract_expired': contract_expired,
'contract_extension': contract_extension,
}
players_stats.append(data_stats)
players_stats = []
df = pd.DataFrame(players_stats)
print(df.head())
df.to_csv('players.csv', index=False)
【问题讨论】:
-
你忘了问你的问题。
-
更新帖子
标签: python web-scraping beautifulsoup html-table