【发布时间】:2023-03-05 07:19:02
【问题描述】:
我正在尝试抓取 http://www.basketball-reference.com/awards/all_league.html 进行一些分析,我的目标如下所示
0 1 马克加索尔 2014-2015
1 第一届安东尼戴维斯 2014-2015
2 2014-2015 年第 1 位勒布朗·詹姆斯
3 1 詹姆斯哈登 2014-2015
4 第一名斯蒂芬库里 2014-2015
5 第二保罗加索尔 2014-2015 等等
这是我到目前为止的代码,有没有办法做到这一点?非常感谢任何建议/帮助。
r = requests.get('http://www.basketball-reference.com/awards/all_league.html')
soup=BeautifulSoup(r.text.replace(' ','').replace('>','').encode('ascii','ignore'),"html.parser")
all_league_data = pd.DataFrame(columns = ['year','team','player'])
stw_list = soup.findAll('div', attrs={'class': 'stw'}) # Find all 'stw's'
for stw in stw_list:
table = stw.find('table', attrs = {'class':'no_highlight stats_table'})
for row in table.findAll('tr'):
col = row.findAll('td')
if col:
year = col[0].find(text=True)
team = col[2].find(text=True)
player = col[3].find(text=True)
all_league_data.loc[len(all_league_data)] = [team, player, year]
all_league_data
【问题讨论】:
标签: python python-2.7 pandas web-scraping beautifulsoup