【问题标题】:Convert multiple columns into single based on another column value python基于另一列值python将多列转换为单列
【发布时间】:2023-03-05 07:19:02
【问题描述】:

我正在尝试抓取 http://www.basketball-reference.com/awards/all_league.html 进行一些分析,我的目标如下所示

0 1 马克加索尔 2014-2015
1 第一届安东尼戴维斯 2014-2015
2 2014-2015 年第 1 位勒布朗·詹姆斯
3 1 詹姆斯哈登 2014-2015
4 第一名斯蒂芬库里 2014-2015
5 第二保罗加索尔 2014-2015 等等

这是我到目前为止的代码,有没有办法做到这一点?非常感谢任何建议/帮助。

r = requests.get('http://www.basketball-reference.com/awards/all_league.html')
soup=BeautifulSoup(r.text.replace(' ','').replace('>','').encode('ascii','ignore'),"html.parser")
all_league_data = pd.DataFrame(columns = ['year','team','player']) 


stw_list = soup.findAll('div', attrs={'class': 'stw'}) # Find all 'stw's'
for stw in stw_list:
    table = stw.find('table', attrs = {'class':'no_highlight stats_table'})
    for row in table.findAll('tr'):
        col = row.findAll('td')
        if col:
            year = col[0].find(text=True)
            team = col[2].find(text=True)
            player = col[3].find(text=True)
            all_league_data.loc[len(all_league_data)] = [team, player, year]
    all_league_data

【问题讨论】:

    标签: python python-2.7 pandas web-scraping beautifulsoup


    【解决方案1】:

    看起来您的代码应该可以正常工作,但这里有一个没有 pandas 的工作版本:

    import requests
    from bs4 import BeautifulSoup
    
    r = requests.get('http://www.basketball-reference.com/awards/all_league.html')
    soup=BeautifulSoup(r.text.replace(' ','').replace('>','').encode('ascii','ignore'),"html.parser")
    all_league_data = []
    
    stw_list = soup.findAll('div', attrs={'class': 'stw'}) # Find all 'stw's'
    for stw in stw_list:
        table = stw.find('table', attrs = {'class':'no_highlight stats_table'})
        for row in table.findAll('tr'):
            col = row.findAll('td')
            if col:
                year = col[0].find(text=True)
                team = col[2].find(text=True)
                player = col[3].find(text=True)
                all_league_data.append([team, player, year])
    
    for i, line in enumerate(all_league_data):
        print(i, *line)
    

    【讨论】:

      【解决方案2】:

      您已经在使用 pandas,所以请使用 read_html

      import pandas as pd
      
      all_league_data = pd.read_html('http://www.basketball-reference.com/awards/all_league.html')
      print(all_league_data)
      

      这将为您提供数据框中的所有表格数据:

        In [7]:  print(all_league_data[0].dropna().head(5))
               0    1    2                 3                   4  \
      0  2014-15  NBA  1st      Marc Gasol C     Anthony Davis F   
      1  2014-15  NBA  2nd       Pau Gasol C  DeMarcus Cousins C   
      2  2014-15  NBA  3rd  DeAndre Jordan C        Tim Duncan F   
      4  2013-14  NBA  1st     Joakim Noah C      LeBron James F   
      5  2013-14  NBA  2nd   Dwight Howard C     Blake Griffin F   
      
                           5                6                    7  
      0       LeBron James F   James Harden G      Stephen Curry G  
      1  LaMarcus Aldridge F     Chris Paul G  Russell Westbrook G  
      2      Blake Griffin F   Kyrie Irving G      Klay Thompson G  
      4       Kevin Durant F   James Harden G         Chris Paul G  
      5         Kevin Love F  Stephen Curry G        Tony Parker G  
      

      根据您的喜好重新排列或删除某些列将是微不足道的,read_html 需要一些参数,例如您也可以应用的 attrs,它们都在链接中。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2019-05-19
        • 2021-05-07
        • 2020-03-24
        • 1970-01-01
        • 2014-03-06
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多