【问题标题】:Scraping and appending data while looping through html tables在遍历 html 表时抓取和附加数据
【发布时间】:2020-11-19 09:14:23
【问题描述】:

我想在 Yahoo Finance 中循环查看 30 个日期的收益发布信息,同时将每个日期的信息附加到单个数据框中。我已使用此代码,但无法获得包含所有信息的合并数据框 - 它仅显示最后日期的信息:

import datetime
import pandas as pd

date = (datetime.date.today() + datetime.timedelta(days=1)).isoformat() #get tomorrow in iso format as needed'''

for i in range(30):
    try: 
        date = (datetime.date.today() + datetime.timedelta(days = i )).isoformat() #get tomorrow in iso format as needed'''
        pd.set_option('display.max_column',None)
        url = pd.read_html("https://finance.yahoo.com/calendar/earnings?day="+date, header=0)
        table = url[0]
        table.append(table)
        print(table)
    except ValueError:
        continue

【问题讨论】:

    标签: python pandas selenium web-scraping


    【解决方案1】:

    您每次迭代都会覆盖您的表。初始化一个列表,然后将每个表附加到该列表中。最后将数据帧列表迭代为 1 个数据帧:

    import datetime
    import pandas as pd
    
    date = (datetime.date.today() + datetime.timedelta(days=1)).isoformat() #get tomorrow in iso format as needed'''
    
    tables = [] #<-- initialize an empty list to store your tables
    for i in range(30):
        try: 
            date = (datetime.date.today() + datetime.timedelta(days = i )).isoformat() #get tomorrow in iso format as needed'''
            pd.set_option('display.max_column',None)
            url = pd.read_html("https://finance.yahoo.com/calendar/earnings?day="+date, header=0)
            table = url[0]
            tables.append(table) #<-- apend each table into your list of tables
            print(table)
        except ValueError:
            print ('Error')
            continue
    
    
    df = pd.concat(tables) #<-- take your list of tables into 1 final dataframe
    

    【讨论】:

      【解决方案2】:

      在循环外创建一个列表以保存您的表格,然后将表格附加到 lst_tables:

      lst_tables = list()
      for i in range(30):
          try: 
              date = (datetime.date.today() + datetime.timedelta(days = i )).isoformat() #get tomorrow in iso format as needed'''
              pd.set_option('display.max_column',None)
              url = pd.read_html("https://finance.yahoo.com/calendar/earnings?day="+date, header=0)
              lst_tables.append(url[0])
              print(url[0]) # if you really need to print it out for checking
          except ValueError:
              continue
      

      您可以使用 pd.concat(lst_tables) 来获得完整的数据帧

      df_completed = pd.concat(lst_tables)
      

      【讨论】:

      • 谢谢!但是,我看到这会创建一个列表,其中包含一堆数据框作为生活在其中的项目,而不是一个数据框,所有数据都堆叠在另一个之上...
      • @JuneSmith,您好,我编辑了答案,您可以使用 pd.concat 将所有数据堆叠在一个数据帧中
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2017-10-23
      • 2012-05-20
      • 2018-10-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多