【问题标题】:How to scrape the yahoo earnings calendar with beautifulsoup如何用beautifulsoup刮掉雅虎收益日历
【发布时间】:2019-10-28 14:32:07
【问题描述】:
如何从雅虎收入日历中提取日期?
这适用于 python 3。
from bs4 import BeautifulSoup as soup
import urllib
url = 'https://finance.yahoo.com/calendar/earnings?day=2019-06-13&symbol=ibm'
response = urllib.request.urlopen(url)
html = response.read()
page_soup = soup(html,'lxml')
table = page_soup.find('p')
print(table)
输出为“无”
【问题讨论】:
标签:
python
html
web-scraping
beautifulsoup
【解决方案1】:
这里有两种简洁的方式
import requests
from bs4 import BeautifulSoup as bs
r = requests.get('https://finance.yahoo.com/calendar/earnings?day=2019-06-13&symbol=ibm&guccounter=1')
soup = bs(r.content, 'lxml')
# using attribute = value selector
dates = [td.text for td in soup.select('[aria-label="Earnings Date"]')]
#using nth-of-type to get column
dates = [td.text for td in soup.select('#cal-res-table td:nth-of-type(3)')]
【解决方案2】:
可能跑题了,但由于您想从网页中获取表格,您可以考虑使用可用于两行的 pandas:
import pandas as pd
earnings = pd.read_html('https://finance.yahoo.com/calendar/earnings?day=2019-06-13&symbol=ibm')[0]
【解决方案3】:
Beautiful Soup 有一些 find 函数可以用来检查 DOM,请参考documentation
from bs4 import BeautifulSoup as soup
import urllib.request
url = 'https://finance.yahoo.com/calendar/earnings?day=2019-06-13&symbol=ibm'
response = urllib.request.urlopen(url)
html = response.read()
page_soup = soup(html,'lxml')
table = page_soup.find_all('td')
Dates = []
for something in table:
try:
if something['aria-label'] == "Earnings Date":
Dates.append(something.text)
except:
print('')
print(Dates)