【发布时间】:2015-10-23 17:49:59
【问题描述】:
我正在尝试解析在http://www.swiftcodesbic.com 找到的表格,并且我正在使用Pandas 自动抓取表格。在大多数情况下,这工作正常,但有一个表有两个<tbody> 标签,我认为它会导致打嗝。故障表可以在here找到。
我用来将 html 解析为 pandas.DataFrame 的代码是:
pandas.read_html(countryPage.text, attrs={"id":"t2"}, skiprows=1)[0]
其中countryPage 是requests.get() 对象。我可以在 pandas 调用中添加任何内容来告诉它获取第二个 <tbody> 标签吗?或者,如果这不是问题,有人可以解释可能导致它返回“未找到表”错误的原因吗?提前致谢。
编辑
这是我目前正在使用的解决方案,但我仍然想知道一个更“pythonic”的方法。
try:
tempDataFrame = pd.read_html(countryPage.text, attrs={"id":"t2"}, skiprows=1)[0]
except:
if "france" is in url: #pseudo-code
soup = BeautifulSoup(countryPage.text)
table = soup.find_all("table")[2].findAll('tbody')[1] #this will vary based on your situation
table = "<table>" + str(table) + "</table>" #pandas needs the table tag to recognize a table
tempDataFrame = pd.read_html(table)[0]
同样,我很想知道如何以更有效的方式做到这一点。
【问题讨论】: