【问题标题】:Getting pandas.read_html( ) to work when HTML table contains more than one <tbody> tag当 HTML 表格包含多个 <tbody> 标签时,让 pandas.read_html() 工作
【发布时间】:2015-10-23 17:49:59
【问题描述】:

我正在尝试解析在http://www.swiftcodesbic.com 找到的表格,并且我正在使用Pandas 自动抓取表格。在大多数情况下,这工作正常,但有一个表有两个&lt;tbody&gt; 标签,我认为它会导致打嗝。故障表可以在here找到。

我用来将 html 解析为 pandas.DataFrame 的代码是:

pandas.read_html(countryPage.text, attrs={"id":"t2"}, skiprows=1)[0]

其中countryPage 是requests.get() 对象。我可以在 pandas 调用中添加任何内容来告诉它获取第二个 &lt;tbody&gt; 标签吗?或者,如果这不是问题,有人可以解释可能导致它返回“未找到表”错误的原因吗?提前致谢。

编辑

这是我目前正在使用的解决方案,但我仍然想知道一个更“pythonic”的方法。

try:
  tempDataFrame = pd.read_html(countryPage.text, attrs={"id":"t2"}, skiprows=1)[0]
except:
  if "france" is in url: #pseudo-code
    soup = BeautifulSoup(countryPage.text)
    table = soup.find_all("table")[2].findAll('tbody')[1] #this will vary based on your situation
    table = "<table>" + str(table) + "</table>" #pandas needs the table tag to recognize a table
    tempDataFrame = pd.read_html(table)[0]

同样,我很想知道如何以更有效的方式做到这一点。

【问题讨论】:

    标签: python html pandas


    【解决方案1】:

    使用match 参数应该可以解决问题。来自 pandas.read_html 文档:

    match : str or compiled regular expression, optional
    
        The set of tables containing text matching this regex or string will be returned. Unless the HTML is extremely simple you will probably need to pass a non-empty string here. Defaults to ‘.+’ (match any non-empty string). The default value will return all tables contained on a page. This value is converted to a regular expression so that there is consistent behavior between Beautiful Soup and lxml.
    

    试试这种方法

    tempDataFrame = pd.read_html(countryPage.text, match='foo', skiprows=1)
    

    其中 foo 是表中包含的字符串

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-02-06
      • 1970-01-01
      • 2021-08-21
      • 2012-03-24
      • 2016-04-28
      • 1970-01-01
      • 1970-01-01
      • 2011-06-19
      相关资源
      最近更新 更多