【问题标题】:Beautiful Soup: get contents of search result tagBeautiful Soup:获取搜索结果标签的内容
【发布时间】:2017-02-19 11:37:12
【问题描述】:

尝试使用漂亮的汤(它是一个“标签”对象)获取这种类型的html sn-p的内容。

<span class="font5"> arrives at this calculation from the Torah’s report that the deluge (rains) began on the 17<sup>th</sup> day of the second month </span>

我试过了:

soup.contents.find_all('span')
soup.find_all('span')
soup.find_all(re.compile("font[0-9]+"))
soup.string
soup.child

而且这些似乎都不起作用。我能做什么?

【问题讨论】:

  • 你是说soup.find_all('span') 什么都没找到,太多了,还是找到了一些但不包括你想要的东西?
  • 所有这些要么不返回任何内容,要么返回错误。
  • 那么您要查找的内容不在源代码中。

标签: python beautifulsoup


【解决方案1】:

soup.find_all('span') 确实有效;返回所有span 标签。

如果您想获得带有font&lt;N&gt; 类的span 标记,请将模式指定为关键字参数class_:

soup.find_all('span', class_=re.compile('font[0-9]+'))

【讨论】:

  • 这两个都返回[],没有内容...请注意这是一个标签对象,不是汤对象。
  • @EsterLin,您能否提供重现您的问题的最小工作代码?
【解决方案2】:

如果以 font 开头的足够独特,您还可以使用 css 选择器 来寻找以 font 开头的类:

soup.select("span[class^=font]")

【讨论】:

    【解决方案3】:
    print ''.join(soup.findAll(text=True))
    

    (已回复here)

    【讨论】:

      猜你喜欢
      • 2014-07-11
      • 2021-02-22
      • 1970-01-01
      • 1970-01-01
      • 2020-08-24
      • 2017-12-31
      • 1970-01-01
      • 2021-12-04
      • 1970-01-01
      相关资源
      最近更新 更多