【问题标题】:Reformat line/div extracted from html重新格式化从 html 中提取的 line/div
【发布时间】:2020-08-14 15:21:48
【问题描述】:

我目前无法重新格式化从网站提取的 div。

这是我目前拥有的:

<div class=" frame frame-default frame-type-textmedia frame-layout-0" id="c47903"><a id="c47904"/><div class="ce-textpic ce-left ce-above"><div class="ce-bodytext"><p>The latest data of the evolution of COVID-19 over the past 24hours <strong>in Québec</strong> reveal:</p><ul><li>87new cases, bringing the total number of infected persons to61,004;</li><li>no deaths have occurred in the past 24hours, to which are added 3deaths which occurred between August7 and12, for a total of5,718;</li><li>the number of hospitalizations increased by2 compared to the previous day, for a cumulative total of151. Of these, 25were in intensive care, an increase of2;</li><li>18,596tests were performed on August12, for a cumulative total of1,428,286.</li></ul></div></div></div> 

但我想要类似的东西:

魁北克过去 24 小时 COVID-19 演变的最新数据:87 例新病例,使感染者总数达到 61,004 人;过去 24 小时内未发生死亡,加上 8 月 7 日至 12 日之间发生的 3 例死亡,总数为 5,718;住院人数较前一天增加2,累计达151人。其中 25 人在重症监护室,增加了 2 人;8 月 12 日进行了 18,596 次检查,累计总数为 1,428,286 次。

我手动删除了它,但是有没有一些耗时较少的东西?

【问题讨论】:

  • 你试过用get_text()
  • 这能回答你的问题吗? Extracting text from HTML file using Python
  • @TheLazyScripter 只是一堆 str(div).replace("我想要的东西","")
  • 如果您已经捕获了带有bs4 的标签,请尝试在标签上使用get_text()

标签: python html python-3.x beautifulsoup extract


【解决方案1】:

试试这个

text = r'<div class=" frame frame-default frame-type-textmedia frame-layout-0" id="c47903"><a id="c47904"/><div class="ce-textpic ce-left ce-above"><div class="ce-bodytext"><p>The latest data of the evolution of COVID-19 over the past 24hours <strong>in Québec</strong> reveal:</p><ul><li>87new cases, bringing the total number of infected persons to61,004;</li><li>no deaths have occurred in the past 24hours, to which are added 3deaths which occurred between August7 and12, for a total of5,718;</li><li>the number of hospitalizations increased by2 compared to the previous day, for a cumulative total of151. Of these, 25were in intensive care, an increase of2;</li><li>18,596tests were performed on August12, for a cumulative total of1,428,286.</li></ul></div></div></div>'
import re
print(re.sub(r'<[^<>]*>', ' ', text))

【讨论】:

  • 你可以在 **Python 3 中尝试 f"{variableforthevalue}" ** 对于 Python 2 你必须使用 "% s"%(变量值)
【解决方案2】:

尝试类似:

soup.select_one('div[class="ce-bodytext"]').text.strip()

这应该会得到您预期的输出。

【讨论】:

    【解决方案3】:

    试试

    str(bs4_obj.select('div')[0].text)
    

    我不知道如何从 unicode 转换它, 但它摆脱了 html 标签。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2019-05-15
      • 1970-01-01
      • 1970-01-01
      • 2011-01-27
      • 2013-05-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多