【问题标题】:Removing particular content from result parces using beautifulsoup使用 beautifulsoup 从结果包中删除特定内容
【发布时间】:2015-12-05 06:35:22
【问题描述】:
def get_description(link):
    redditFile = urllib2.urlopen(link)
    redditHtml = redditFile.read()
    redditFile.close()
    soup = BeautifulSoup(redditHtml)
    desc = soup.find('div', attrs={'class': 'op_gd14 FL'}).text
    return desc

这是给我这个 html 文本的代码

    <div class="op_gd14 FL">
    <p><span class="bigT">P</span>restige Estates Projects Ltd has informed BSE that the 18th Annual General Meeting (AGM) of the Company will be held on September 30, 2015.Source : BSE<br><br>  
<a href="../../company-notices/nestleindia/notices/PEP02">Read all announcements in Prestige Estate</a>  </p><p>                                                </p>

</div>

这个结果对我来说很好,我只是想排除

的内容

&lt;a href="../../company-notices/nestleindia/notices/PEP02"&gt;Read all announcements in Prestige Estate&lt;/a&gt;

来自结果,即我的脚本中的desc,如果它存在,则忽略它不存在。我该怎么做?

【问题讨论】:

    标签: python regex web-scraping beautifulsoup


    【解决方案1】:

    只需对最后一行进行一些更改并添加 re 模块

    ...
    return re.sub(r'<a(.*)</a>','',desc)
    

    输出:

    '<div class="op_gd14 FL">\n    <p><span class="bigT">P</span>restige Estates Projects Ltd has informed BSE that the 18th Annual General Meeting (AGM) of the Company will be held on September 30, 2015.Source : BSE<br><br>  \n  </p><p> 
    

    【讨论】:

    【解决方案2】:

    您可以使用extract() 从find() 结果中删除不必要的标签:

    descItem = soup.find('div', attrs={'class': 'op_gd14 FL'}) # get the DIV
    [s.extract() for s in descItem('a')]                       # remove <a> tags
    return descItem.get_text()                                 # return the text
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2013-05-01
      • 1970-01-01
      • 2015-02-20
      • 1970-01-01
      • 2016-02-12
      • 2023-03-27
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多