【问题标题】:Extracting content of div with BeautifulSoup用 BeautifulSoup 提取 div 的内容
【发布时间】:2013-04-29 22:29:34
【问题描述】:

起初我想说,我已经找到了 the same question 的答案,但我无法让它们工作。我尝试从评论中提取数据,现在评论的内容和有用性。总的来说,我是 BeautifulSoup 和 Python 的新手。

目前,我使用 findAll 方法来获取包含评论的 div 列表,例如,一些对产品有意见的随机网站:

import urllib2
from BeautifulSoup import BeautifulSoup
turl = "http://www.amazon.com/The-Great-Gatsby-Scott-Fitzgerald/product-reviews/0743273567/ref=cm_cr_pr_hist_5?ie=UTF8&filterBy=addFiveStar&showViewpoints=0"
page= urllib2.urlopen(turl);
soup = BeautifulSoup(page);
products = soup.findAll("div", style = "margin-left:0.5em;")
print products[0]

这样我得到这样的输出:

<div style="margin-left:0.5em;">
<div style="margin-bottom:0.5em;">
        335 of 368 people found the following review helpful
      </div>
<div style="margin-bottom:0.5em;">
<span style="margin-right:5px;"><span class="swSprite s_star_5_0 " title="5.0 out of 5 stars"><span>5.0 out of 5 stars</span></span> </span>
<span style="vertical-align:middle;"><b>Decades later, still great but on different terms.</b>, <nobr>August 24, 2001</nobr></span>
</div>
<div style="margin-bottom:0.5em;">
<div><div style="float:left;">By&nbsp;</div><div style="float:left;"><a href="http://www.amazon.com/gp/pdp/profile/A1IKD6BDEE18CI"><span style="font-weight: bold;">mirope "mirope"</span></a>  - <a href="http://www.amazon.com/gp/cdp/member-reviews/A1IKD6BDEE18CI?ie=UTF8&amp;sort_by=MostRecentReview">See all my reviews</a><br />
<a href="http://www.amazon.com/gp/help/customer/display.html?ie=UTF8&amp;nodeId=14279681&amp;pop-up=1#VN" target="AmazonHelp" onclick="return amz_js_PopWin(this.href,'AmazonHelp','width=340,height=340,resizable=1,scrollbars=1,toolbar=1,status=1');"><span class="cmtySprite s_BadgeVineVoice "><span>(VINE VOICE)</span></span></a>
&nbsp;&nbsp;


</div></div><div style="clear:both;"></div>
</div>
<div class="tiny" style="margin-bottom:0.5em;">
<span class="crVerifiedStripe"><b class="h3color tiny" style="margin-right: 0.5em;">Amazon Verified Purchase</b><span class="tiny verifyWhatsThis">(<a href="http://www.amazon.com/gp/community-help/amazon-verified-purchase" target="AmazonHelp" onclick="amz_js_PopWin('http://www.amazon.com/gp/community-help/amazon-verified-purchase', 'AmazonHelp', 'width=400,height=500,resizable=1,scrollbars=1,toolbar=0,status=1');return false; ">What's this?</a>)</span></span>
</div>
<div class="tiny" style="margin-bottom:0.5em;">
<b><span class="h3color tiny">This review is from: </span><a href="https://rads.stackoverflow.com/amzn/click/com/0684801523" rel="nofollow noreferrer">The Great Gatsby (Paperback)</a></b>
</div>

Having reread this book for the first time in 20 years, I can confirm that there's a reason that it's considered one of the very best American novels. However, my reaction to the story was different than when I first read it in high school. I recall that back then I was hoping that Daisy and Gatsby's love story would ultimately yield a happy ending. Now, I found them both to be such shallow creatures that they inspired no pity. While I considered the characters to be emotionally stunted, that dooesn't mean I was not impressed with Fitzergerald's skillful rendering. As in most forms of art, in literature it is more difficult to accurately and interestingly portray nothingness than to describe a richly endowed subject. At this more cynical age, I found Daisy to be a remarkable emotional void, and Gatsby's quest to pour all of his hopes and dreams into such a shallow cauldron only confirmed his own vapidity. One thing that hasn't changed in all these years is my amazement at Fitzgerald's ability to set a scene. His descriptive passages are truly poetic, and his command of word choice in unparalleled. All this made for a stimulating and delightful read.
      <div style="padding-top: 10px; clear: both; width: 100%;">
<div class="reviews-voting-stripe" style="float:left; padding-right:15px; border-right:1px solid #CCCCCC"><div style="padding-bottom:5px;"><b class="tiny" style="color:#666666;white-space:nowrap;">Help other customers find the most helpful reviews</b>&nbsp;</div><div style="width:300px;">
<a name="R3KCIEAV000FPG.2115.Helpful.Reviews" style="font-size:1px;"> </a><span class="crVotingButtons"><nobr><span class="votingPrompt">Was this review helpful to you?&nbsp;</span><a rel="nofollow" class="votingButtonReviews votingButton-yes" href="http://www.amazon.com/gp/voting/cast/Reviews/2115/R3KCIEAV000FPG/Helpful/1?ie=UTF8&amp;target=aHR0cDovL3d3dy5hbWF6b24uY29tL3Jldmlldy8wNzQzMjczNTY3&amp;token=9BE8627F650F9D873DB4042D67CB37FA98AFD161&amp;voteAnchorName=R3KCIEAV000FPG.2115.Helpful.Reviews&amp;voteSessionID=000-0000000-0000000"><span class="cmtySprite s_largeYes "><span>Yes</span></span></a>
<a rel="nofollow" class="votingButtonReviews votingButton-no" href="http://www.amazon.com/gp/voting/cast/Reviews/2115/R3KCIEAV000FPG/Helpful/-1?ie=UTF8&amp;target=aHR0cDovL3d3dy5hbWF6b24uY29tL3Jldmlldy8wNzQzMjczNTY3&amp;token=B35087155FEB75AC5155B500CE8518AEFD4ADBAC&amp;voteAnchorName=R3KCIEAV000FPG.2115.Helpful.Reviews&amp;voteSessionID=000-0000000-0000000"><span class="cmtySprite s_largeNo "><span>No</span></span></a></nobr> <span class="votingMessage"></span></span>
</div></div><div style="float:left;"><div style="padding-left:15px;"><div style="white-space:nowrap;"><span class="tiny">
<a name="R3KCIEAV000FPG.2115.Inappropriate.Reviews" style="font-size:1px;"> </a><span class="reportingButton"><nobr><a rel="nofollow" class="reportingButton" href="http://www.amazon.com/gp/voting/cast/Reviews/2115/R3KCIEAV000FPG/Inappropriate/1?ie=UTF8&amp;target=aHR0cDovL3d3dy5hbWF6b24uY29tL3Jldmlldy8wNzQzMjczNTY3&amp;token=414B10F161A63A55D269D6EE7DC174FF22482F7E&amp;voteAnchorName=R3KCIEAV000FPG.2115.Inappropriate.Reviews&amp;voteSessionID=000-0000000-0000000">Report abuse</a></nobr></span>
</span> <span style="color:#CCCCCC;">|</span> <span class="tiny"><a href="http://www.amazon.com/review/R3KCIEAV000FPG">Permalink</a></span></div><div style="white-space:nowrap;padding-left:-5px;padding-top:5px;"><a href="http://www.amazon.com/review/R3KCIEAV000FPG"><span class="swSprite s_comment "><span>Comment</span></span></a>&nbsp;<a href="http://www.amazon.com/review/R3KCIEAV000FPG">Comments (19)</a></div></div></div><div style="clear:both;"></div>
</div>
<br />
</div>

我想从这个输出中提取两个整数 - 335 和 368(有多少人认为它有用)和包含评论本身的评论文本(只是单词,没有标签和新行)的字符串,放置在主 div 中,下 5 个子 div。我怎样才能得到这个 div 的一部分而没有其余部分,处理标签?

编辑:

我将 BeautifulSoap 返回的对象转换为字符串并加载回汤 - 还有其他方法吗?好像不太好看然后我使用你的方法,但是我得到很多空行,我尝试将它们删除条带和使用条件,但它们仍然存在:

import urllib2
from BeautifulSoup import BeautifulSoup
turl = "http://www.amazon.com/The-Great-Gatsby-Scott-Fitzgerald/product-reviews/0743273567/ref=cm_cr_pr_hist_5?ie=UTF8&filterBy=addFiveStar&showViewpoints=0"
toppage = urllib2.urlopen(turl);
soup = BeautifulSoup(toppage);
products = soup.findAll("div", style = "margin-left:0.5em;")

for (counter,i) in enumerate(products):
    soup2 = BeautifulSoup(str(products[counter]))
    for (counter2,x) in enumerate(soup2.div):
        if x.string:
            if x.string.isspace:
                print "empty string"
            else:
                print "string number " + str(counter) + " " + x.string.strip().lstrip() 
**

【问题讨论】:

  • 欢迎来到 StackOverflow!请务必对您认为有用的问题和答案进行投票,并确保您接受可以解决您的问题的答案。如果您需要,请随时要求更多说明。

标签: python beautifulsoup


【解决方案1】:

完整的最小工作示例

使用您的源网页,这是一个完整的示例

import urllib2, re
from BeautifulSoup import BeautifulSoup   

turl = "http://rads.stackoverflow.com/amzn/click/0743273567"
toppage = urllib2.urlopen(turl)
soup = BeautifulSoup(toppage)

review_tag  = {'class':re.compile("mt9 reviewText")}
helpful_tag = {'class':re.compile("hlp")}

all_reviews = soup.findAll(attrs=review_tag)
all_helpful = soup.findAll(attrs=helpful_tag)

for text,info in zip(all_reviews, all_helpful):
    print info.string.strip()
    print '\n'.join(text.findAll(text=True)).strip()
    print "*******************************************"

这给了

337 of 370 people found the following review helpful
Having reread this book for the first time in 20 years, I can confirm that there's a reason that it's [...]
*******************************************
114 of 123 people found the following review helpful
It's difficult to give any even-handed critique F. Scott Fitzgerald's standard-setting Jazz Age [...]
*******************************************
54 of 60 people found the following review helpful
Scott Fitzgerald, a monumental talent who only occasionally got things working right, made Gatsby great by the extraordinary invention of Nick Carraway.  Carraway as

旧版

这是在编辑帖子之前完成的:

假设您已将数据加载到一个名为soup 的汤中,令人难以置信

for x in soup.body.div:
    if x.string:
        print x.string.strip()

给予:

335 of 368 people found the following review helpful

Having reread this book for the first time in 20 years, [... more here]

您要查找哪些字符串。

就这么简单吗?

html 可能是一团糟,所以让我给您一些提示,帮助您在新网页中爬行。首先我找到了文字:

import re
x = soup.find(text=re.compile('Having reread this book'))

然后我通过父母来了解我在调查什么:

print x.parent
print x.parent.parent
print x.parent.parent.parent

从那里我看到所有内容都作为字符串包含在主 div 中。然后简单地循环遍历我正在寻找的内容!

【讨论】:

  • 有道理,但products 不是soup 对象,它只是一个列表,我不能为此调用.body.div 方法。
  • @sirVir 我为您提供了一个完整的最小示例,可以满足您对编辑的要求。
  • 非常感谢您的帮助,但它对我来说仍然非常缓慢...我想访问 reviews、helpfulness 和评论日期 - 由date_tag = {'class':re.compile("inlineblock txtsmall")} 和评论标题(仅使用title_tag = {'class':re.compile("txtlarge gl3 gr4 reviewTitle valignMiddle")} 似乎不起作用)以字符串的形式找到。我不明白您的打印机制究竟是如何工作的,而且我很难在其中添加任何内容。您能给我一些指导吗?
猜你喜欢
  • 2021-04-15
  • 2014-11-29
  • 2015-01-05
  • 1970-01-01
  • 2014-10-26
  • 1970-01-01
  • 2017-01-29
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多