【问题标题】:text scraping with beautiful soup and python is not working用漂亮的汤和 python 刮文字不起作用
【发布时间】:2017-09-22 20:21:34
【问题描述】:

我正在尝试从这个 html 中抓取文本

  <table class="table table-hover table-condensed">
<tbody>
    <tr style="background-color: aliceblue">
    </tr>

    <tr>
        <td class="text-center" style="width: 10%">
            <img class="img-thumbnail" src="/Files/image/placeholder100.png" style="width: 100px">
        </td>

        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/NSN/1520-00-087-7637">1520-00-087-7637</a></td>
        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/PartNumber/UH1H">UH1H</a></td>

        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/CAGE/97499">97499</a></td>
        <td class="text-center" style="vertical-align: middle; width: 10%"><a href="/CAGE/97499"><img class="img-thumbnail" src="/Files/cage/90/97499.jpg" title="CAGE 97499" alt="CAGE 97499"></a></td>
        <td nowrap="" style="vertical-align: middle">

            <h4>&emsp;&emsp; MAMA,BEAR</h4>


            <p>
                <em>&emsp;&emsp;&emsp;&emsp;Alternate References: <a href="/NSN/1520-00-087-7637">1520-00-087-7637</a>, <a href="/NSN/1520-00-087-7637">000877637</a></em>
            </p>
        </td>
    </tr>

我必须分别提取以下值

妈妈,熊

1520-00-087-7637

所以我尝试使用此代码

tablecontainer = page_soup1.find_all("tr")

    for container in tablecontainer:

            NSN = container.find("td", {"class": "text-center"}).a.text

            print(NSN)

当我运行代码时,我得到了

AttributeError: ResultSet 对象没有属性'find'

我做错了什么以及如何提取值

【问题讨论】:

  • 您要准确删除哪个1520-00-087-7637?
  • @coder 文本不是 标记中的文本
  • 就在导致错误的行之前,打印container。它是什么,是你所期望的吗?

标签: python beautifulsoup


【解决方案1】:

试试这个:

#!/usr/bin/env python

from bs4 import BeautifulSoup 

data = '''
<table class="table table-hover table-condensed">
<tbody>
    <tr style="background-color: aliceblue">
    </tr>

    <tr>
        <td class="text-center" style="width: 10%">
            <img class="img-thumbnail" src="/Files/image/placeholder100.png" style="width: 100px">
        </td>

        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/NSN/1520-00-087-7637">1520-00-087-7637</a></td>
        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/PartNumber/UH1H">UH1H</a></td>

        <td class="text-center" nowrap="" style="vertical-align: middle; width: 10%"><a href="/CAGE/97499">97499</a></td>
        <td class="text-center" style="vertical-align: middle; width: 10%"><a href="/CAGE/97499"><img class="img-thumbnail" src="/Files/cage/90/97499.jpg" title="CAGE 97499" alt="CAGE 97499"></a></td>
        <td nowrap="" style="vertical-align: middle">

            <h4>&emsp;&emsp; MAMA,BEAR</h4>


            <p>
                <em>&emsp;&emsp;&emsp;&emsp;Alternate References: <a href="/NSN/1520-00-087-7637">1520-00-087-7637</a>, <a href="/NSN/1520-00-087-7637">000877637</a></em>
            </p>
        </td>
    </tr>
'''

soup = BeautifulSoup(data, 'html.parser')
# extract the text from <h4> tag
for i in soup.find_all('h4'): print i.text.strip()
# extract the text from the link which contains the word 'NSN' in the href attribute
print soup.find(lambda tag: tag.name=='a' and 'NSN' in tag.attrs['href']).text

另外,如果您不喜欢使用lambda,您可以使用regex,如下所示:

import re
re.compile(r'[NSN]*.')
print soup.find('a', attrs={'href':pt}).text

【讨论】:

  • 就像 print 语句是一个无效的语法 print i.text.strip() ^ SyntaxError: invalid syntax
  • @learner101 我没有这样的问题。检查任何缩进错误。也尝试按原样 cp 粘贴我的答案并再次检查。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2018-01-13
  • 2017-12-23
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-06-14
  • 2017-08-15
相关资源
最近更新 更多