【问题标题】:Beautiful Soup: Get text data from htmlBeautiful Soup:从 html 中获取文本数据
【发布时间】:2015-06-04 06:57:09
【问题描述】:

这是我的 html 代码,现在我想使用美丽的汤从以下 html 代码中提取数据

<tr class="tr-option">
<td class="td-option"><a href="">A.</a></td>
<td class="td-option">120 m</td>
<td class="td-option"><a href="">B.</a></td>
<td class="td-option">240 m</td>
<td class="td-option"><a href="">C.</a></td>
<td class="td-option" >300 m</td>
<td class="td-option"><a href="">D.</a></td>
<td class="td-option" >None of these</td>
</tr>

这是我漂亮的汤代码

soup = BeautifulSoup(html_doc)
for option in soup.find_all('td', attrs={'class':"td-option"}):
    print option.text

以上代码的输出:

A.
120 m
B.
240 m
C.
300 m
D.
None of these

但我想要以下输出

A.120 m
B.240 m
C.300 m
D.None of these

我该怎么办?

【问题讨论】:

    标签: python html beautifulsoup


    【解决方案1】:

    由于find_all 返回选项列表,您可以使用列表推导式来获得您期望的答案

    >>> a_list = [ option.text for option in soup.find_all('td', attrs={'class':"td-option"}) ]
    >>> new_list = [ a_list[i] + a_list[i+1] for i in range(0,len(a_list),2) ]
    >>> for option in new_list:
    ...     print option
    ... 
    A.120 m
    B.240 m
    C.300 m
    D.None of these
    

    它有什么作用?

    • [ a_list[i] + a_list[i+1] for i in range(0,len(a_list),2) ]a_list 获取相邻元素并附加它们。

    【讨论】:

      【解决方案2】:
      soup = BeautifulSoup(html_doc) 
      options = soup.find_all('td', attrs={'class': "td-option"}) 
      texts = [o.text for o in options] 
      lines = [] 
      # Add every two-element pair as a concatenated item
      for a, b in zip(texts[0::2], texts[1::2]): 
          lines.append(a + b)
      for l in lines:
          print(l)
      

      A.120 m
      B.240 m
      C.300 m
      D.None of these
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2015-03-20
        • 1970-01-01
        • 1970-01-01
        • 2015-07-17
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多