【问题标题】:findall returning only the last attributefindall 只返回最后一个属性
【发布时间】:2018-11-23 13:26:29
【问题描述】:

我搜索了类似的问题,但没有找到我需要的。

我正在网上搜索两个属性,在这种情况下是 red 和 green in span

from urllib.request import urlopen
from bs4 import BeautifulSoup
html=urlopen('http://www.pythonscraping.com/pages/warandpeace.html')
soup=BeautifulSoup(html,'html.parser')
nameList=soup.findAll("span",{"class":"red","class":"green"})
print(nameList)

但是我只获得了绿色属性,我尝试使用

nameList,nameList2=soup.findAll("span",{"class":"red","class":"green"})

但我收到错误ValueError: too many values to unpack (expected 2) 有没有办法同时打印并将每个属性存储在名单中(不使用多个findAll)

【问题讨论】:

    标签: python web-scraping findall


    【解决方案1】:

    您可以尝试使用 CSS 选择器来匹配 span 和两个类名,如下所示:

    nameList = soup.select("span.red, span.green")
    

    如果你还想用findAll,试试

    nameList = soup.findAll("span",{"class":["red", "green"]})
    

    【讨论】:

    • 它有效,有没有办法单独存储红色和绿色?
    • @timmy , soup.select("span.red, span.green") 两个都选,soup.select("span.green") 只选绿色,soup.select("span.red") 只选红色
    • @Andersson 我明白,这是一个很好的方法。但是为什么findAll 不存储这两个属性?
    • @timmy ,因为您应该将类​​名作为列表传递:soup.findAll("span",{"class":["red","green"]})
    • 这种方法有效,但没有办法单独存储每个人吗?
    【解决方案2】:

    因为你的红色和绿色是唯一的类属性,你可以用类属性检查 span

    from urllib.request import urlopen
    from bs4 import BeautifulSoup
    html=urlopen('http://www.pythonscraping.com/pages/warandpeace.html')
    soup=BeautifulSoup(html,'html.parser')
    nameList=soup.select("span[class]")
    print(nameList)
    

    要拥有单独的列表,您可以按类名使用 2 个选择:

    reds = soup.select('span.red')
    greens = soup.select('span.green')
    print(reds,greens)
    

    【讨论】:

    • 如何在不使用多选的情况下单独存储红色和绿色?
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-07-15
    • 1970-01-01
    • 1970-01-01
    • 2023-01-15
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多