【问题标题】:Why is value "External_links" and not items scraped from website?为什么价值“External_links”而不是从网站上抓取的项目?
【发布时间】:2018-12-29 21:31:23
【问题描述】:

我的代码在下面,但是为什么brand值输出External_links而不是我拉的项目列表。

from bs4 import BeautifulSoup as soup
from urllib.request import urlopen as uReq


my_url = 'https://en.wikipedia.org/wiki/Harry_Potter'
uClient = uReq(my_url)
page_html = uClient.read()
uClient.close()

page_soup = soup(page_html,"html.parser")
headline = page_soup.findAll("span",{"class":"mw-headline"})

for item in headline:
    brand = item["id"] # Outputs "External_links"

【问题讨论】:

    标签: python web-scraping beautifulsoup urllib


    【解决方案1】:

    在您的for 循环中,您将遍历页面中的每个标题,然后将标题值分配给变量brand。循环结束后,brand 的值将成为最后一个标题(“External_links”)。

    如果您修改代码以打印出每个标题的值,您将看到您得到了您正在寻找的值。

    >>> for item in headline:
    ...    print(item["id"])
    ...
    Plot
    Early_years
    Voldemort_returns
    Supplementary_works
    Harry_Potter_and_the_Cursed_Child
    In-universe_books
    Pottermore_website
    Structure_and_genre
    Themes
    Origins
    Publishing_history
    Translations
    Completion_of_the_series
    Cover_art
    Achievements
    Cultural_impact
    Commercial_success
    Awards,_honours,_and_recognition
    Reception
    Literary_criticism
    Social_impact
    Controversies
    Adaptations
    Films
    Spin-off_prequels
    Games
    Audiobooks
    Stage_production
    Attractions
    The_Wizarding_World_of_Harry_Potter
    The_Making_of_Harry_Potter
    References
    Further_reading
    External_links
    

    【讨论】:

    • 但我需要将它分配给一个变量,以便我可以导出到 csv,这可能吗?
    • 您可以创建一个列表,然后将每个标题添加到列表中,然后将列表的内容写入文件。
    【解决方案2】:

    您的brand 变量需要是一个列表,例如代码可能如下所示:

    from bs4 import BeautifulSoup as soup
    from urllib.request import urlopen as uReq
    from pprint import pprint
    
    my_url = 'https://en.wikipedia.org/wiki/Harry_Potter'
    with uReq(my_url) as uClient:
        page_html = uClient.read()
        page_soup = soup(page_html, "xml")
    
    brand = []
    for item in page_soup.find_all('span', {'class': 'mw-headline'}):
        brand.append(item["id"])
    
    pprint(brand)
    

    打印:

    ['Plot',
     'Early_years',
     'Voldemort_returns',
     'Supplementary_works',
     'Harry_Potter_and_the_Cursed_Child',
     'In-universe_books',
     'Pottermore_website',
     'Structure_and_genre',
     'Themes',
     'Origins',
     'Publishing_history',
     'Translations',
     'Completion_of_the_series',
     'Cover_art',
     'Achievements',
     'Cultural_impact',
     'Commercial_success',
     'Awards,_honours,_and_recognition',
     'Reception',
     'Literary_criticism',
     'Social_impact',
     'Controversies',
     'Adaptations',
     'Films',
     'Spin-off_prequels',
     'Games',
     'Audiobooks',
     'Stage_production',
     'Attractions',
     'The_Wizarding_World_of_Harry_Potter',
     'The_Making_of_Harry_Potter',
     'References',
     'Further_reading',
     'External_links']
    

    【讨论】:

      【解决方案3】:

      使用列表推导实现同样的目的:

      import requests
      from bs4 import BeautifulSoup
      from pprint import pprint
      
      url = 'https://en.wikipedia.org/wiki/Harry_Potter'
      
      soup = BeautifulSoup(requests.get(url).text, "lxml")
      items = [item.get('id') for item in soup.find_all('span',class_='mw-headline')]
      pprint(items)
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2022-01-08
        • 2020-03-10
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多