【问题标题】:Web scraping for rottentomatoes got errorrottentomatoes 的网页抓取出错了
【发布时间】:2021-06-01 13:45:19
【问题描述】:

我想从rottentomatoes抓取一个页面。

页面截图是

根据图片span class= descriptor是a class的父类,div class = info director是Directed By的gradparent。

我想刮掉导演的名字

headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.182 Safari/537.36", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}
url= 'https://editorial.rottentomatoes.com/guide/best-sci-fi-movies-of-all-time/'
r = requests.get(url, headers=headers)#, proxies=proxies)
content = r.content
soup = BeautifulSoup(content)
director = []
people1 = soup.find_all('div',{'class':'info director'})
for d in people1:
    Dir = d.find('a').text
    director.append(Dir)

我收到了这个错误

AttributeError: 'NoneType' object has no attribute 'text'

【问题讨论】:

  • 请包含您用于抓取的语言的标签。显示的代码既不是 HTML 也不是 CSS。
  • 好的。我正在使用 Python。
  • 我认为你的意思是“刮”。报废意味着扔掉。

标签: python html css web-scraping


【解决方案1】:

使用“info director”类定位 div,并使用单行将所有 href 文本转储到列表中

import requests
from bs4 import BeautifulSoup

headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.182 Safari/537.36", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}
url = 'https://editorial.rottentomatoes.com/guide/best-sci-fi-movies-of-all-time/'
r = requests.get(url, headers=headers)
soup = BeautifulSoup(r.content, 'html5lib')
directors = [a.text for a in (d.find('a') for d in soup.find_all('div', attrs={'class': 'info director'})) if a]
for x in range(len(directors)):
    print(directors[x])  # output directors

# alternative loop
directors = []
for d in soup.find_all('div', attrs={'class': 'info director'}):
    for a in d.find('a'):
        directors.append(a)
        print(a)

【讨论】:

  • 这里您没有使用 span class= 描述符。但是,a 在 span 和 div 标签下。 a for a 是什么?
  • 我没有使用跨度,因为您的目标是 a href 标签,它不是跨度的子级。这也是您一开始就收到错误的原因,因为您是说找到所有 span 的标签子项,但没有。
  • a for a 只是为了完成所有事情。它是一个嵌套循环。内部循环查找所有带有“info director”类的 span 标签,外部循环查找 a 标签。
  • 所以标签是div class = info director的子标签。如果我在info director 中找到a,我会遇到同样的错误。
  • 你能用简单的循环改变一个线性吗?
猜你喜欢
  • 2018-12-01
  • 2011-12-24
  • 2014-04-06
  • 1970-01-01
  • 1970-01-01
  • 2020-06-18
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多