【发布时间】:2021-06-01 13:45:19
【问题描述】:
我想从rottentomatoes抓取一个页面。
页面截图是
根据图片span class= descriptor是a class的父类,div class = info director是Directed By的gradparent。
我想刮掉导演的名字
headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.182 Safari/537.36", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}
url= 'https://editorial.rottentomatoes.com/guide/best-sci-fi-movies-of-all-time/'
r = requests.get(url, headers=headers)#, proxies=proxies)
content = r.content
soup = BeautifulSoup(content)
director = []
people1 = soup.find_all('div',{'class':'info director'})
for d in people1:
Dir = d.find('a').text
director.append(Dir)
我收到了这个错误
AttributeError: 'NoneType' object has no attribute 'text'
【问题讨论】:
-
请包含您用于抓取的语言的标签。显示的代码既不是 HTML 也不是 CSS。
-
好的。我正在使用 Python。
-
我认为你的意思是“刮”。报废意味着扔掉。
标签: python html css web-scraping