【问题标题】:Python BeautifulSoup webcrawling: Appending piece of data to listPython BeautifulSoup 网络爬虫:将一条数据附加到列表中
【发布时间】:2015-09-12 19:47:59
【问题描述】:

我要抓取的网站是http://www.boxofficemojo.com/yearly/chart/?yr=2013&p=.htm。我现在关注的特定页面是http://www.boxofficemojo.com/movies/?id=catchingfire.htm。

我需要获得“外国总收入”金额(在 Total Lifetime Grosses 下),但由于某种原因,我无法通过循环获得它,以便它通过所有电影,但它适用于我输入的单个链接。

这是我为每部电影获取此金额的函数。

def getForeign(item_url):
    s = urlopen(item_url).read()
    soup = BeautifulSoup(s)
    return soup.find(text="Foreign:").find_parent("td").find_next_sibling("td").get_text(strip = True)

这是遍历每个链接的函数

def spider(max_pages):
    page = 1
    while page <= max_pages:
        url = 'http://www.boxofficemojo.com/yearly/chart/?page=' + str(page) + '&view=releasedate&view2=domestic&yr=2013&p=.htm'
        source_code = requests.get(url)
        plain_text = source_code.text
        soup = BeautifulSoup(plain_text)
        for link in soup.select('td > b > font > a[href^=/movies/?]'):
            href = 'http://www.boxofficemojo.com' + link.get('href')
            details(href)
            listOfDirectors.append(getDirectors(href))
            str(listOfDirectors).replace('[','').replace(']','')
            #getActors(href)
            title = link.string
            listOfTitles.append(title)
        page += 1

我有一个名为 listOfForeign = [] 的列表,我想将每部电影的外国总金额附加到该列表中。 问题是,如果我使用输入的单个完整链接调用 getForeign(item_url),例如:

print listOfForeign.append(getForeign(http://www.boxofficemojo.com/movies/?id=catchingfire.htm))

后来

print listOfForeign

它打印出一个正确的数量。

但是当我运行函数 spider(max_pages) 并添加:

listOfForeign.append(getForeign(href)) 

在for循环中,稍后尝试将listOfForeign打印出来,我得到一个错误

AttributeError: 'NoneType' object has no attribute 'find_parent'

为什么我无法在蜘蛛函数中为每部电影成功添加此数量?在 spider(max_pages) 函数中,我在变量“href”中获取每个电影的链接,并且基本上做的事情与分别添加每个单独的链接相同。

完整代码:

import requests
from bs4 import BeautifulSoup
from urllib import urlopen
import xlwt
import csv
from tempfile import TemporaryFile

listOfTitles = []
listOfGenre = []
listOfRuntime = []
listOfRatings = []
listOfBudget = []
listOfDirectors = []
listOfActors = []
listOfForeign = []
resultFile = open("movies.csv",'wb')
wr = csv.writer(resultFile, dialect='excel')

def spider(max_pages):
    page = 1
    while page <= max_pages:
        url = 'http://www.boxofficemojo.com/yearly/chart/?page=' + str(page) + '&view=releasedate&view2=domestic&yr=2013&p=.htm'
        source_code = requests.get(url)
        plain_text = source_code.text
        soup = BeautifulSoup(plain_text)
        for link in soup.select('td > b > font > a[href^=/movies/?]'):
            href = 'http://www.boxofficemojo.com' + link.get('href')
            details(href)
            listOfForeign.append(getForeign(href))
            listOfDirectors.append(getDirectors(href))
            str(listOfDirectors).replace('[','').replace(']','')
            #getActors(href)
            title = link.string
            listOfTitles.append(title)
        page += 1


def getDirectors(item_url):
    source_code = requests.get(item_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text)
    tempDirector = []
    for director in soup.select('td > font > a[href^=/people/chart/?view=Director]'):
        tempDirector.append(str(director.string))
    return tempDirector

def getActors(item_url):
    source_code = requests.get(item_url)
    plain_text = source_code.text
    soup = BeautifulSoup(plain_text)
    tempActors = []
    print soup.find(text="Actors:").find_parent("tr").text[7:]



def details(href):
    response = requests.get(href)
    soup = BeautifulSoup(response.content)
    genre = soup.find(text="Genre: ").next_sibling.text
    rating = soup.find(text='MPAA Rating: ').next_sibling.text
    runtime = soup.find(text='Runtime: ').next_sibling.text
    budget = soup.find(text='Production Budget: ').next_sibling.text

    listOfGenre.append(genre)
    listOfRuntime.append(runtime)
    listOfRatings.append(rating)
    listOfBudget.append(budget)


def getForeign(item_url):
    s = urlopen(item_url).read()
    soup = BeautifulSoup(s)
    try:
        return     soup.find(text="Foreign:").find_parent("td").find_next_sibling("td").get_text(strip = True)
    except AttributeError:
        return "$0"

spider(1)

print listOfForeign
wr.writerow(listOfTitles)
wr.writerow(listOfGenre)
wr.writerow(listOfRuntime)
wr.writerow(listOfRatings)
wr.writerow(listOfBudget)
for item in listOfDirectors:
    wr.writerow(item)

【问题讨论】:

  • 第一页会出现这种情况吗?

标签: python web-scraping beautifulsoup web-crawler html-parsing


【解决方案1】:

代码一旦进入电影页面就会失败没有外国收入,例如42。你应该处理这样的情况。例如,捕获异常并将其设置为$0。

您也正在体验differences between parsers - specify the lxml or html5lib parser explicitly(您需要拥有lxml 或html5lib installed)。

另外,为什么不使用requests 来解析电影页面呢:

def getForeign(item_url):
    response = requests.get(item_url)
    soup = BeautifulSoup(response.content, "lxml")  # or BeautifulSoup(response.content, "html5lib")
    try:
        return soup.find(text="Foreign:").find_parent("td").find_next_sibling("td").get_text(strip = True)
    except AttributeError:
        return "$0"

附带说明一下,由于脚本的阻塞性质,总体而言,您拥有的代码变得相当复杂和缓慢,请求是按顺序逐个发送的。切换到Scrapy web-scraping framework 可能是个好主意,这除了使代码更快之外,还有助于将其组织成逻辑组——你会有一个蜘蛛,里面有抓取逻辑,项目类定义你的提取数据模型,用于将提取的数据写入数据库的管道(如果需要)等等。

【讨论】:

  • 当我尝试打印 listOfForeign 时,它现在正在打印一个空列表。我将 listOfForeign.append(getForeign(href)) 保留在蜘蛛函数中,并将 getForeign 方法更改为您上面写的。
  • @alphamonkey 它不应该产生这样的效果。请发布您当前拥有的完整代码。
猜你喜欢
  • 2015-05-12
  • 1970-01-01
  • 2020-07-27
  • 1970-01-01
  • 1970-01-01
  • 2021-05-26
  • 1970-01-01
  • 2023-01-26
  • 2017-06-09
相关资源
最近更新 更多