【问题标题】:python scrapy extract data from websitepython scrapy从网站中提取数据
【发布时间】:2015-03-14 22:38:26
【问题描述】:

我想从this page 抓取数据。这是我当前的代码:

buf = cStringIO.StringIO()
c = pycurl.Curl()
c.setopt(c.URL, "http://www.guardalo.org/99407/")
c.setopt(c.VERBOSE, 0)
c.setopt(c.WRITEFUNCTION, buf.write)
c.setopt(c.CONNECTTIMEOUT, 15)
c.setopt(c.TIMEOUT, 15)
c.setopt(c.SSL_VERIFYPEER, 0)
c.setopt(c.SSL_VERIFYHOST, 0)
c.setopt(c.USERAGENT, 'Mozilla/5.0 (Windows NT 6.1; WOW64; rv:8.0) Gecko/20100101 Firefox/8.0')
c.perform()
body = buf.getvalue()
c.close()

response = HtmlResponse(url='http://www.guardalo.org/99407/', body=body)
print Selector(response=response).xpath('//edindex/text()').extract()

它有效,但我需要标题、视频链接和描述作为单独的变量。我怎样才能做到这一点?

【问题讨论】:

    标签: python web-scraping scrapy


    【解决方案1】:

    标题可以使用//title/text()提取,视频源链接通过//video/source/@src:

    selector = Selector(response=response)
    
    title = selector.xpath('//title/text()').extract()[0]
    description = selector.xpath('//edindex/text()').extract()
    video_sources = selector.xpath('//video/source/@src').extract()[0]
    
    code_url = selector.xpath('//meta[@name="EdImage"]/@content').extract()[0]
    code = re.search(r'(\w+)-play-small.jpg$', code_url).group(1)
    
    print title
    print description
    print video_sources
    print code
    

    打印:

    Best Babies Laughing Video Compilation 2012 [HD] - Guardalo
    [u'Best Babies Laughing Video Compilation 2012 [HD]', u"Ciao a tutti amici di guardalo,quello che propongo oggi \xe8 un video sui neonati buffi con risate travolgenti, facce molto buffe,iniziamo con una coppia di gemellini che se la ridono fra loro,per passare subito con una biondina che si squaqqera dalle risate al suono dello strappo della carta ed \xe8 solo l'inizio.", u'\r\nBuone risate a tutti', u'Elia ride', u'Funny Triplet Babies Laughing Compilation 2014 [NEW HD]', u'Real Talent Little girl Singing Listen by Beyonce .', u'Bimbo Napoletano alle Prese con il Distributore di Benzina', u'Telecamera nascosta al figlio guardate che fa,video bambini divertenti,video bambini divertentissimi']
    http://static.guardalo.org/video_image/pre-roll-guardalo.mp4
    L49VXZwfup8
    

    【讨论】:

    • 对于视频需要捕获此代码:L49VXZwfup8,是来自youtube的代码视频!
    • @pythoncoder 好的,更新了答案,这是您要问的吗?谢谢。
    • @pythoncoder 还注意到 Alex Martelli 在这里有一个有效的观点——如果你使用 Scrapy 从这个单一的 URL 中提取数据——那么这是一个巨大的开销。我假设您要将解决方案扩展到此类多个 URL。
    【解决方案2】:

    不需要scrapy 来获取单个 URL —— 只需使用更简单的工具(甚至最简单的urllib.urlopen(theurl).read()!)获取单个页面的 HTML,然后使用 BeautifulSoup 分析 HTML。从一个简单的“查看源代码”来看,您正在寻找:

    <title>Best Babies Laughing Video Compilation 2012 [HD] - Guardalo</title>
    

    (标题),三者之一:

    <source src="http://static.guardalo.org/video_image/pre-roll-guardalo.mp4" type='video/mp4'>
    <source src="http://static.guardalo.org/video_image/pre-roll-guardalo.webm" type='video/webm'>
    <source src="http://static.guardalo.org/video_image/pre-roll-guardalo.ogv" type='video/ogg'>
    

    (视频链接S,复数,我不能选择一个,因为你没有告诉我们你喜欢哪种格式!-),和

    <meta name="description" content="Ciao a tutti amici di guardalo,quello che propongo oggi è un video sui neonati buffi con risate" />
    

    (描述)。 BeautifulSoup 使得获取每一个变得非常简单,例如在需要的导入之后

    html = urllib.urlopen('http://www.guardalo.org/99407/').read()
    soup = BeautifulSoup(html)
    title = soup.find('title').text
    

    等等(但你必须选择一个视频链接——我在他们的来源中看到它们被称为“前贴片广告”,所以实际上可能是指向实际非广告视频的链接不在页面上,但只有在登录后才能访问或其他)。

    【讨论】:

    • 需要为视频捕获此代码:L49VXZwfup8 因为这是视频 youtube 的代码
    猜你喜欢
    • 1970-01-01
    • 2015-01-15
    • 1970-01-01
    • 2017-01-23
    • 2019-04-17
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-04-03
    相关资源
    最近更新 更多