【问题标题】:What is the mistake here?这里有什么错误?
【发布时间】:2016-10-12 15:31:07
【问题描述】:

这些是我的代码,但它似乎是正确的,但它不起作用,请帮助

HEADER_XPATH = ['//h1[@class="story-body__h1"]//text()']    
AUTHOR_XPATH = ['//span[@class="byline__name"]//text()']   
PUBDATE_XPATH = ['//div/@data-datetime']  
WTAGS_XPATH = ['']   
CATEGORY_XPATH = ['//span[@rev="news|source""]//text()']    
TEXT = ['//div[@property="articleBody"]//p//text()']   
INTERLINKS = ['//div[@class="story-body__link"]//p//a/@href']  
DATE_FORMAT_STRING = '%Y-%m-%d'

class BBCSpider(Spider):
    name = "bbc"
    allowed_domains = ["bbc.com"]
    sitemap_urls = [
        'http://Www.bbc.com/news/sitemap/',
        'http://www.bbc.com/news/technology/',
        'http://www.bbc.com/news/science_and_environment/']

    def parse_page(self, response):
        items = []
        item = ContentItems()
        item['title'] = process_singular_item(self, response, HEADER_XPATH, single=True)
        item['resource'] = urlparse(response.url).hostname
        item['author'] = process_array_item(self, response, AUTHOR_XPATH, single=False)
        item['pubdate'] = process_date_item(self, response, PUBDATE_XPATH, DATE_FORMAT_STRING, single=True)
        item['tags'] = process_array_item(self, response, TAGS_XPATH, single=False)
        item['category'] = process_array_item(self, response, CATEGORY_XPATH, single=False)
        item['article_text'] = process_article_text(self, response, TEXT)
        item['external_links'] = process_external_links(self, response, INTERLINKS, single=False)
        item['link'] = response.url
        items.append(item)
        return items

【问题讨论】:

  • 有什么问题?也许解释一下问题是什么?输入?渴望输出?你在做什么?
  • 问题是当我运行我的代码时没有发生任何事情。它不通过页面!我认为我的错误在于变量@MooingRawr

标签: python python-3.x scrapy web-crawler scrapy-spider


【解决方案1】:

你的蜘蛛只是结构很糟糕,因此它什么也没做。
scrapy.Spider 蜘蛛需要 start_urls 类属性,该属性应包含蜘蛛将用于开始抓取的 url 列表,所有这些 url 将回调到类方法 parse 这意味着它也是必需的。

你的蜘蛛有 sitemap_urls 类属性,它没有在任何地方使用,而且你的蜘蛛有 parse_page 类方法,也从未在任何地方使用过。
所以简而言之,你的蜘蛛应该看起来像这样:

class BBCSpider(Spider):
    name = "bbc"
    allowed_domains = ["bbc.com"]
    start_urls = [
        'http://Www.bbc.com/news/sitemap/',
        'http://www.bbc.com/news/technology/',
        'http://www.bbc.com/news/science_and_environment/']

    def parse(self, response):
        # This is a page with all of the articles
        article_urls = # find article urls in the pages
        for url in article_urls:
            yield Request(url, self.parse_page)

    def parse_page(self, response):
        # This is an article page
        items = []
        item = ContentItems()
        # populate item
        return item

【讨论】:

  • 非常感谢
  • @nik 太棒了!如果它解决了您的问题,请随时点击它左侧的“接受问题按钮”。
  • ,当然,我会的。你能在articles_urls中告诉我我必须提供什么作为例子吗?(因为我要处理很多url)
  • 你需要找出一个 xpath,像这样:response.xpath("//a[contains(@class,'title-link')]/@href").extract() 找到一些。
猜你喜欢
  • 2011-08-09
  • 1970-01-01
  • 2022-01-24
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2022-08-15
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多