【问题标题】:How to fetch all pages of specific website with Ruby on Rails如何使用 Ruby on Rails 获取特定网站的所有页面
【发布时间】:2015-07-28 10:16:11
【问题描述】:

我目前正在使用 Ruby on Rails(Ruby:2.2.1,Rails:4.2.1)构建网站,并希望从特定网站提取数据,然后将其显示出来。我使用 Nokogiri 来获取网页的内容。我正在寻找的是获取该网站的所有页面并获取其内容。

在我的代码下面:

doc = Nokogiri::HTML(open("www.google.com").read)
puts doc.at_css('title').text
puts doc.to_html

【问题讨论】:

  • 你需要的代码很复杂,你用它写了大约 1% 的东西。您基本上需要遍历页面上的所有链接,当您获取时,过滤掉外部链接,并存储一组已获取的页面,以避免重复调用。
  • 你应该搜索 Stack Overflow。沿着这条线有很多问题。以下是一些提示:stackoverflow.com/a/4981595/128421

标签: ruby-on-rails ruby nokogiri


【解决方案1】:

这是您需要的非常近似的要点:

class Parser
  attr_accessor :pages

  def fetch_all(host)
    @host = host

    fetch(@host)
  end

  private

  def fetch(url)
    return if pages.any? { |page| page.url == url }
    parse_page(Nokogiri::HTML(open(url).read))
  end

  def parse_page(document)
    links = extract_links(document)

    pages << Page.new(
      url: url,
      title: document.at_css('title').text,
      content: document.to_html,
      links: links
    )

    links.each { |link| fetch(@host + link) }
  end

  def extract_links(document)
    document.css('a').map do |link|
      href = link['href'].gsub(@host, '')
      href if href.start_with?('/')
    end.compact.uniq
  end
end

class Page
  attr_accessor :url, :title, :html_content, :links

  def initialize(url:, title:, html_content:, links:)
    @url = url
    @title = title
    @html_content = html_content
    @links = links
  end
end

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-07-26
    • 1970-01-01
    • 2010-09-15
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多