【问题标题】:Select previous td and clicking link with Mechanize and Nokogiri选择上一个 td 并单击带有 Mechanize 和 Nokogiri 的链接
【发布时间】:2015-06-06 17:18:39
【问题描述】:

您好,我正在废弃一个带有 mechanize 和 nokogiri 的网页。我正在选择一系列链接<a></a>

 html_body = Nokogiri::HTML(body)
    links = html_body.css('.L1').xpath("//table/tbody/tr/td[2]/a[1]")

然后我需要检查每个链接的内容(<a>content</a>,而不是href)是否与我的数据库中的某些内容匹配。我正在这样做:

       links.each do |link|
          if link = @tournament.homologation_number

如果实现了我的条件,我需要选择我检查过的链接的<td> 之前的<td></td>,然后单击其中的链接。

<td><a href="link I want to click if condition is true"></a></td>
<td><a href="">content I check with my condition</a></td>

如何使用 Mechanize 和 nokogiri 实现这一目标?

【问题讨论】:

    标签: ruby web-scraping nokogiri mechanize


    【解决方案1】:

    我会迭代第一个 td,因为它比以前的元素更容易获得后续元素(无论如何使用 css)

    page.search('td[1]').each do |td|
      if td.at('+ td a').text == 'foo'
        page2 = agent.get td.at('a')[:href]
      end
    end
    

    【讨论】:

      【解决方案2】:

      首先你必须全选&lt;td&gt;&lt;/td&gt;,下面的xpath //table/tbody/tr/td[2]/a[1]只选择第一个&lt;a&gt;&lt;/a&gt;元素,所以你可以试试//table/tbody/tr/td之类的,但这要视情况而定。

      拥有&lt;td&gt;&lt;/td&gt; 数组后,您可以像这样访问他们的链接:

      tds.each do |td|
        link = td.children.first             # Select the first children
        if condition_is_matched(link.html)   # Only consider the html part of the link, if matched follow the previous link
          previous_td   = td.previous
          previous_url = previous_td.children.first.href
          goto_url previous_url
        end
      end
      

      【讨论】:

      • 好吧,我设法获得了第一个 td,但 td.previous 没有捕获任何东西
      猜你喜欢
      • 1970-01-01
      • 2012-07-23
      • 2014-03-27
      • 2011-11-21
      • 2014-03-15
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-03-19
      相关资源
      最近更新 更多