【问题标题】:Grab everything between b elements with Nokogiri使用 Nokogiri 抓取 b 元素之间的所有内容
【发布时间】:2015-04-20 09:42:05
【问题描述】:

这里是 HTML:

<tr class="level2">
    <td> 
        <b>word</b>
        "Text I need"
        <b>word</b>
        "Text I need"
        <b>word</b>
        "Text I need"
        <b>word</b>
        "Text I need"
        <i>blabla</>
        "Text I need"
        <b>word</b>
        "Text I need"
        <i>blabla</>
        "Text I need"
        <i>blabla</>
        <b>word</b>

    </td>
</tr>

我想选择&lt;b&gt; 元素之间的每个节点,然后在以后遍历它们中的每一个。目前我有:

translations = page.xpath('//text()[preceding-sibling::b]')

当&lt;b&gt; 元素之间只有文本时,它可以正常工作。但是,当&lt;b&gt; 元素之间出现一个或多个&lt;i&gt; 标签时,我只得到节点中的第一个文本。节点中的剩余文本转到下一个节点。 我想要输出:

node 1: Text I need 
node 2: Text I need 
node 3: Text I need 
node 4: Text I need 
        Text I need 
node 5: Text I need 
        Text I need 

这是代码:

require 'rubygems'
require 'open-uri'
require 'nokogiri' #parse html
require 'csv'

DATA_DIR = "words"
Dir.mkdir(DATA_DIR) unless File.exists?(DATA_DIR) # making directory
BASE_LINK = "http://dict.ibs.ee/translate.cgi?word=" 
LANGUAGE = "&language=English"
WILDCARD = "*"
SLEEP_TIME = 0.1 # sleep between web requests in seconds
counter = 1 #counter for file name
i = 1
name = "IBSwords"+"#{counter}"+".csv"

alphabet = %w[a b c d e f g h i j k l m n o p q r s t u v w x y z]
four_letter_combinations = alphabet.product(alphabet, alphabet, alphabet).map(&:join)
#combination from 4 letters
for combination in four_letter_combinations
  begin
    i += 1
      if (i % 150000 ) == 0
        counter += 1
        name = "IBSwords"+"#{counter}"+".csv" 
      end
    sleep (SLEEP_TIME) 
    link = BASE_LINK+"about"+LANGUAGE
    page = Nokogiri::HTML(open(link)) #retry in 60 sec if no connection
  rescue StandardError=>e
    puts "#{e} No Connection, retrying..."
    sleep 60
  retry
  else 
    unless page.css('body > div > center > table > tbody > tr > td > div > center > table > tbody > tr > td > blockquote > dl > dd > b').nil?
      puts "*****************#{i} #{combination}***********"
      en_words = page.css('blockquote > dl > dd > b')
      #ee_words = page.css('blockquote > dl > dd').to_s.split(/<b>.*<\/b>/)
      ee_words = page.xpath('//text()[preceding-sibling::b]') 
      # iterating through 
      en_words.zip(ee_words).each  do |word, ee_word|
      en_word = word.text.chomp.strip
      ee_trans = ee_word.text.chomp.strip
      #en_desc = word.xpath('td[2]/node()[not(self::strong)]').text
      puts "#{en_word}"
      puts "#{ee_trans}"
      puts "*******************************"
      i += 1
      #writing to csv 
      CSV.open("words/#{name}", "ab") do |row| # write to CSV
          row << [
          en_word,
          #en_desc,
          ee_trans,
          #ee_desc
        ]
      end
    end
  end
end
end

【问题讨论】:

  • 如果您只查找td 节点中的文本,为什么不使用page.css('td').text?
  • 与“b 标签之间的文本”相比,您想要的更好的描述是“b 元素之间的文本”。 “标签”表示&lt;b&gt; 和&lt;/b&gt; – 开始和结束标签。 “标签之间”是指b元素的内容,这不是你想要的。
  • 你看,我也在使用words = page.css('tr td b') 在每个 b 标记中都有一个单词,并且该 b 标记之后的所有以下文本都是翻译。稍后我将每个翻译映射到单词:words.zip(translations).each do |x, y|
  • 在提出问题时,将代码剥离到能证明问题的最低限度。除此之外的任何事情都会浪费我们的时间来帮助您解决问题。

标签: ruby nokogiri


【解决方案1】:

我将您的 HTML 缩减为不那么冗长。它无需额外的文本即可实现相同的效果。

我会这样做:

require 'nokogiri'

doc = Nokogiri::HTML(<<EOT)
<tr class="level2">
    <td> 
        <b>word</b>
        "Text I need"
        <b>word</b>
        "Text I need"
        <i>blabla</i>
        "Text I need"
        <b>word</b>
        "Text I need"
        <i>blabla</i>
        "Text I need"
        <i>blabla</i>
        <b>word</b>
    </td>
</tr>
EOT

doc.search('td i').remove

由于不需要 &lt;i&gt; 节点,因此只需剥离它们即可。生成的 doc 看起来像:

puts doc.to_html
# >> <!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN" "http://www.w3.org/TR/REC-html40/loose.dtd">
# >> <html><body>
# >> <tr class="level2">
# >>     <td> 
# >>         <b>word</b>
# >>         "Text I need"
# >>         <b>word</b>
# >>         "Text I need"
# >>         
# >>         "Text I need"
# >>         <b>word</b>
# >>         "Text I need"
# >>         
# >>         "Text I need"
# >>         
# >>         <b>word</b>
# >> 
# >>     </td>
# >> </tr>
# >> </body></html>

一旦&lt;i&gt; 节点消失,就可以遍历&lt;td&gt; 的内容并处理它们的文本:

text = doc.at('td').children.reject { |n| n.text.strip == '' }.slice_before { |n| n.name == 'b' }.map{ |a| a.map { |n| n.text.strip }}

此时text 包含:

text
# => [["word", "\"Text I need\""],
#     ["word", "\"Text I need\"", "\"Text I need\""],
#     ["word", "\"Text I need\"", "\"Text I need\""],
#     ["word"]]

请注意,后面有一个“单词”,它模仿了您提供的示例 HTML。如果您知道您不会保留任何尾随文本,您可以简单地 pop 关闭该元素。如果您认为有些元素只是单个项目,您可以遍历列表以查找单项并拒绝它们。如何处理由您自己决定。

【讨论】:

    【解决方案2】:

    您可能正在寻找xpath-only 解决方案,但这里是使用 ruby​​ 枚举器的解决方案:

    xml.xpath('//td').children.inject({}) do |memo, node|
      case node.name
      when 'b' then memo["#{node.children.first}"] = ""
      when 'text' 
        memo["#{memo.keys.last}"] << "#{node}" unless memo.length.zero?
      else # just skip
      end 
    
      memo
    end
    

    这给了:

    #⇒ {
    #  "word 1" => "\n        \"Text I need 1\"\n        ",
    #  "word 2" => "\n        \"Text I need 2\"\n        ",
    #  "word 3" => "\n        \"Text I need 3\"\n        ",
    #  "word 4" => "\n        \"Text I need 41\"\n        \n        \"Text I need 42\"\n        ",
    #  "word 5" => "\n        \"Text I need 51\"\n        \n        \"Text I need 52\"\n        \n        ",
    #  "word 6" => "\n\n    "
    # }
    

    希望对您有所帮助。

    【讨论】:

    • 感谢您的回答,但我真的不知道如何正确实现您的代码。如您所见,我正在尝试打印 EN 单词+ EE 翻译对(添加完整代码)
    • 我不明白问题出在哪里。上面的代码生成哈希,您可以像以前一样简单地使用 .each do |word, ee_word| 对其进行迭代。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多