【发布时间】:2015-04-20 09:42:05
【问题描述】:
这里是 HTML:
<tr class="level2">
<td>
<b>word</b>
"Text I need"
<b>word</b>
"Text I need"
<b>word</b>
"Text I need"
<b>word</b>
"Text I need"
<i>blabla</>
"Text I need"
<b>word</b>
"Text I need"
<i>blabla</>
"Text I need"
<i>blabla</>
<b>word</b>
</td>
</tr>
我想选择<b> 元素之间的每个节点,然后在以后遍历它们中的每一个。目前我有:
translations = page.xpath('//text()[preceding-sibling::b]')
当<b> 元素之间只有文本时,它可以正常工作。但是,当<b> 元素之间出现一个或多个<i> 标签时,我只得到节点中的第一个文本。节点中的剩余文本转到下一个节点。
我想要输出:
node 1: Text I need
node 2: Text I need
node 3: Text I need
node 4: Text I need
Text I need
node 5: Text I need
Text I need
这是代码:
require 'rubygems'
require 'open-uri'
require 'nokogiri' #parse html
require 'csv'
DATA_DIR = "words"
Dir.mkdir(DATA_DIR) unless File.exists?(DATA_DIR) # making directory
BASE_LINK = "http://dict.ibs.ee/translate.cgi?word="
LANGUAGE = "&language=English"
WILDCARD = "*"
SLEEP_TIME = 0.1 # sleep between web requests in seconds
counter = 1 #counter for file name
i = 1
name = "IBSwords"+"#{counter}"+".csv"
alphabet = %w[a b c d e f g h i j k l m n o p q r s t u v w x y z]
four_letter_combinations = alphabet.product(alphabet, alphabet, alphabet).map(&:join)
#combination from 4 letters
for combination in four_letter_combinations
begin
i += 1
if (i % 150000 ) == 0
counter += 1
name = "IBSwords"+"#{counter}"+".csv"
end
sleep (SLEEP_TIME)
link = BASE_LINK+"about"+LANGUAGE
page = Nokogiri::HTML(open(link)) #retry in 60 sec if no connection
rescue StandardError=>e
puts "#{e} No Connection, retrying..."
sleep 60
retry
else
unless page.css('body > div > center > table > tbody > tr > td > div > center > table > tbody > tr > td > blockquote > dl > dd > b').nil?
puts "*****************#{i} #{combination}***********"
en_words = page.css('blockquote > dl > dd > b')
#ee_words = page.css('blockquote > dl > dd').to_s.split(/<b>.*<\/b>/)
ee_words = page.xpath('//text()[preceding-sibling::b]')
# iterating through
en_words.zip(ee_words).each do |word, ee_word|
en_word = word.text.chomp.strip
ee_trans = ee_word.text.chomp.strip
#en_desc = word.xpath('td[2]/node()[not(self::strong)]').text
puts "#{en_word}"
puts "#{ee_trans}"
puts "*******************************"
i += 1
#writing to csv
CSV.open("words/#{name}", "ab") do |row| # write to CSV
row << [
en_word,
#en_desc,
ee_trans,
#ee_desc
]
end
end
end
end
end
【问题讨论】:
-
如果您只查找
td节点中的文本,为什么不使用page.css('td').text? -
与“
b标签之间的文本”相比,您想要的更好的描述是“b元素之间的文本”。 “标签”表示<b>和</b>– 开始和结束标签。 “标签之间”是指b元素的内容,这不是你想要的。 -
你看,我也在使用
words = page.css('tr td b')在每个 b 标记中都有一个单词,并且该 b 标记之后的所有以下文本都是翻译。稍后我将每个翻译映射到单词:words.zip(translations).each do |x, y| -
在提出问题时,将代码剥离到能证明问题的最低限度。除此之外的任何事情都会浪费我们的时间来帮助您解决问题。