您的代码离目标不远。让我们来看看吧。
测试数据
以下是一些用于测试的数据:
oceans = ["Pacific; 1", "Atlantic; 2", "Indian; 3",
"Pacific; 2", "Atlantic; 1", "Pacific; 3"]
我没有读取文件的行,而是通过读取字符串数组来简化事情。一旦代码运行起来,就很容易将其更改为从文件中读取。
现在我们有了一些输入数据,我们可以显示我们想要的预期结果
hash =
{ 'Pacific' => ['1', '2', '3'],
'Atlantic' => ['2', '1'],
'Indian' => ['1'] }
或
hash =
{ 'Pacific' => "'1', '2', '3'",
'Atlantic' => "'2', '1'",
'Indian' => "'1'" }
我们将使用第一种,因为它最容易处理,如果我们想要第二种形式,我们可以很容易地从第一种计算出来:
hash.keys.each { |k| hash[k] = hash[k].join(',') }
#=> ["Pacific", "Atlantic", "Indian"]
但是,等等,这不是返回的哈希值。不,是hash.keys。我们要的是hash的新值:
hash #=> {"Pacific"=>"1,2,3", "Atlantic"=>"2,1", "Indian"=>"1"}
另外:向 SO 发布问题时,将一些说明性输入数据与预期结果一起包含通常会很有帮助。这往往会澄清并节省文字。尝试使用尽可能少的数据。
您的代码
这是您的代码,用数组 oceans 替换文件的读取:
descriptor_code_hash = Hash.new
oceans.each do |file_line|
file_line = file_line.chomp
mesh_descriptor, tree_code = file_line.split(/\;/)
descriptor_code_hash[mesh_descriptor] = tree_code
if descriptor_code_hash.has_key? mesh_descriptor
tree_code << "," << tree_code
else
descriptor_code_hash[mesh_descriptor]
end
end
主要问题是行:
descriptor_code_hash[mesh_descriptor] = tree_code
每次循环时,键 mesh_descriptor 的 descriptor_code_hash 的值被重置为 oceans 的当前元素的 tree_code 的值(代表文件的一行)。您需要删除此行。
接下来,我们需要修改你的if/else/end声明,如下:
if descriptor_code_hash.has_key? mesh_descriptor
descriptor_code_hash[mesh_descriptor] << tree_code
else
descriptor_code_hash[mesh_descriptor] = [tree_code]
end
这将为您提供以下信息:
descriptor_code_hash = Hash.new
oceans.each do |file_line|
file_line = file_line.chomp
mesh_descriptor, tree_code = file_line.split(/\;/)
if descriptor_code_hash.has_key? mesh_descriptor
descriptor_code_hash[mesh_descriptor] << tree_code
else
descriptor_code_hash[mesh_descriptor] = [tree_code]
end
end
当我们运行它时,我们得到:
descriptor_code_hash
#=> {"Pacific"=>[" 1", " 2", " 3"], "Atlantic"=>[" 2", " 1"],
# "Indian"=>[" 3"]}
如您所见,结果是正确的,除了有一个小的格式问题。我们可以通过改变来解决这个问题:
file_line.split(/\;/)
到
file_line.split(/\;/).map { |w| w.strip }
可以通过两种方式简化:
file_line.split(';').map(&:strip)
让我们试试吧。假设:
file_line = "Pacific; 1\n"
然后
file_line.split(';').map(&:strip) #=> ["Pacific", "1"]
这是期望的结果。请注意,我在字符串的末尾包含了一个换行符。那是为了向您展示 strip 删除它以及空格。这意味着您不需要上一行:
file_line = file_line.chomp
(file_line.chomp.split(/\s*;\s*/) 也可以。)
您的代码现在简化为:
descriptor_code_hash = Hash.new
oceans.each do |file_line|
mesh_descriptor, tree_code = file_line.split(';').map(&:strip)
if descriptor_code_hash.has_key? mesh_descriptor
descriptor_code_hash[mesh_descriptor] << tree_code
else
descriptor_code_hash[mesh_descriptor] = [tree_code]
end
end
抛光
现在考虑如何使代码更像 Ruby。首先,查看@BroiSatse 给出的答案中使用的以下行(代替您的if/else/end 构造):
(descriptor_code_hash[mesh_descriptor] ||= []) << tree_code
对于任何变量a,a ||= [] 与a = (a || []) 相同。如果没有定义a,它将等于nil,所以(nil || []) => []。如果a 被分配了一个(非零)值,(a || []) => a。也就是说,如果descriptor_code_hash没有键mesh_descriptor(意思是descriptor_code_hash[mesh_descriptor] => nil),则descriptor_code_hash[mesh_descriptor]会被赋值为[];否则,它将被自己分配(即,它不会改变)。
之后
descriptor_code_hash[mesh_descriptor] ||= []
被执行,descriptor_code_hash[mesh_descriptor] 将等于一个数组,空或其他。 << tree_code 然后将 tree_code 附加到哈希值(一个数组)。最后,我们可以使用{} 而不是Hash.new,但这纯粹是一种风格选择。
您的代码现在如下所示:
descriptor_code_hash = {}
oceans.each do |file_line|
mesh_descriptor, tree_code = file_line.split(';').map(&:strip)
(descriptor_code_hash[mesh_descriptor] ||= []) << tree_code
end
现在让我们把它变成一个方法并做更多的改变:
def descriptor_code_hash(oceans)
oceans.each_with_object({}) do |line, hash|
mesh_descriptor, tree_code = line.split(';').map(&:strip)
(hash[mesh_descriptor] ||= []) << tree_code
end
end
descriptor_code_hash(oceans)
#=> {"Pacific"=>["1", "2", "3"], "Atlantic"=>["2", "1"], "Indian"=>["3"]}
我已经简化了一些变量名称,因为方法的目的是通过它的名称来描述的。通读 Enumerable#each_with_object 的文档(从 1.9 版开始提供),了解它是如何使用的。
您可能希望文件名作为方法参数。
最后一件事:您可以改写如下:
def descriptor_code_hash(oceans)
oceans.each_with_object(Hash.new {|k,h| h[k] = {} }) do |line, hash|
mesh_descriptor, tree_code = line.split(';').map(&:strip)
hash[mesh_descriptor] << tree_code
end
end
这里对象被初始化为:
Hash.new {|k,h| h[k] = {} }
这使得默认值(当将新键添加到哈希时)为空哈希。这就是倒数第三行可以简化为所示的原因。