【问题标题】:How to return the row from a data frame based on the maximum value of the data frame in R?如何根据R中数据框的最大值从数据框中返回行?
【发布时间】:2015-04-17 00:11:12
【问题描述】:

假设我有一个如下所示的数据框。我在 Stackoverflow 上找到的大多数建议旨在从 one 列中获取最大值,然后返回行索引。 我想知道是否有办法通过扫描两个或更多列的最大值来返回数据帧的行索引。

总结一下,从下面的例子中,我想得到行:

11 building_footprint_sum 0.003 0.470

保存数据帧的最大值

+----+-------------------------+--------------------+-------------------+
| id |        plot_name        | rsquare_allotments | rsquare_block_dev |
+----+-------------------------+--------------------+-------------------+
|  6 | building_footprint_max  | 0.002              | 0.421             |
|  7 | building_footprint_mean | 0.002              | 0.354             |
|  8 | building_footprint_med  | 0.002              | 0.350             |
|  9 | building_footprint_min  | 0.002              | 0.278             |
| 10 | building_footprint_sd   | 0.003              | 0.052             |
| 11 | building_footprint_sum  | 0.003              | 0.470             |
+----+-------------------------+--------------------+-------------------+

有没有一种相当简单的方法来实现这一点?

【问题讨论】:

  • 是否确定每一列的最大值都在同一行?如果不是,则其中一列必须优先,否则您将需要一些决策规则,例如使用两列的总和。

标签: r dataframe max


【解决方案1】:

您正在寻找矩阵达到最大值的行索引。您可以通过使用 which() 和 arr.ind=TRUE 选项来做到这一点:

> set.seed(1)
> foo <- matrix(rnorm(6),3,2)
> which(foo==max(foo),arr.ind=TRUE)
     row col
[1,]   1   2

所以在这种情况下,您需要第 1 行。(并且您可以丢弃 col 输出。)

如果您走这条路,请注意浮点运算和==(请参阅常见问题解答 7.31)。最好这样做:

> which(foo>max(foo)-0.01,arr.ind=TRUE)
     row col
[1,]   1   2

您使用适当的小值代替 0.01。

【讨论】:

  • 感谢您的建议,但在处理不仅包含数值的数据帧时似乎存在问题。当我尝试将您的代码用于我上面提到的数据框时,它说:“仅为数据框定义,只有数字变量”
  • (1) 您能否编辑您的帖子以包含一个可重现的示例,并指出如果数字和非数字数据混合在一起,您希望看到max 操作的结果? (并在此处发表评论,以便我注意到您的编辑)(2)我怀疑您需要对 data.frame 的相关列进行子集化,例如 as.matrix(your.data.frame[,3:4]) - 也许您可以尝试一下,看看它是否有帮助。
【解决方案2】:

尝试使用pmax

?pmax    
pmax and pmin take one or more vectors (or matrices) as arguments and
return a single vector giving the ‘parallel’ maxima (or minima) of the vectors.

我建议分两步完成

# make a new column that compares column 3 and column 4 and returns the larger value
> df$new <- pmax(df$rsquare_allotments, df$rsquare_block_dev)

# look for the row, where the new variable has the largest value
> df[(df$new == max(df$new)), ][3:4]

考虑如果最大值出现不止一次,您的结果将有不止一行

【讨论】:

  • 显然我没有足够准确地指出我的问题。我想扫描上面示例数据框的最后两列以获取最大值,而不是最大和。当找到数据框的最大值时,应返回包含最大值的行。
  • 好的。但是具有多列的 data.frame 的最大值是多少?
  • 和上面的例子一样,两列的最大值是0.470(第6行,第4列)。所以必须找到这个值并返回出现最大值的行,以便我可以将返回的行存储在新的 df 中。
  • 感谢@muc8 的启发。但是命令:df$new &lt;- pmax(df[3:4]) 以某种方式返回了一个向量,该向量写入新列的前两行。然后df$new &lt;- pmax(df$rsquare_allotments, df$rsquare_block_dev) 完成了这项工作。非常感谢
  • @Axel.Foley 只是为了澄清不同的方法:我的解决方案提供了行,正如您在问题的所需输出中所示 (I want to get the row)。但是,您的标题暗示您正在寻找行索引。因此,为了让一切清楚,请编辑您的问题(不再使用索引一词),或者如果这是您要找的,请接受 Stephan 的答案。
猜你喜欢
  • 1970-01-01
  • 2023-02-07
  • 2021-11-04
  • 1970-01-01
  • 1970-01-01
  • 2021-09-09
  • 2020-11-13
  • 2019-10-02
  • 2021-11-22
相关资源
最近更新 更多