【问题标题】:How to compare every row of dataframe to dataframe in R?如何将每一行数据帧与 R 中的数据帧进行比较?
【发布时间】:2021-04-02 01:59:19
【问题描述】:

我想获得与数据框中所有其他行相等的值的数量:

library(tidyverse)

df <- tibble(
  a = c(1, 1, 5, 1),
  b = c(2, 3, 2, 8),
  c = c(2, 6, 2, 2)
)

想要的输出:

# A tibble: 4 x 4
      a     b     c desired_column
  <dbl> <dbl> <dbl> <list>        
1     1     2     2 <dbl [4]>     
2     1     3     6 <dbl [4]>     
3     5     2     2 <dbl [4]>     
4     1     8     2 <dbl [4]> 

在“desired_column”列中: 第一行:3、1、2、2:

3:是因为第一行的三个值与其自身相比是相同的

1:是因为在两行和同一列(第一和第二)中都有一个具有相同值的值:

2:第一行和第三行同一列有两个相等的值:

2:第一行和第四行同一列有两个相等的值:

“desired_column”的第二、三、四行是同一个过程的结果: 结果中的ith 数字是当前行和ith 行之间共有值的数量

【问题讨论】:

  • 我不明白为什么在结果中,第一行第一个数字是3(“与自身相比,三个值相同),但第四行(输入1, 8, 2)有一个@ 987654335@作为第一个结果编号。第4行没有重复值,为什么第一个结果编号是2?
  • 哦,我想我明白了。结果中的ith 数字是当前行和ith 行之间共有值的数量!
  • desired_column 的第四行第一个结果数为 2 是因为比较第 4 行 (1, 8, 2) 和第 1 行 (1, 2, 2) 时,只有两个值位于两行:1 和 2。

标签: r dplyr row tidyverse


【解决方案1】:

我的方法是将数据连接到自身,制作一个表格,将每个值与每个原始行中该列的值进行比较。然后我们计算比赛并再次扩大范围。

df %>%
  rowid_to_column() %>%
  pivot_longer(-rowid) -> df2

left_join(df2, df2, by = "name") %>%
  count(rowid.x, rowid.y, wt = value.x == value.y) %>%     # Edit - shorter
  pivot_wider(names_from = rowid.y, values_from = n) %>%
  nest(desired_column = c(`1`:`4`)) %>%
  select(-rowid.x) -> matches

bind_cols(df, matches)


# A tibble: 4 x 4
      a     b     c desired_column  
  <dbl> <dbl> <dbl> <list>          
1     1     2     2 <tibble [1 × 4]>
2     1     3     6 <tibble [1 × 4]>
3     5     2     2 <tibble [1 × 4]>
4     1     8     2 <tibble [1 × 4]>


> matches %>%
+   unnest(cols = c(desired_column))
# A tibble: 4 x 4
    `1`   `2`   `3`   `4`
  <int> <int> <int> <int>
1     3     1     2     2
2     1     3     0     1
3     2     0     3     1
4     2     1     1     3

【讨论】:

    【解决方案2】:

    您可以这样做:简而言之,对于数据框的每一行,复制它以创建一个新的数据框,并将所有值更改为该行,并将该数据框与原始数据框进行比较(值是否相同)。 rowSums 每一个比较都会给你你想要的向量。

    # Create the desired output in list 
    lst <- 
      lapply(1:nrow(df), function(nr) {
         rowSums(replicate(nrow(df), df[nr, ], simplify = FALSE) %>% 
                 do.call("rbind", .) == df)})
    
    # To create the desired dataframe
    df %>% tibble(desired_column = I(lst))
    

    在最后一行tibble调用中,I()用于将列表输出作为一列放入。

    【讨论】:

      【解决方案3】:

      另一种方法是使用几个 for 循环来创建一个函数:

      count_combs <- function(df){
      output <- list()
      vector <- NULL
      for(i in 1:nrow(df)){
        for(j in 1:nrow(df)){
        vector[j] <- sum(df[i,] %in% df[j,])
      }
      output[[i]] <- vector
      }
      return(output)
      }
      
      df$desired_column<- count_combs(df)
      

      这里 count_combs 函数计算每行的组合,每行被 i 迭代一次,再次被 j 迭代,每次行元素是比较行的 %in% 时求和。

      【讨论】:

        猜你喜欢
        • 2021-12-24
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2016-08-26
        • 2018-11-19
        • 1970-01-01
        • 2021-09-16
        • 1970-01-01
        相关资源
        最近更新 更多