【问题标题】:Pipe output of one data.frame to another using dplyr使用 dplyr 将一个 data.frame 的输出通过管道传输到另一个
【发布时间】:2016-11-05 05:46:02
【问题描述】:

我有两个 data.frames--一个查找表,它告诉我一组包含在一个组中的产品。每个组有至少一种类型 1 和类型 2 的产品。

第二个 data.frame 告诉我有关交易的详细信息。每笔交易可以有以下产品之一:

a) 仅来自其中一个组的类型 1 的产品s

b) 仅来自其中一个组的类型 2 的产品s

c) 类型 1 和类型 2 的产品来自同一组

对于我的分析,我有兴趣找出上面的 c),即有多少交易有类型 1和类型 2(来自同一组)的产品已售出。如果类型 1 的产品和类型 2 的产品来自不同组的产品在同一交易中出售,我们将完全忽略该交易。

因此,类型 1 或类型 2 的每个产品必须属于同一组。

这是我的查找表:

> P_Lookup
   Group ProductID1 ProductID2
  Group1          A          1
  Group1          B          2
  Group1          B          3
  Group2          C          4
  Group2          C          5
  Group2          C          6
  Group3          D          7
  Group3          C          8
  Group3          C          9
  Group4          E         10
  Group4          F         11
  Group4          G         12
  Group5          H         13
  Group5          H         14
  Group5          H         15 

例如,我不会将产品 G 和产品 15 放在一个交易中,因为它们属于不同的组。

这里是交易:

  TransactionID ProductID ProductType
             a1         A           1
             a1         B           1
             a1         1           2
             a2         C           1
             a2         4           2
             a2         5           2
             a3         D           1
             a3         C           1
             a3         7           2
             a3         8           2
             a4         H           1
             a5         1           2
             a5         2           2
             a5         3           2
             a5         3           2
             a5         1           2
             a6         H           1
             a6        15           2

我的代码:

现在,我可以使用dplyr 编写代码来筛选一组交易。但是,我不确定如何为 all 组矢量化我的代码。

这是我的代码:

P_Groups<-unique(P_Lookup$Group)
Chosen_Group<-P_Groups[5]

P_Group_Ind <- P_Trans %>%
group_by(TransactionID)%>%
dplyr::filter((ProductID %in% unique(P_Lookup[P_Lookup$Group==Chosen_Group,]$ProductID1)) | 
(ProductID %in% unique(P_Lookup[P_Lookup$Group==Chosen_Group,]$ProductID2)) ) %>%
mutate(No_of_PIDs = n_distinct(ProductType)) %>%
mutate(Group_Name = Chosen_Group)

P_Group_Ind<-P_Group_Ind[P_Group_Ind$No_of_PIDs>1,]

只要我手动选择每个组,即通过设置Chosen_Group,它就可以很好地工作。但是,我不确定如何实现自动化。我在想的一种方法是使用 for 循环,但我知道 R 的美妙之处在于矢量化,所以我想远离使用 for 循环。

我真诚地感谢任何帮助。我花了将近两天的时间。我查看了using dplyr in for loop in r,但似乎这个帖子在谈论一个不同的问题。


数据: 这是dput 代表P_Trans:

structure(list(TransactionID = c("a1", "a1", "a1", "a2", "a2", 
"a2", "a3", "a3", "a3", "a3", "a4", "a5", "a5", "a5", "a5", "a5", 
"a6", "a6"), ProductID = c("A", "B", "1", "C", "4", "5", "D", 
"C", "7", "8", "H", "1", "2", "3", "3", "1", "H", "15"), ProductType = c(1, 
1, 2, 1, 2, 2, 1, 1, 2, 2, 1, 2, 2, 2, 2, 2, 1, 2)), .Names = c("TransactionID", 
"ProductID", "ProductType"), row.names = c(NA, 18L), class = "data.frame")

这是dput P_Lookup:

structure(list(Group = c("Group1", "Group1", "Group1", "Group2", 
"Group2", "Group2", "Group3", "Group3", "Group3", "Group4", "Group4", 
"Group4", "Group5", "Group5", "Group5"), ProductID1 = c("A", 
"B", "B", "C", "C", "C", "D", "C", "C", "E", "F", "G", "H", "H", 
"H"), ProductID2 = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 
14, 15)), .Names = c("Group", "ProductID1", "ProductID2"), row.names = c(NA, 
15L), class = "data.frame")

这是将产品添加到 P_Trans 后查找表中不存在的 dput():

structure(list(TransactionID = c("a1", "a1", "a1", "a2", "a2", 
"a2", "a3", "a3", "a3", "a3", "a4", "a5", "a5", "a5", "a5", "a5", 
"a6", "a6", "a7"), ProductID = c("A", "B", "1", "C", "4", "5", 
"D", "C", "7", "8", "H", "1", "2", "3", "3", "1", "H", "15", 
"22"), ProductType = c(1, 1, 2, 1, 2, 2, 1, 1, 2, 2, 1, 2, 2, 
2, 2, 2, 1, 2, 3)), .Names = c("TransactionID", "ProductID", 
"ProductType"), row.names = c(NA, 19L), class = "data.frame")

【问题讨论】:

    标签: r nested dplyr


    【解决方案1】:

    以下是一个 tidyverse(dplyr、tidyr 和 purrr)解决方案,希望对您有所帮助。

    请注意,在最后一行中使用map_df 会将所有结果作为数据框返回。如果您希望它成为每个组的列表对象,则只需使用map。

    library(dplyr)
    library(tidyr)
    library(purrr)
    
    # Save unique groups for later use
    P_Groups <- unique(P_Lookup$Group)
    
    # Convert lookup table to product IDs and Groups
    P_Lookup <- P_Lookup %>% 
                  gather(ProductIDn, ProductID, ProductID1, ProductID2) %>% 
                  select(ProductID, Group) %>% 
                  distinct() %>% 
                  nest(-ProductID, .key = Group)
    
    # Bind Group information to transactions
    # and group for next analysis
    P_Trans <- P_Trans %>%
                 left_join(P_Lookup) %>%
                 filter(!map_lgl(Group, is.null)) %>%  
                 unnest(Group) %>% 
                 group_by(TransactionID)
    
    # Iterate through Groups to produce results
    map(P_Groups, ~ filter(P_Trans, Group == .)) %>% 
      map(~ mutate(., No_of_PIDs = n_distinct(ProductType))) %>% 
      map_df(~ filter(., No_of_PIDs > 1))
    #> Source: local data frame [12 x 5]
    #> Groups: TransactionID [4]
    #> 
    #>    TransactionID ProductID ProductType  Group No_of_PIDs
    #>            <chr>     <chr>       <dbl>  <chr>      <int>
    #> 1             a1         A           1 Group1          2
    #> 2             a1         B           1 Group1          2
    #> 3             a1         1           2 Group1          2
    #> 4             a2         C           1 Group2          2
    #> 5             a2         4           2 Group2          2
    #> 6             a2         5           2 Group2          2
    #> 7             a3         D           1 Group3          2
    #> 8             a3         C           1 Group3          2
    #> 9             a3         7           2 Group3          2
    #> 10            a3         8           2 Group3          2
    #> 11            a6         H           1 Group5          2
    #> 12            a6        15           2 Group5          2
    

    【讨论】:

    • 这是一个很棒的回应。我是初学者,所以我有一个简单的问题:你能解释一下你为什么这样做mutate(i = 1:n()) 和as_data_frame()。我尝试在没有这两行的情况下运行代码,它运行良好。所以,我很好奇。
    • 好消息@watchtower!这些台词只是为了帮助我在进行过程中弄清楚事情。我已将它们从答案中删除。
    • 感谢您的澄清。我有一个快速的问题——我的原始数据集P_Trans 有一些ProductID 不在查找表P_Lookup 中。在这种情况下,我收到此错误Error: Each column must either be a list of vectors or a list of data frames [Group]。我是初学者,所以我不知道如何解决这个问题。你认为你能帮助我吗?根据 SO 的政策,我不想创建一个新线程。我在帖子中添加了dput()。非常感谢您的帮助
    • @watchtower 感谢您的接受。我进行了一项更改,将删除查找表中未包含的所有产品:filter(!map_lgl(Group, is.null))。这发生在与P_Trans 的绑定中。
    【解决方案2】:

    这里是单管dplyr解决方案:

    P_DualGroupTransactionsCount <- 
        P_Lookup %>% # data needing single column map of Keys
        gather(IDnum, ProductID, ProductID1:ProductID2) %>% # produce long single map of Keys for GroupID (tidyr::)
        right_join(P_trans) %>% # join transactions to groupID info
        group_by(TransactionID, Group) %>% # organize for same transaction & same group
        mutate(DualGroup = ifelse(n_distinct(ProductType)==2, T, F)) %>% # flag groups with both groups in a single transaction
        filter(DualGroup == T) %>% # choose only doubles
        select(TransactionID, Group) %>% # remove excess columns
        distinct %>%  # remove excess rows
        nrow # count of unique transaction ID's
    
    # P_DualGroupTransactions
    # Source: local data frame [4 x 2]
    # Groups: TransactionID, Group [4]
    #     
    # TransactionID  Group
    #           <chr>  <chr>
    # 1            a1 Group1
    # 2            a2 Group2
    # 3            a3 Group3
    # 4            a6 Group5
    
    
    # P_DualGroupTransactionsCount
     [1] 4
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2018-11-17
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2019-03-09
      • 2012-01-11
      • 2013-10-13
      相关资源
      最近更新 更多