【问题标题】:R function(): how to pass parameters which contain characters and regular expressionR函数():如何传递包含字符和正则表达式的参数
【发布时间】:2014-01-09 09:15:54
【问题描述】:

我的数据如下:

>df2

  id   calmonth        product
1 101       01           apple
2 102       01 apple&nokia&htc
3 103       01             htc
4 104       01       apple&htc
5 104       02           nokia

现在我想计算当calmonth='01' 时products 包含both 'apple' and 'htc' 的ids 的数量。因为我需要的不仅是'apple'和'htc',我还需要'apple'和'nokia'等。 所以我想通过这样的功能来实现这一点:

xandy=function(a,b) data.frame(product=paste(a,b,sep='&'),
                               csum=length(grep('a.*b',x=df2$product))
                              )

另外,我制作了一个这样的参数列表:

para=c('apple','htc','nokia')

但问题就在这里。当我传递像

这样的参数时
xandy(para[1],para[2])

结果如下:

  product    csum
1 apple&htc    0

我的预期结果应该是什么

  product    csum   calmonth
1 apple&htc    2     01
2 apple&htc    0     02

那么参数传递的问题在哪里呢? 还有,我怎样才能正确地将calmonth 添加到函数()xandy 中? 仅供参考。这个问题源于我之前的另一个问题 What's the R statement responding to SQL's 'in' statement


评论后编辑

我的预测结果是:

product    csum   calmonth
 1 apple&htc    2     01
 2 apple&htc    0     02

【问题讨论】:

    标签: regex r


    【解决方案1】:

    可能的回答是解决问题的另一种方式。

    library(stringr)
    

    函数contains将根据split字符拆分字符串向量的元素,并评估是否包含所有目标词。

    contains <- function(x, target, split="&") {
      l <- str_split(x, split)
      sapply(l, function(x, y) all(y %in% x), y=target)  
    }
    
    contains(d$product, c("apple", "htc")) 
    [1] FALSE  TRUE FALSE  TRUE FALSE
    

    剩下的只是子集和总结

    get_data <- function(a, b) {
      e <- subset(d, contains(product, c(a, b)))
      e$product2 <- paste(a, b, sep="&")
      ddply(e, .(calmonth, product2), summarise, csum=length(id))
    }
    

    使用下面的数据,订单现在不再起作用(见下面的评论)。

    get_data("apple", "htc")
    
      calmonth  product2 csum
    1        1 apple&htc    1
    2        2 apple&htc    2
    
    get_data("htc", "apple")
    
      calmonth  product2 csum
    1        1 htc&apple    1
    2        2 htc&apple    2
    

    我知道这不是您问题的直接答案,但我觉得这种方法很干净。

    评论后编辑

    您得到csum=0 的原因仅仅是您正在搜索错误的正则表达式模式,即a something in between b 而不是apple ... htc。您需要构建正确的正则表达式模式,即paste0(a, ".*", b).

    这里有一个完整的解决方案。我不会称它为漂亮的代码,但无论如何(请注意,我更改了数据以表明它可以概括几个月)。

    library(plyr)
    
    df2 <- read.table(text="
      id   calmonth        product
     101       01           apple
     102       01 apple&nokia&htc
     103       01             htc
     104       02       apple&htc
     104       02       apple&htc",  header=T)
    
    xandy <- function(a, b) {
      pattern <- paste0(a, ".*", b)
      d1 <- df2[grep(pattern, df2$product), ]
      d1$product <- paste0(a,"&", b)
      ddply(d1, .(calmonth), summarise, 
            csum=length(calmonth),
            product=unique(product))  
    }
    xandy("apple", "htc")
    
      calmonth csum   product
    1        1    1 apple&htc
    2        2    2 apple&htc
    

    【讨论】:

    • @Mark Heckmann谢谢!但是为什么不尝试df2[grep("apple.*htc",df2$product),] 来获得上述结果呢?其实我真正想做的是计算相应的 id 包含响应产品。
    • 再次感谢!但是当我传递参数xandy('htc','nokia')时,出现错误:replacement has 1 row, data has 0 ,因为没有任何返回值。而且,我想知道ddply()是否不能包含参数,我发现参数不能正确过去。
    • 是的,因为您选择的方法取决于顺序。您必须以相反的顺序指定请求,即xandy("nokia", "htc") 才能工作。正则表达式模式"htc.*nokia" 不会像示例中那样找到apple&amp;nokia&amp;htc。因此它返回长度为零的数据帧并发生错误。我首先描述的方法没有这个缺点。然而,我知道它需要一些修改以满足您的需求。
    • 是的,我明白了,我的意思是如果一个 id 只使用产品 'htc' 但在月份 '01' 和 '02' 中没有使用 'nokia',那么结果将是两个零线,isn不是吗?发生这种情况时会出现错误。并且,我提出了一个类似这样的新问题,期待您的回答,谢谢![link]stackoverflow.com/questions/21041252/…
    • 我先完成了我建议的方法。
    猜你喜欢
    • 2020-10-10
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-05-19
    • 1970-01-01
    • 2021-12-31
    • 1970-01-01
    相关资源
    最近更新 更多