【问题标题】:Manipulate char vectors inside a data.table object in R在 R 中的 data.table 对象中操作字符向量
【发布时间】:2015-11-18 16:40:55
【问题描述】:

我对使用 data.table 并了解其所有细微之处还有些陌生。 我查看了文档和 SO 中的其他示例,但找不到我想要的,所以请帮忙!

我有一个 data.table,它基本上是一个字符向量(每个条目都是一个句子)

DT=c("I love you","she loves me")
DT=as.data.table(DT)
colnames(DT) <- "text"
setkey(DT,text)

# > DT
#            text
# 1:   I love you
# 2: she loves me

我想做的是能够在 DT 对象内执行一些基本的字符串操作。例如,添加一个新列,其中我将有一个 char 向量,其中每个条目都是“text”列中字符串中的一个 WORD。

所以我想有一个新的列 charvec where

> DT[1]$charvec
[1] "I" "love "you"

当然,我想用data.table的方式来做,超快,因为我需要在>1Go文件的fils上做这种事情,并且使用更复杂和计算量大的函数。所以不要使用 APPLY、LAPPLY 和 MAPPLY

我最接近的尝试如下:

myfun1 <- function(sentence){strsplit(sentence," ")}
DU1 <- DT[,myfun1(text),by=text]
DU2 <- DU1[,list(charvec=list(V1)),by=text]
# > DU2
#            text      charvec
# 1:   I love you   I,love,you
# 2: she loves me she,loves,me

例如,要制作一个删除每个句子的第一个单词的函数,我这样做了

myfun2 <- function(l){l[[1]][-1]}
DV1 <- DU2[,myfun2(charvec),by=text]
DV2 <- DV1[,list(charvec=list(V1)),by=text]
# > DV2
#            text  charvec
# 1:   I love you love,you
# 2: she loves me loves,me

问题是,在 charvec 列中,我有一个列表而不是向量...

> str(DU2[1]$charvec)
# List of 1
# $ : chr [1:3] "I" "love" "you"

1) 我怎样才能做我想做的事? 我正在考虑使用的其他类型的函数是对 char 向量进行子集化,或者对其应用一些哈希等。

2) 顺便说一句,我可以用一条线而不是两条线到达 DU2 或 DV2 吗? 3) 我不太了解 data.table 的语法。为什么使用 [..] 内的命令 list() 时,V1 列消失了? 4) 在另一个线程上,我读到了一些关于函数cSplit 的内容。

。有什么好处吗?是适配data.table对象的函数吗?

非常感谢

更新

感谢@Ananda Mahto 也许我应该让自己更清楚我的最终目标 我有一个存储为字符串的 10,000,000 个句子的巨大文件。 作为该项目的第一步,我想对每个句子的前 5 个单词进行哈希处理。 10,000,000 个句子甚至不会进入我的记忆,所以我首先将 1,000,000 个句子分成 10 个文件,大约是 10x 1Go 文件。 以下代码在我的笔记本电脑上只为一个文件需要几分钟。

library(data.table); library(digest);
num_row=1000000
DT <- fread("sentences.txt",nrows=num_row,header=FALSE,sep="\t",colClasses="character")
DT=as.data.table(DT)
colnames(DT) <- "text"
setkey(DT,text)
rawdata <- DT

hash2 <- function(word){ #using library(digest)
        as.numeric(paste("0x",digest(word,algo="murmur32"),sep=""))
}

那么,

print(system.time({ 

        colnames(rawdata) <- "sentence"
        rawdata <- lapply(rawdata,strsplit," ")

        sentences_begin <- lapply(rawdata$sentence,function(x){x[2:6]})
        hash_list <- sapply(sentences_begin,hash2)
        # remove(rawdata)
})) ## end of print system.time for loading the data

我知道我将 R 推到了极限,但我正在努力寻找更快的实现,我正在考虑 data.table 功能......因此我的所有问题

这是一个不包括 lapply 的实现,但它实际上更慢!

print(system.time({
myfun1 <- function(sentence){strsplit(sentence," ")}
DU1 <- DT[,myfun1(text),by=text]
DU2 <- DU1[,list(charvec=list(V1)),by=text]

myfun2 <- function(l){l[[1]][2:6]}
DV1 <- DU2[,myfun2(charvec),by=text]
DV2 <- DV1[,list(charvec=list(V1)),by=text]

rebuildsentence <- function(S){
        paste(S,collapse=" ") }

myfun3 <- function(l){hash2(rebuildsentence(l[[1]]))}

DW1 <- DV2[,myfun3(charvec),by=text]

})) #end of system.time

在这个使用数据文件的实现中,没有 lapply,所以我希望散列会更快。但是,因为在每一列中我都有一个列表而不是一个 char 向量,所以这可能会显着减慢(?)整个事情。

使用上面的第一个代码(lapply/sapply)在我的笔记本电脑上花费了 1 个多小时。我希望通过更有效的数据结构来加快速度?使用 Python、Java 等的人......在几秒钟内完成类似的工作。

当然,另一种方法是找到更快的哈希函数,但我假设 digest 包中的那个已经优化了。

【问题讨论】:

  • 你的最终目标是什么?
  • 如上所述,主要目标是避免使用 lapply、sapply、mapply,仅使用 data.table 语法非常快速地操作大文件的字符串 data.table(超过 1Go,基本上内存和速度允许的最大值....)。我希望能够将字符串列转换为 char 向量列到同一个表中,然后取它的一个子集,比如前 5 个单词或最后 5 个单词,我想计算这 5 个词,或者整个句子的散列......我已经使用 data.table 和 lapply 样式完成了 not ,它花了几分钟时间。太长了。
  • 对于我的教育,从表 DU2 及其第二列(每个元素都是一个列表)你将如何重建原始文本(单个字符串)?
  • 您是否尝试过正则表达式解决方案来提取前五个单词并直接对其进行哈希处理?
  • @AnandaMahto 什么是正则表达式?

标签: r string data.table strsplit


【解决方案1】:

我不太确定你在追求什么,但你可以试试我的“splitstackshape”包中的cSplit_l 以进入你的列表列:

library(splitstackshape)
DU <- cSplit_l(DT, "DT", " ")

然后,您可以编写如下函数来从列表列中删除值:

RemovePos <- function(inList, pos = 1) {
  lapply(inList, function(x) x[-c(pos[pos <= length(x)])])
}

示例用法:

DU[, list(RemovePos(DT_list, 1)), by = DT]
#              DT       V1
# 1:   I love you love,you
# 2: she loves me loves,me
DU[, list(RemovePos(DT_list, 2)), by = DT]
#              DT     V1
# 1:   I love you  I,you
# 2: she loves me she,me
DU[, list(RemovePos(DT_list, c(1, 2))), by = DT]
#              DT  V1
# 1:   I love you you
# 2: she loves me  me

更新

基于您对 `lapply 的厌恶,也许您可​​以尝试以下方法:

## make a copy of your "text" column
DT[, vals := text]

## Use `cSplit` to create a "long" dataset. 
## Add a column to indicate the word's position in the text.
DTL <- cSplit(DT, "vals", " ", "long")[, ind := sequence(.N), by = text][]
DTL
#            text  vals ind
# 1:   I love you     I   1
# 2:   I love you  love   2
# 3:   I love you   you   3
# 4: she loves me   she   1
# 5: she loves me loves   2
# 6: she loves me    me   3

## Now, you can extract values easily
DTL[ind == 1]
#            text vals ind
# 1:   I love you    I   1
# 2: she loves me  she   1
DTL[ind %in% c(1, 3)]
#            text vals ind
# 1:   I love you    I   1
# 2:   I love you  you   3
# 3: she loves me  she   1
# 4: she loves me   me   3

更新 2

我不知道你得到的是什么类型的时间,但正如我在评论中提到的,你也许可以尝试使用正则表达式,这样你就不必拆分然后将字符串重新粘贴在一起。

这是一个示例......

设置一些数据来玩:

library(data.table)
DT <- data.table(
  text = c("This is a sentence with a lot of words.",
           "This is a sentence with some more words.",
           "Words and words and even some more words.",
           "But, I don't know how you want to deal with punctuation...",
           "Just one more sentence, for easy multiplication.")
)

DT2 <- rbindlist(replicate(10000/nrow(DT), DT, FALSE))
DT3 <- rbindlist(replicate(1000000/nrow(DT), DT, FALSE))

测试 gsub 模式以从每个句子中提取 5 个单词....

## Regex to extract first five words -- this should work....
patt <- "^((?:\\S+\\s+){4}\\S+).*"

## Check out some of the timings
system.time(temp <- DT2[, gsub(patt, "\\1", text)])
#    user  system elapsed 
#    0.03    0.00    0.03 
system.time(temp2 <- DT3[, gsub(patt, "\\1", text)])
#    user  system elapsed 
#       3       0       3 
head(temp)
# [1] "This is a sentence with"     "This is a sentence with"     "Words and words and even"   
# [4] "But, I don't know how"       "Just one more sentence, for" "This is a sentence with" 

我猜你想要做什么......

## I'm assuming you want something like this....
## Takes about a minute on my system. 
## ... but note the system time for the creation of "temp2" (without digest)
## Not sure if I interpreted your hash requirement correctly....
system.time(out <- DT3[
  , firstFive := gsub(patt, "\\1", text)][
  , firstFiveHash := hash2(firstFive), by = 1:nrow(DT3)][])
#    user  system elapsed 
#   62.14    0.05   62.20 

head(out)
#                                                          text                   firstFive firstFiveHash
# 1:                    This is a sentence with a lot of words.     This is a sentence with    4179639471
# 2:                   This is a sentence with some more words.     This is a sentence with    4179639471
# 3:                  Words and words and even some more words.    Words and words and even    2556713080
# 4: But, I don't know how you want to deal with punctuation...       But, I don't know how    3765680401
# 5:           Just one more sentence, for easy multiplication. Just one more sentence, for     298317689
# 6:                    This is a sentence with a lot of words.     This is a sentence with    4179639471

【讨论】:

  • 我不想要任何使用 lapply、mapply 或 sapply 的解决方案。我想使用 data.table 语法。据记录,这些函数在 data.table 中已过时,并且在寻找性能(或“纯高效的 R 代码”)时应该永远使用查看cran.r-project.org/web/packages/data.table/vignettes/…
  • @FaguiCurtain,使用“长”方向的cSplit,使用.N添加数字索引,然后您可以根据.N创建的值进行子集化。我不确定您从哪里得到lapply 已被“data.table”废弃并且应该永远与“data.table”一起使用的想法。
  • 确实,*apply 函数是一种在 within data.table 语法中完成工作的有效方法。您可能需要检查s3.amazonaws.com/assets.datacamp.com/img/blog/…
  • @AnandaMahto 感谢您抽出宝贵时间。我不知道你想用命令 gsub 做什么?我也不明白你最后一行的语法。在前五列中,我看到超过 5 个单词。另外,为了使问题更简单,在我使用的原始文件中,只有空格,没有其他标点符号。
  • 三列:::是什么意思?
猜你喜欢
  • 2018-06-01
  • 1970-01-01
  • 2018-12-18
  • 1970-01-01
  • 2011-11-10
  • 1970-01-01
  • 2017-09-04
  • 2016-03-30
  • 1970-01-01
相关资源
最近更新 更多