【问题标题】:Text mining in R to search and extract informationR中的文本挖掘以搜索和提取信息
【发布时间】:2016-11-26 09:14:08
【问题描述】:

有没有一种方法可以在数据行中搜索模式,然后将它们存储在新表的单独列中?例如,如果我需要从下面的身体中提取金额、钞票和硬币,你认为在 R 上可以达到预期的结果吗

user_id   |        ts |                 body                    |  address |    
3633|      2016-09-29|  A wallet with amount = $ 100 has been found with 4 bills and 5 coins|   TEST |    
4266|      2016-07-20|  A purse having amount = $ 150 has been found with 40 bills and 15 coins|    NAME |
7566|      2016-07-20|  A pocket having amount = $ 200 has been found with 4 bills and 5 coins| HELLO |

(这是想要的结果)

user_id   | Amount | Bills| Coins|
3633      | $100   |    4  |     5|
4266      | $150   |    40 |    15|
7566      | $200   |    10 |    10|

【问题讨论】:

  • 是的,有可能。您将需要使用正则表达式。见?regex。 effect of this 的东西。

标签: r text pattern-matching analysis text-mining


【解决方案1】:

这是stringr 和lapply 的一种解决方案,但肯定还有更多。第一个子集仅包含 user.id 和 body 列,以提供如下内容:

df <- data.frame(user.id = c(3633, 4266, 7566),
      body = c("A wallet with amount = $ 100 has been found with 4 bills and 5 coins",
               "A purse having amount = $ 150 has been found with 40 bills and 15 coins",
               "A pocket having amount = $ 200 has been found with 4 bills and 5 coins"))

现在我们将正则表达式应用于df 的所有行,以将数字提取到列表中,取消列出,转换为指定列名的矩阵,将原始数据框中的cbind 和cbind 转置为user.id .

library(stringr)
mat <- t(matrix(unlist(lapply(df, str_match_all, "[0-9]+")[2]), nrow = nrow(df)))
colnames(mat) <- c("Amount", "Bills", "Coins")
outputdf <- cbind(df[1], mat)

这给了:

> outputdf
#  user.id Amount Bills Coins
#1    3633    100     4     5
#2    4266    150    40    15
#3    7566    200     4     5

我敢肯定也有更简洁的方法。

【讨论】:

    猜你喜欢
    • 2013-06-19
    • 2018-09-21
    • 2019-05-05
    • 2020-04-15
    • 1970-01-01
    • 2017-04-10
    • 1970-01-01
    • 2015-07-17
    • 1970-01-01
    相关资源
    最近更新 更多