【问题标题】:How to extract all numbers in a string as a vector如何将字符串中的所有数字提取为向量
【发布时间】:2021-02-24 23:02:35
【问题描述】:

有没有办法将字符串中的所有数字提取为向量?我有一个不遵循任何特定模式的大型数据集,因此使用 extract + regex 模式不一定会提取所有数字。因此,例如对于如下所示的每一行数据框:

c("3.2% 1ST $100000 AND 1.1% BALANCE", "3.3% 1ST $100000 AND 1.2% BALANCE AND $3000 BONUS FULL PRICE ONLY", 
"$4000", "3.3% 1ST $100000 AND 1.2% BALANCE", "3.3% 1ST $100000 AND 1.2% BALANCE", 
"3.2 - $100000")

[1] "3.2% 1ST $100000 AND 1.1% BALANCE"                                
[2] "3.3% 1ST $100000 AND 1.2% BALANCE AND $3000 BONUS FULL PRICE ONLY"
[3] "$4000"                                                            
[4] "3.3% 1ST $100000 AND 1.2% BALANCE"                                
[5] "3.3% 1ST $100000 AND 1.2% BALANCE"                                
[6] "3.2 - $100000"   

我想要这样的输出:

[1] "3.2 100000 1.1"                                
[2] "3.3 100000 1.2 3000"
[3] "4000"                                                            
[4] "3.3 100000 1.2 "                                
[5] "3.3 100000 1.2 "                                
[6] "3.2 100000 "   

我查看了资源,找到了这个链接:https://statisticsglobe.com/extract-numbers-from-character-string-vector-in-r

regmatches(x, gregexpr("[[:digit:]]+", x))

上面的函数似乎有效,但它不能同时对各种数字执行此任务。我知道"[[:digit:]]+" 只查找整数,但我们如何更改它以使其涵盖各种数字?

【问题讨论】:

标签: r regex extract


【解决方案1】:

我们还需要在匹配模式中添加.

sapply(regmatches(x, gregexpr("\\b[[:digit:].]+\\b", x)), paste, collapse= ' ')
#[1] "3.2 100000 1.1"    
#[2] "3.3 100000 1.2 3000" 
#[3] "4000"              
#[4] "3.3 100000 1.2"   
#[5] "3.3 100000 1.2"     
#[6] "3.2 100000"   

【讨论】:

  • 感谢@akrun,但随后它会在 1ST 中提取 1。我只是在寻找纯数字
  • @Roozbeh_you 抱歉,已更新单词边界。你能检查一下吗
【解决方案2】:

Akrun 的答案是完美的,但只是为了添加另一个解决方案,使用一个包来创建我最近发现的正则表达式模式。

library(stringr)
library(rebus)
library(magrittr)

pattern = one_or_more(DIGIT) %R% optional(DOT) %R% optional(one_or_more(DIGIT))

str_remove(x, "1ST") %>% 
str_match_all( pattern = pattern) %>% 
  lapply( function(x) paste(as.vector(x), collapse = " ")) %>% 
  unlist()

【讨论】:

  • 谢谢约翰。你的回答也是正确的;但是,我应该补充一点,并非所有字符串都具有 1ST,其中一些具有例如 1-ST。这可能会导致您在使用 str-remove 时出现问题,对吧?
【解决方案3】:

您可以使用负前瞻正则表达式:

stringr::str_extract_all(x, '\\d+(\\.\\d+)?(?![A-Z])')

#[[1]]
#[1] "3.2"    "100000" "1.1"   

#[[2]]
#[1] "3.3"    "100000" "1.2"    "3000"  

#[[3]]
#[1] "4000"

#[[4]]
#[1] "3.3"    "100000" "1.2"   

#[[5]]
#[1] "3.3"    "100000" "1.2"   

#[[6]]
#[1] "3.2"    "100000"

如果您希望输出为一个字符串:

sapply(stringr::str_extract_all(x, '\\d+(\\.\\d+)?(?![A-Z])'), paste, collapse = ' ')
#[1] "3.2 100000 1.1"      "3.3 100000 1.2 3000" "4000"               
#[4] "3.3 100000 1.2"      "3.3 100000 1.2"      "3.2 100000"  

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2013-01-10
    • 1970-01-01
    • 1970-01-01
    • 2017-02-19
    • 2021-07-01
    • 2016-06-24
    相关资源
    最近更新 更多