【问题标题】:Replace strings in one column which match another column替换一列中与另一列匹配的字符串
【发布时间】:2021-07-03 16:20:38
【问题描述】:

我有一个看起来像这样的数据框(但适用于每个美国县)

Countyname Neighbour County Neighbour State
Autauga County, AL Chilton County AL
Autauga County, AL Dallas County AL
Baldwin County, AL Escambia County FL
Catron County, NM Apache County AZ

对于所有县,如果相邻县处于同一州,我想用缺失值 NA 替换“邻州”列中的值,如果它处于不同的州,我想保持不变。 IE。我想得到这样的结果:

Countyname Neighbour County Neighbour State
Autauga County, AL Chilton County NA
Autauga County, AL Dallas County NA
Baldwin County, AL Escambia County FL
Catron County, NM Apache County AZ

我正在考虑遍历每一行,如果“Countyname”列包含“Neighbour State”列中的条目,请将条目替换为 NA(例如,如果“Autauga County, AL”包含“AL”,则为第一行',将 'AL' 替换为 NA)。我该怎么做(或者有没有更有效的方法,因为这感觉很笨重)?

【问题讨论】:

  • 考虑对 CountyName 中的最后两个字符进行子串化,并使用 ifelseNeighbor State 进行比较。试一试,然后返回具体问题。
  • “我正在考虑循环遍历每一行” - R 是一种矢量化语言,因此它的函数默认执行此操作。
  • 我同意@Phil 的观点,循环尤其是 for 循环在 R 中大多是可以避免的,这就是这种语言的美妙之处。

标签: r string if-statement


【解决方案1】:

您可以使用str_detect()ifelse()

library(dplyr)
library(stringr)

df%>%
        mutate(Neighbour_State=ifelse(str_detect(Countyname, Neighbour_State), NA, Neighbour_State))

或者,最好是str_detectreplace()

df%>%
        mutate(Neighbour_State=replace(Neighbour_State, str_detect(Countyname, Neighbour_State), NA))

#Or with pipes:

df%>%
        mutate(Neighbour_State=Neighbour_State%>%replace(., str_detect(Countyname, .), NA))
          Countyname Neighbour_County Neighbour_State
1 Autauga County, AL   Chilton County            <NA>
2 Autauga County, AL    Dallas County            <NA>
3 Baldwin County, AL  Escambia County              FL
4  Catron County, NM    Apache County              AZ

【讨论】:

  • 谢谢!我不确定为什么其他用户的建议不起作用(表格保持不变),但使用替换有效。
  • 很高兴我能帮上忙。你可以检查这个:stackoverflow.com/help/someone-answers
【解决方案2】:

您还可以使用 regex 从县中查找 stateName,并使用函数 na_if 替换与 neighbour_county 匹配的值。像这样

  • regex '.*\\,\\s' 将匹配逗号和空格的所有内容,并将其替换为 '',即什么都没有。
  • 将 mutate 与 ifelse na_if replace case_when 或任何条件函数结合使用。
df <- read.table(header = T, text = "Countyname Neighbour_County    Neighbour_State
'Autauga County, AL'    'Chilton County'    AL
'Autauga County, AL'    'Dallas County' AL
'Baldwin County, AL'    'Escambia County'   FL
'Catron County, NM' 'Apache County' AZ")

library(dplyr, warn.conflicts = F)

df %>% 
  mutate(Neighbour_State = na_if(gsub('.*\\,\\s','', Countyname, perl = T), Neighbour_State))
#>           Countyname Neighbour_County Neighbour_State
#> 1 Autauga County, AL   Chilton County            <NA>
#> 2 Autauga County, AL    Dallas County            <NA>
#> 3 Baldwin County, AL  Escambia County              AL
#> 4  Catron County, NM    Apache County              NM

reprex package (v2.0.0) 于 2021 年 7 月 3 日创建

【讨论】:

    【解决方案3】:
    library(stringr)
    dat %>% 
        dplyr::mutate(
            Neighbour.State = dplyr::if_else(
                str_trim(unlist(str_split(dat$Countyname[1], ","))[2]) == Neighbour.State, "NA", Neighbour.State)
            )
    
              Countyname Neighbour.County Neighbour.State
    1 Autauga County, AL   Chilton County              NA
    2 Autauga County, AL    Dallas County              NA
    3 Baldwin County, AL  Escambia County              FL
    4  Catron County, NM    Apache County              AZ
    

    【讨论】:

      【解决方案4】:

      你不需要循环。您可以改用dplyr 包:

      data <- data %>%
        mutate(`Neighbour State` = ifelse(substr(Countyname, nchar(Countyname)-1, nchar(Countyname))== `Neighbour State`, NA,`Neighbour State`))
      

      【讨论】:

      • 哦,是的,忘记 substr 是基础 R,我会编辑!
      猜你喜欢
      • 2018-09-13
      • 1970-01-01
      • 2022-01-07
      • 2019-01-12
      • 1970-01-01
      • 2022-07-25
      • 2015-08-17
      • 1970-01-01
      • 2021-04-19
      相关资源
      最近更新 更多