【问题标题】:Extracting state abbreviation from string with 5 and 9 digit zip code in R从R中具有5位和9位邮政编码的字符串中提取状态缩写
【发布时间】:2021-12-16 15:33:36
【问题描述】:

我正在尝试从具有不同格式的数据框中的地址列中提取状态缩写。示例:

"123 Any St., Some City, IL 65234 United States"
"456 Any Other St That Town, CA 62626-1234 US"

我使用的代码适用于具有 5 位邮政编码的字符串,但不适用于具有 9 位邮政编码的字符串:

df$state <- str_extract(df$address, "\\b[A-Z]{2}(?=\\s+\\d{5}$)")

如何更改它以提取状态,然后是 5 位和 9 位邮政编码?

【问题讨论】:

  • 匹配状态缩写与正则表达式(?&lt;=, )(?:AL|AK|AS|AZ|AR|CA|CO|CT|DE|DC|FM|FL|GA|GU|HI|ID|IL|IN|IA|KS|KY|LA|ME|MH|MD|MA|MI|MN|MS|MO|MT|NE|NV|NH|NJ|NM|NY|NC|ND|MP|OH|OK|OR|PW|PA|PR|RI|SC|SD|TN|TX|UT|VT|VI|VA|WA|WV|WI|WY)(?=\d{5}(?:-\d{4})? 。请注意,正则表达式以空格结尾。

标签: r regex string


【解决方案1】:

您可以使用 tidyr::extract 函数,该函数在您使用 tibble/dataframe 时效果特别好。在您的情况下,我正在执行以下操作:将数据放入名为 df 的数据框/tibble 中,使用 tidyr::extract 将信息拉入两列 - zipcode 和 state。

tidyr::extract 函数使用括号来区分您想要哪些信息在哪些列中。因此,由于我要提取到两个不同的列,因此这里有两组括号,其中包含正则表达式。第一个正则表达式是\\d{5}|\\d{5}-\\d{4},表示恰好匹配 5 位数字或 5 位数字,后跟一个破折号,然后是 4 位数字。下一个正则表达式是.{1,},它匹配任意字符 1 到任意次数。在我找到这两个表达式之前,我会根据需要多次匹配任何字符,直到找到带有.{1,} 的邮政编码。我用\\s(空格)将这两列分开。

library(tidyverse)

df <- tibble(zipcodes = c("123 Any St., Some City, IL 65234 United States",
"456 Any Other St That Town, CA 62626-1234 US"))

df %>% 
  tidyr::extract(zipcodes, into = c("zipcodes", "state"),
          ".{1,}(\\d{5}|\\d{5}-\\d{4})\\s(.{1,})")

【讨论】:

    【解决方案2】:

    当我将您的代码用于示例字符串上的 5 位邮政编码时,它不起作用并返回 NAs。

    如果我们删除最后一个$,那么它适用于 5 位和 9 位邮政编码:

    teststr <- c("123 Any St., Some City, IL 65234 United States",
                 "456 Any Other St That Town, CA 62626-1234 US")
    
    stringr::str_extract(teststr, "\\b[A-Z]{2}(?=\\s+\\d{5})")
    #> [1] "IL" "CA"
    

    由reprex package (v2.0.1) 于 2021 年 11 月 2 日创建

    【讨论】:

    • 谢谢!这非常有效。
    猜你喜欢
    • 1970-01-01
    • 2013-02-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-01-31
    • 1970-01-01
    • 1970-01-01
    • 2023-02-21
    相关资源
    最近更新 更多