【问题标题】:extract comma separated strings提取逗号分隔的字符串
【发布时间】:2014-11-27 13:36:57
【问题描述】:

我有如下数据框。这是一个具有统一外观模式的样本集数据,但整个数据不是很统一:

locationid      address     
1073744023  525 East 68th Street, New York, NY      10065, USA
1073744022  270 Park Avenue, New York, NY 10017, USA      
1073744025  Rockefeller Center, 50 Rockefeller Plaza, New York, NY 10020, USA 
1073744024  1251 Avenue of the Americas, New York, NY 10020, USA
1073744021  1301 Avenue of the Americas, New York, NY 10019, USA 
1073744026  44 West 45th Street, New York, NY 10036, USA

我需要从这个地址中找到城市和国家名称。我尝试了以下方法:

1) strsplit 这给了我一个列表,但我无法从中访问最后一个或倒数第三个元素。

2) 正则表达式 找国家很容易

str_sub(str_extract(address, "\\d{5},\\s.*"),8,11)

但对于城市

str_sub(str_extract(address, ",\\s.+,\\s.+\\d{5}"),3,comma_pos)

我找不到comma_pos,因为它让我再次遇到同样的问题。 我相信有一种更有效的方法可以使用上述任何一种方法来解决这个问题。

【问题讨论】:

    标签: regex r string strsplit


    【解决方案1】:

    试试这个代码:

    library(gsubfn)
    
    cn <- c("Id", "Address", "City", "State", "Zip", "Country")
    
    pat <- "(\\d+) (.+), (.+), (..) (\\d+), (.+)"
    read.pattern(text = Lines, pattern = pat, col.names = cn, as.is = TRUE)
    

    给出以下data.frame,从中可以轻松挑选出组件:

              Id                                  Address     City State   Zip Country
    1 1073744023                     525 East 68th Street New York    NY 10065     USA
    2 1073744022                          270 Park Avenue New York    NY 10017     USA
    3 1073744025 Rockefeller Center, 50 Rockefeller Plaza New York    NY 10020     USA
    4 1073744024              1251 Avenue of the Americas New York    NY 10020     USA
    5 1073744021              1301 Avenue of the Americas New York    NY 10019     USA
    6 1073744026                      44 West 45th Street New York    NY 10036     USA
    

    解释它使用这种模式(引号内的反斜杠必须加倍):

    (\d+) (.+), (.+), (..) (\d+), (.+)
    

    通过以下 debuggex 铁路图进行可视化 - 更多信息请参见 Debuggex Demo

    并用文字解释如下:

    • "(\\d+)" - 一位或多位数字(代表Id)后跟
    • " " 后跟一个空格
    • "(.+)" - 任何非空字符串(代表Address)后跟
    • ", " - 一个逗号和一个空格,后跟
    • "(.+)" - 任何非空字符串(代表City)后跟
    • ", " - 一个逗号和一个空格,后跟
    • "(..)" - 两个字符(代表State)后跟
    • " " - 后跟一个空格
    • "(\\d+)" - 一位或多位数字(代表Zip)后跟
    • ", " - 一个逗号和一个空格,后跟
    • "(.+)" - 任何非空字符串(代表Country

    它之所以有效,是因为正则表达式是贪婪的,每次正则表达式的后续部分无法匹配时,它总是试图找到可以匹配回溯的最长字符串。

    这种方法的优点是正则表达式非常简单直接,整个代码非常简洁,一个read.pattern 语句就可以完成所有工作:

    注意:我们将此用于Lines

    Lines <- "1073744023 525 East 68th Street, New York, NY 10065, USA
    1073744022 270 Park Avenue, New York, NY 10017, USA
    1073744025 Rockefeller Center, 50 Rockefeller Plaza, New York, NY 10020, USA
    1073744024 1251 Avenue of the Americas, New York, NY 10020, USA
    1073744021 1301 Avenue of the Americas, New York, NY 10019, USA
    1073744026 44 West 45th Street, New York, NY 10036, USA"
    

    【讨论】:

    • 我更喜欢那个演示。
    • 其中有不少。我已经列出了我在底部附近的 gsubfn 主页上找到的那些。 gsubfn.googlecode.com
    【解决方案2】:

    拆分数据

     ss <- strsplit(data,",")`
    

    然后

    n <- sapply(s,len)
    

    将给出元素的数量(因此您可以向后工作)。那么

    mapply(ss,"[[",n)
    

    给你最后一个元素。或者你可以这样做

    sapply(ss,tail,1)
    

    获取最后一个元素。

    要获得您需要的倒数第二个(或更一般地说)

    sapply(ss,function(x) tail(x,2)[1])
    

    【讨论】:

    • sapply(ss,tail,1) 有效,但 sapply(ss,tail,2) 给我错误:错误:错误的结果大小 (12),预期为 6 或 1
    【解决方案3】:

    这是一种使用 tidyr 包的方法。就个人而言,我只是使用 tidyr 包的extract 将整个内容拆分为所有不同的元素。这使用了正则表达式,但与您要求的方式不同。

    library(tidyr)
    
    extract(x, address, c("address", "city", "state", "zip", "state"), 
        "([^,]+),\\s([^,]+),\\s+([A-Z]+)\\s+(\\d+),\\s+([A-Z]+)")
    
    ##   locationid                       address     city state   zip state
    ## 1 1073744023          525 East 68th Street New York    NY 10065   USA
    ## 2 1073744022               270 Park Avenue New York    NY 10017   USA
    ## 3 1073744025          50 Rockefeller Plaza New York    NY 10020   USA
    ## 4 1073744024   1251 Avenue of the Americas New York    NY 10020   USA
    ## 5 1073744021   1301 Avenue of the Americas New York    NY 10019   USA
    ## 6 1073744026           44 West 45th Street New York    NY 10036   USA
    

    这是对来自http://www.regexper.com/ 的正则表达式的直观解释:

    【讨论】:

      【解决方案4】:

      我想你想要这样的东西。

      > x <- "1073744026 44 West 45th Street, New York, NY 10036, USA"
      > regmatches(x, gregexpr('^[^,]+, *\\K[^,]+', x, perl=T))[[1]]
      [1] "New York"
      > regmatches(x, gregexpr('^[^,]+, *[^,]+, *[^,]+, *\\K[^\n,]+', x, perl=T))[[1]]
      [1] "USA"
      

      正则表达式解释:

      • ^ 断言我们处于起步阶段。
      • [^,]+ 匹配任何字符,但不匹配 , 一次或多次。如果您的数据框包含空字段,请将其更改为 [^,]*
      • , 匹配文字 ,
      • &lt;space&gt;* 匹配零个或多个空格。
      • \K 从打印中丢弃以前匹配的字符。与\K 后面的模式匹配的字符将显示为输出。

      【讨论】:

      • 嗨。我是正则表达式的初学者。你能解释一下这些是什么意思吗?
      【解决方案5】:

      这个模式怎么样:

      ,\s(?<city>[^,]+?),\s(?<shortCity>[^,]+?)(?i:\d{5},)(?<country>\s.*)
      

      此模式将匹配这三个组:

      1. “组”:“城市”,“价值”:“纽约”
      2. “group”:“shortCity”,“value”:“NY”
      3. “组”:“国家”,“值”:“美国”

      【讨论】:

        【解决方案6】:

        使用rex 构造正则表达式可能会使这类任务更简单一些。

        x <- data.frame(
          locationid = c(
            1073744023,
            1073744022,
            1073744025,
            1073744024,
            1073744021,
            1073744026
            ),
          address = c(
            '525 East 68th Street, New York, NY      10065, USA',
            '270 Park Avenue, New York, NY 10017, USA',
            'Rockefeller Center, 50 Rockefeller Plaza, New York, NY 10020, USA',
            '1251 Avenue of the Americas, New York, NY 10020, USA',
            '1301 Avenue of the Americas, New York, NY 10019, USA',
            '44 West 45th Street, New York, NY 10036, USA'
            ))
        
        library(rex)
        
        sep <- rex(",", spaces)
        
        re <-
          rex(
            capture(name = "address",
              except_some_of(",")
            ),
            sep,
            capture(name = "city",
              except_some_of(",")
            ),
            sep,
            capture(name = "state",
              uppers
            ),
            spaces,
            capture(name = "zip",
              some_of(digit, "-")
            ),
            sep,
            capture(name = "country",
              something
            ))
        
        re_matches(x$address, re)
        #>                      address     city state   zip country
        #>1        525 East 68th Street New York    NY 10065     USA
        #>2             270 Park Avenue New York    NY 10017     USA
        #>3        50 Rockefeller Plaza New York    NY 10020     USA
        #>4 1251 Avenue of the Americas New York    NY 10020     USA
        #>5 1301 Avenue of the Americas New York    NY 10019     USA
        #>6         44 West 45th Street New York    NY 10036     USA
        

        此正则表达式还将处理 9 位邮政编码 (12345-1234) 和美国以外的国家/地区。

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多