【问题标题】:How do I unpack tuple format in R?如何在 R 中解压缩元组格式?
【发布时间】:2022-03-12 16:23:57
【问题描述】:

这是数据集。

library(data.table)

x <- structure(list(id = c("A", "B" ),
                    segment_stemming = c("[('Brownie', 'Noun'), ('From', 'Josa'), ('Pi', 'Noun')]", 
                                          "[('Dung-caroon-gye', 'Noun'), ('in', 'Josa'), ('innovation', 'Noun')]" )), 
               row.names = c(NA, -2L), 
               class = c("data.table", "data.frame" ))

x
# id                                                     segment_stemming
# 1:  A               [('Brownie', 'Noun'), ('From', 'Josa'), ('Pi', 'Noun')]
# 2:  B [('Dung-caroon-gye', 'Noun'), ('in', 'Josa'), ('innovation', 'Noun')]

我想将元组拆分成行。这是我的预期结果。

id             segment_stemming
A              ('Brownie', 'Noun')
A              ('From', 'Josa')
A              ('Pi', 'Noun')
B              ('Dung-caroon-gye', 'Noun')
B              ('in', 'Josa')
B              ('innovation', 'Noun')

我已经使用 R 搜索了元组格式,但找不到任何线索来得出结果。

【问题讨论】:

  • 你想保留括号"()"吗? ('Brownie', 'Noun') 而不是 'Brownie', 'Noun'
  • 先用Python怎么样?

标签: r dataframe data.table tuples


【解决方案1】:

data.table接近

这是一个使用data.table + reticulate的选项

library(reticulate)
library(data.table)
setDT(x)[
  ,
  segment_stemming := gsub("(\\(.*?\\))", '\"\\1\"', segment_stemming)
][
  ,
  lapply(.SD, py_eval),
  id
]

给了

   id            segment_stemming
1:  A         ('Brownie', 'Noun')
2:  A            ('From', 'Josa')
3:  A              ('Pi', 'Noun')
4:  B ('Dung-caroon-gye', 'Noun')
5:  B              ('in', 'Josa')
6:  B      ('innovation', 'Noun')

另一个data.table 选项使用strsplit + trimws,如下所示

library(data.table)
setDT(x)[
  ,
  .(segment_stemming = trimws(
    unlist(strsplit(segment_stemming, "(?<=\\)),\\s+(?=\\()", perl = TRUE)),
    whitespace = "\\[|\\]"
  )),
  id
]

给予

   id            segment_stemming
1:  A         ('Brownie', 'Noun')
2:  A            ('From', 'Josa')
3:  A              ('Pi', 'Noun')
4:  B ('Dung-caroon-gye', 'Noun')
5:  B              ('in', 'Josa')
6:  B      ('innovation', 'Noun')

基础 R

一些基本的 R 选项也应该可以工作

with(
  x,
  setNames(
    rev(
      stack(
        tapply(
          segment_stemming,
          id,
          function(v) {
            trimws(
              unlist(strsplit(v, "(?<=\\)),\\s+(?=\\()", perl = TRUE)),
              whitespace = "\\[|\\]"
            )
          }
        )
      )
    ),
    names(x)
  )
)

或

with(
  x,
  setNames(
    rev(
      stack(
        setNames(
          regmatches(segment_stemming, gregexpr("\\(.*?\\)", segment_stemming)),
          id
        )
      )
    ),
    names(x)
  )
)

【讨论】:

  • reticulate::py_eval 很有趣。作为旁注,对 OP 的建议是,一种更方便的存储和操作此类数据的方法将是一个“平面”data.frame,例如,由以下人员返回:rrapply::rrapply(setNames(lapply(x$segment_stemming, reticulate::py_eval), x$id), how = "melt")
  • @alexis_laz 不错的建议,谢谢!
【解决方案2】:

这是一种使用separate_rows的方法:

library(tidyverse)

x %>% 
  mutate(segment_stemming = gsub("\\[|\\]", "", segment_stemming)) %>% 
  separate_rows(segment_stemming, sep = ",\\s*(?![^()]*\\))")

# A tibble: 6 x 2
  id    segment_stemming           
  <chr> <chr>                      
1 A     ('Brownie', 'Noun')        
2 A     ('From', 'Josa')           
3 A     ('Pi', 'Noun')             
4 B     ('Dung-caroon-gye', 'Noun')
5 B     ('in', 'Josa')             
6 B     ('innovation', 'Noun') 

一种获得更好结果的方法,通过一些操作(unnest_wider 不是必需的)。

x %>% 
  mutate(segment_stemming = gsub("\\[|\\]", "", segment_stemming)) %>% 
  separate_rows(segment_stemming, sep = ",\\s*(?![^()]*\\))") %>% 
  mutate(segment_stemming = segment_stemming %>% 
           str_remove_all("[()',]") %>% 
           str_split(" ")) %>% 
  unnest_wider(segment_stemming)

# A tibble: 6 x 3
  id    ...1            ...2 
  <chr> <chr>           <chr>
1 A     Brownie         Noun 
2 A     From            Josa 
3 A     Pi              Noun 
4 B     Dung-caroon-gye Noun 
5 B     in              Josa 
6 B     innovation      Noun 

【讨论】:

    【解决方案3】:

    这是另一个潜在的选择:

    library(data.table)
    
    dt <- structure(list(id = c("A", "B" ), segement_stemming = c("[('Brownie', 'Noun'), ('From', 'Josa'), ('Pi', 'Noun')]", "[('Dung-caroon-gye', 'Noun'), ('in', 'Josa'), ('innovation', 'Noun')]" )), row.names = c(NA, -2L), class = c("data.table", "data.frame" ))
    
    dt2 <- dt[, c(segement_stemming = strsplit(segement_stemming, "(?<=[^']),", perl = TRUE)), by = id]
    dt2[, names(dt2) := lapply(.SD, function(x) gsub("\\[|\\]", "", x))]
    dt2
    #>    id           segement_stemming
    #> 1:  A         ('Brownie', 'Noun')
    #> 2:  A            ('From', 'Josa')
    #> 3:  A              ('Pi', 'Noun')
    #> 4:  B ('Dung-caroon-gye', 'Noun')
    #> 5:  B              ('in', 'Josa')
    #> 6:  B      ('innovation', 'Noun')
    

    由reprex package (v2.0.1) 于 2022-03-11 创建

    【讨论】:

      【解决方案4】:
      x[,.(segment_stemming = unlist(str_extract_all(segment_stemming, "\\(.*?\\)"))), by = id]
      

      或者您可以使用tidyr::unnest。这样一来str_extract_all就只有一个电话了:

      x[, segment_stemming := str_extract_all(segment_stemming, "\\(.*?\\)")]
      unnest(x, segment_stemming)
      

      【讨论】:

      • 伟大的洞察力,一次提取而不是分裂。
      【解决方案5】:

      data.table 方式如下:

      library(stringr)
      
      x [, segment_stemming:=gsub("\\[|\\]", "", segment_stemming, perl = T)] #remove brackets
      x [, parsed := str_split(segment_stemming, "\\),")]                     # split string
      out <- x[, .(unlist(parsed, recursive = F)), by = .(id)]                # unlist elements
      out [ , V1  := gsub("\\)?$",")", V1)][]                                 # adjust format
      
             id                          V1
         <char>                      <char>
      1:      A         ('Brownie', 'Noun')
      2:      A            ('From', 'Josa')
      3:      A              ('Pi', 'Noun')
      4:      B ('Dung-caroon-gye', 'Noun')
      5:      B              ('in', 'Josa')
      6:      B      ('innovation', 'Noun')
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2012-02-28
        • 2022-03-18
        • 1970-01-01
        • 1970-01-01
        • 2010-10-27
        • 2014-08-26
        相关资源
        最近更新 更多