【问题标题】:How to Construct Nested For Loops in R如何在 R 中构造嵌套的 For 循环
【发布时间】:2018-12-18 01:14:36
【问题描述】:

我正在使用 R 来匹配两个不同数据集中的名称。我想比较字符串。我基本上有两个字符串数据框,都包含位置 ID(不是唯一的)以及人的全名。对于某些人来说,一个数据框的全名可能包含两个姓氏。另一个数据框具有相同的位置代码(不是唯一的),但姓氏将只有两者之一(总是随机的两者中的哪一个)。

我想做的是做一个grep(),逐行第一个数据帧到并得到第二个输出的搜索结果。我的方法是这样做:

  1. 使用 paste() 函数,粘贴位置 ID 和名字。这将有助于匹配。但我真的需要匹配姓氏(可以是任何一个姓氏)。我们称这个新向量为location_first

  2. 在姓氏列上使用函数strsplit()。列表中的某些元素将只有一项,而其他元素(即具有两个姓氏的人)将在该元素中具有两项。我们可以将此列表称为strsplit_ln。

  3. 然后我会以循环的形式进行第二次粘贴:将strsplit_ln 的第一个元素粘贴到location_first,对其执行 grep,然后移动到strplit_ln 的下一个元素并对此进行grep。我想在我的控制台上的下沉文本文件上打印出整个 grep 搜索结果。

这是我想以循环(或嵌套循环)的形式执行的逐步过程

# prepare the test data
names_df1 = data.frame(location = c(1530, 6801, 1530, 6801, 1967),
                       first_name = c("Axel", "Bill", "Carlos", "Flavio", "Jong"),
                       last_name = c("Williams", "Johnson Clarke", "Lopez Gutierrez",  "Mar", "Yoon"), stringsAsFactors = F)

names_df2 = data.frame(location = c(1530, 6801, 1530, 6801, 1967),
                       first_name = c("Axel", "Bill", "Carlos", "Flavio", "Jong"),
                       last_name = c("Williams", "Clarke", "Lopez", "Mar", "Yoon"), stringsAsFactors = F)


# Step 1: paste id and first name. Location ID and First Name are identical in both data frames. I will paste the last name in the second step. 
location_name_df1 = paste(names_df1$location, names_df1$first_name)
location_name_df2 = paste(names_df2$location, names_df2$first_name, names_df2$last_name)


# Step 2: string split the last names in df1. I want a loop to go through each element and subelement of this list. 
last_name_strsplit = strsplit(names_df1$last_name, split = " ")


          # these are what I would be searching. Note that in the loop, I go search through each sub element v of the ith element in the list.
          # paste(location_name_df1[i], last_name_strsplit[[i]][v])
          paste(location_name_df1[1], last_name_strsplit[[1]][1])

          paste(location_name_df1[2], last_name_strsplit[[2]][1])
          paste(location_name_df1[2], last_name_strsplit[[2]][2])

          paste(location_name_df1[3], last_name_strsplit[[3]][1])
          paste(location_name_df1[3], last_name_strsplit[[3]][2])

          paste(location_name_df1[4], last_name_strsplit[[4]][1])

          paste(location_name_df1[5], last_name_strsplit[[5]][1])


    # this is the actual search I would like to do. I paste the location_name_df1 with the last names in last_name_strsplit, going through each element (i), as well as each sub element (v)
    names_df1[grep(paste(location_name_df1[1], last_name_strsplit[[1]][1]),location_name_df2),] # search result successful

    names_df1[grep(paste(location_name_df1[2], last_name_strsplit[[2]][1]),location_name_df2),] # search result NOT successful. Note that this part of the list has two elements. Loop should jump to the second sub element of last_name_strplit
    names_df1[grep(paste(location_name_df1[2], last_name_strsplit[[2]][2]),location_name_df2),] # This search result was successful

    names_df1[grep(paste(location_name_df1[3], last_name_strsplit[[3]][1]),location_name_df2),] # search result successful
    names_df1[grep(paste(location_name_df1[3], last_name_strsplit[[3]][2]),location_name_df2),] # search result NOT successful. Note that this part of the list has two elements. End of sub elements, move on to the next row

    names_df1[grep(paste(location_name_df1[4], last_name_strsplit[[4]][1]),location_name_df2),] # search result successful

    names_df1[grep(paste(location_name_df1[5], last_name_strsplit[[5]][1]),location_name_df2),] # search result successful

我很确定我必须做一个嵌套循环结构,其中我遍历列表的每个元素 (i),然后遍历它的每个子元素 (v)。然而,当我做一个嵌套循环时,往往会发生的是我复制了很多粘贴并且搜索本身出错了。

有人可以给我一些关于如何使用上述步骤创建循环结构的指示吗?我再次使用 R/RStudio 来匹配数据。

谢谢!

【问题讨论】:

  • 对不起!你说的对。我将编辑我的帖子以反映这一点。我正在使用 R。

标签: r loops nested-loops


【解决方案1】:

这是一个更简单的方法。首先,我们对位置和名字进行完全连接,然后我们使用 stringr::str_detect(与 grep 不同,它在字符串 和 模式上进行矢量化)来过滤掉最后一个单姓不是可能的双姓之一:

full = merge(names_df1, names_df2, by = c("location", "first_name"))

library(stringr)
matches = full[str_detect(string = full$last_name.x, pattern = fixed(full$last_name.y)), ]
matches           
#   location first_name     last_name.x last_name.y
# 1     1530       Axel        Williams    Williams
# 2     1530     Carlos Lopez Gutierrez       Lopez
# 3     1967       Jong            Yoon        Yoon
# 4     6801       Bill  Johnson Clarke      Clarke
# 5     6801     Flavio             Mar         Mar

如果你更喜欢dplyr,你可以这样做:

library(dplyr)
full_join(names_df1, names_df2, by = c("location", "first_name")) %>% 
  filter(str_detect(string = last_name.x, pattern = fixed(last_name.y))

【讨论】:

  • 这看起来很有希望!明天我会用我的实际数据集试试这个。您能否详细说明str_detect() 的作用。我看到它将一个姓氏作为参数,然后将第二个姓氏向量作为另一个参数。 fixed() 的 R 文档非常模糊。这是否意味着像“von Gogh”这样的姓氏仍然会与“gogh”这个名字相匹配?如果位置代码和名字相同,但一个姓氏是“van”而另一个姓氏是“evane”怎么办?谢谢!!
  • 我对此进行了测试。我想这正是我想要的。但是,我仍然想知道这样的嵌套循环是否仍然可行,尽管它会效率低下。只是出于好奇。
  • 在str_detect 中使用fixed() 与在grep 中使用参数fixed = TRUE 相同。没有它,将使用正则表达式模式,这意味着您可以选择做各种花哨的东西,. 将匹配任何字符,$ 是字符串的结尾,^ 表示“不是”,具体数量和模式等。使用fixed,只考虑完全匹配,所有花哨的选项都被禁用。它也快得多。
  • 案例是一个单独的问题。按照我的方式,只会考虑完全匹配(包括大小写)。如果你想忽略大小写(这样"gogh" 将匹配"Gogh"),那么你不能使用stringr。您可以使用ignore.case = TRUE 返回grep,但需要额外的矢量化。相反,我建议使用toupper() 或tolower() 将所有内容转换为相同的大小写。
  • 当然循环方法是可能的。编写和调试真的很烦人。你最好使用grepl 而不是grep。并且粘贴在一起似乎完全没有必要。至少帮自己一个忙,从加入位置和名字开始,否则你实际上是在编写一个低效的 merge 特例与字符串匹配。
猜你喜欢
  • 2016-05-03
  • 2020-03-27
  • 1970-01-01
  • 2023-03-02
  • 1970-01-01
  • 2017-04-30
  • 2017-07-04
  • 2021-11-17
  • 1970-01-01
相关资源
最近更新 更多