【发布时间】:2018-05-09 16:51:55
【问题描述】:
我有两个看起来像这样的数据帧(虽然第一个数据帧超过 9000 万行,第二个数据帧超过 1400 万行)另外第二个数据帧是随机排序的
df1 <- data.frame(
datalist = c("wiki/anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/individualism to complete wiki/collectivism",
"strains of anarchism have often been divided into the categories of wiki/social_anarchism and wiki/individualist_anarchism or similar dual classifications",
"the word is composed from the word wiki/anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e",
"anarchy from anarchos meaning one without rulers from the wiki/privative prefix wiki/privative_alpha an- i.e",
"authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/infinitive suffix -izein",
"the first known use of this word was in 1539"),
words = c("anarchist_schools_of_thought individualism collectivism", "social_anarchism individualist_anarchism",
"anarchy -ism", "privative privative_alpha", "infinitive", ""),
stringsAsFactors=FALSE)
df2 <- data.frame(
vocabword = c("anarchist_schools_of_thought", "individualism","collectivism" , "1965-66_nhl_season_by_team","social_anarchism","individualist_anarchism",
"anarchy","-ism","privative","privative_alpha", "1310_the_ticket", "infinitive"),
token = c("Anarchist_schools_of_thought" ,"Individualism", "Collectivism", "1965-66_NHL_season_by_team", "Social_anarchism", "Individualist_anarchism" ,"Anarchy",
"-ism", "Privative" ,"Alpha_privative", "KTCK_(AM)" ,"Infinitive"),
stringsAsFactors = F)
我能够将短语“wiki/”之后的所有单词提取到另一列中。这些单词需要替换为与第二个数据框中的 vocabword 匹配的标记列。因此,例如,我会查看第一个数据帧第一行中 wiki/ 之后的作品“anarchist_schools_of_thought”,然后在第二个数据帧中的词汇单词下找到术语“anarchist_schools_of_thought”,我想用相应的替换它令牌是“Anarchist_schools_of_thought”。
所以它最终应该是这样的:
1 wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism
2 strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications
3 the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e
4 anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative an- i.e
5 authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein
6 the first known use of this word was in 1539
我意识到其中很多只是将单词的第一个字母大写,但其中一些明显不同。我可以做一个 for 循环,但我认为这会花费太多时间,我更喜欢使用 data.table 方式或可能是 stringi 或 stringr 方式。而且我通常只会进行合并,但是由于需要在一行中替换多个单词,这会使事情变得复杂。
提前感谢您的帮助。
【问题讨论】:
-
这与您昨天提出的问题有何不同? stackoverflow.com/q/50241313/5325862
-
我需要替换文本。我想如果我把一些文字分开的话我可以弄明白,但我一直在努力,但一无所获。
-
你能把上一篇文章中的代码加上你从那以后做了什么吗?这样我们就知道您在哪里停下来以及下一步需要什么
-
我用这个:data$words = trimws(gsub("wiki/(\\S+)|(?:(?!wiki/\\S).)+", " \\1 ", data$datalist, perl=TRUE)) 就是这样。就像我说的,我可以做一个 for 循环,但它会非常慢。从昨天开始,我真的没有任何进展
-
这很好,但该代码对于包含在问题正文中很重要。循环肯定会比向量操作慢,所以我认为上一篇文章让你走在了正确的轨道上