【问题标题】:How can I split up a string based on upper case and lower case in R?如何在R中根据大写和小写拆分字符串?
【发布时间】:2022-06-28 22:29:35
【问题描述】:

我有一列名字,其中姓氏都是大写的,名字都是小写的,除了第一个字母。我怎么能把这个分开? 示例:拜登·乔

names <- c("BIDEN Joe", "DE WEERDT Jan", "SCHEPERS Caro")

我想要实现的结果是创建向量/列,其中包含大写字母的单词,因此它变成:

surnames <- c("BIDEN", "DE WEERDT", "SCHEPERS")

还有其他名字:

first_names <- c("Joe", "Jan", "Caro")

提前致谢

【问题讨论】:

  • 如果您向reproducible example 提供可用于测试和验证可能解决方案的示例输入和所需输出,则更容易为您提供帮助。很难从一个例子中推断出来。姓氏或名字中是否有额外的空格?
  • 好的,谢谢您的提示。我在问题中添加了一些额外的示例。
  • 我对由空格隔开的两部分组成的姓氏特别困难。

标签: r string


【解决方案1】:

试试这个:

names <- c("BIDEN Joe", "DE WEERDT Jan", "SCHEPERS Caro")

# Remove capitals followed by a space 
first_names  <- gsub("^[A-Z].+ ", "", names) 
#  "Joe"  "Jan"  "Caro"

# Replace a space followed by a capital followed by a lower case letter
last_names  <- gsub(" [A-Z][a-z].+$", "", names) 
# "BIDEN"     "DE WEERDT" "SCHEPERS"

我也不会调用向量 names,因为它是 base 函数的名称。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2011-03-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-06-05
    相关资源
    最近更新 更多