【问题标题】:Compare two data frames based on common columns基于公共列比较两个数据框
【发布时间】:2017-10-13 09:10:47
【问题描述】:

我有两个 csv 文件:

文件1:

SN  CY  Year    Month   Day Hour    Lat Lon
196101  1   1961    1   14  12  8.3 134.7
196101  1   1961    1   14  18  8.8 133.4
196101  1   1961    1   15  0   9.1 132.5
196101  1   1961    1   15  6   9.3 132.2
196101  1   1961    1   15  12  9.5 132
196101  1   1961    1   15  18  9.9 131.8

文件2:

Year    Month Day RR Hour Lat  Lon
1961    1   14  0   0   14.0917 121.055
1961    1   14  0   6   14.0917 121.055
1961    1   14  0   12  14.0917 121.055
1961    1   14  0   18  14.0917 121.055
1961    1   15  0   0   14.0917 121.055
1961    1   15  0   6   14.0917 121.055

我想在 file2 中添加另一列,如果 file2 中的行存在于 file1 中,只要它们具有相同的年、月、日和小时,则输入“TRUE”,否则为“FALSE”。然后保存为 csv 文件。

想要的输出:

Year    Month Day RR Hour Lat  Lon      com
1961    1   14  0   0   14.0917 121.055 FALSE
1961    1   14  0   6   14.0917 121.055 FALSE
1961    1   14  0   12  14.0917 121.055 TRUE
1961    1   14  0   18  14.0917 121.055 TRUE
1961    1   15  0   0   14.0917 121.055 TRUE
1961    1   15  0   6   14.0917 121.055 TRUE

这是我的脚本:

jtwc <- read.csv("file1.csv",header=T,sep=",")
stn <- read.csv("file2.csv",header=T,sep=",")

if ((jtwc$Year == "stn$YY") & (jtwc$Month == "stn$MM") & (jtwc$Day == "stn$DD") &(jtwc$Hour == "stn$HH")){
stn$com <- "TRUE"
} else {
stn$com <- "FALSE"
}
write.csv(stn,file="test.csv",row.names=T)

这给出了一个错误:

In if ((jtwc$Year == "stn$YY") & (jtwc$Month == "stn$MM") & (jtwc$Day ==  :the condition has length > 1 and only the first element will be used

【问题讨论】:

  • 做一个可重现的例子。比如发布结果 head(dput(YOURDATA))

标签: r


【解决方案1】:

你也可以使用 dplyr/tidyverse:

library(tidyverse)
d2 %>% 
  left_join(select(d1, Year, Month, Day, Hour, Com=Lon)) %>% 
  mutate(Com=ifelse(is.na(Com), FALSE, TRUE))

Joining, by = c("Year", "Month", "Day", "Hour")
  Year Month Day RR Hour     Lat     Lon   Com
1 1961     1  14  0    0 14.0917 121.055 FALSE
2 1961     1  14  0    6 14.0917 121.055 FALSE
3 1961     1  14  0   12 14.0917 121.055  TRUE
4 1961     1  14  0   18 14.0917 121.055  TRUE
5 1961     1  15  0    0 14.0917 121.055  TRUE
6 1961     1  15  0    6 14.0917 121.055  TRUE  

【讨论】:

    【解决方案2】:

    使用data.table 的快速而肮脏的解决方案:

    1. 使用fread 读入文件。
    2. 从file1 中提取想要的列(因为您只对file2 感兴趣)
    3. 使用merge合并文件
    4. 如果没有来自file1 的匹配项,则添加FALSE

    代码:

    library(data.table)
    result <- merge(fread("file2.csv"),
                    fread("file1.csv")[, .(Year, Month, Day, Hour, com = TRUE)], 
                    all.x = TRUE)[is.na(com), com := FALSE]
    
    result
       Year Month Day Hour RR     Lat     Lon   com
    1: 1961     1  14    0  0 14.0917 121.055 FALSE
    2: 1961     1  14    6  0 14.0917 121.055 FALSE
    3: 1961     1  14   12  0 14.0917 121.055  TRUE
    4: 1961     1  14   18  0 14.0917 121.055  TRUE
    5: 1961     1  15    0  0 14.0917 121.055  TRUE
    6: 1961     1  15    6  0 14.0917 121.055  TRUE
    

    【讨论】:

    • 非常感谢!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-01-16
    • 2013-05-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多