【问题标题】:R - Perform function for every two columns in dataR - 对数据中的每两列执行功能
【发布时间】:2018-08-21 17:13:16
【问题描述】:

我有一个包含许多(100+)对坐标的数据框

[lat1] [long2] [lat3] [long4] [..]  [..]
30.12    70.25    32.21    70.25  ..  ..
31.21    71.32    32.32    75.2   ..  ..
32.32    70.25    31.23    75.0   ..  ..

此函数绘制一条连接数据框前两列坐标的线

lines(mapproject(x=data$long2, y=data$lat1), col=3, pch=20, cex=.1)

我需要在每对纬度/经度坐标上执行此功能,以便为每个坐标绘制一条新的/未连接的线

看这个例子 Operate on every two columns in a matrix我想我需要创建一个列表,我该如何为每一列对做呢?

然后,我可以使用这里描述的 lapply https://nicercode.github.io/guides/repeating-things/ - 我应该如何将我的 lines() 调用包装在一个函数中?

完整脚本:

library(maps)
library(mapproj)

data <- read.csv("data.csv")

map('world', proj='orth', fill=TRUE, col="#f2f2f2", border=0, orient=c(90, 0, 0))

lines(mapproject(x=data$long, y=data$lat), col=3, pch=20, cex=.1)

更新问题

在 cmets 的帮助下,我尝试了一种在每两列上执行 sapply 的方法,这似乎可以按预期工作。这有待改进

df <- data.frame(X1 = c(0, 10, 20), 
                   Y2 = c(80, 85, 90), 
                   X3 = c(3, 10, 15), 
                   Y4 = c(93, 100, 105), 
                   X5 = c(16, 20, 35),
                   Y6 = c(100, 105, 130))


map('world', proj='orth', fill=TRUE, col="#f2f2f2", border=0, orient=c(90, 0, 0))

sapply(seq(1,5,by=2),function(i) lines(mapproject(x = (df[,i]), y = (df[,(i+1)])), col = 3))

【问题讨论】:

  • 您可以重塑数据,对数据进行分组,然后在所有组上计算函数
  • 您能否包括您期望的重构数据的样子?

标签: r maps lapply


【解决方案1】:

您的编辑对我有用,但我可能会这样做,因为它可以更好地概括。

library(maps)
library(mapproj)

df <- data.frame(X1 = c(0, 10, 20), 
                 Y2 = c(80, 85, 90), 
                 X3 = c(3, 10, 15), 
                 Y4 = c(93, 100, 105), 
                 X5 = c(16, 20, 35),
                 Y6 = c(100, 105, 130))

draw_lines <- function(df){

#Make sure we get the order right for iterating over length
 x  <- sort(grep("X", names(df), value = TRUE), decreasing = FALSE)
 y  <- sort(grep("Y", names(df), value = TRUE), decreasing = FALSE)

#Check for vector lengths to be the same
stopifnot(length(x) == length(y))

#I dont want to print anything to console
invisible(lapply(1 : length(x), function(j){
         lines(mapproject(x = df[, x[j]], y = df[, y[j]]), col = 3, pch = 20, cex = .1)
         })
         )
}

然后:

map('world', proj='orth', fill=TRUE, col="#f2f2f2", border=0, orient=c(90, 0, 0))
draw_lines(df = df)

希望对你有帮助

【讨论】:

  • 这是一个很好的答案,但我遇到了一个问题,即我有不可用的列名,即 X1、Y2、X3、Y4 ......其中 X1 和 Y2 是匹配的。你觉得这样做怎么样? sapply(seq(1,449,by=2),function(i) lines(mapproject(x = (df[,i]), y = (df[,(i+1)])), col = 3, pch = 20, cex = .1))
  • 应该在这里工作,当您可以保证您的列完全按照您的需要进行排序时。你确定 x = df[, i] 和 y = df[, i+1]。只是问,因为您在问题行中写道(mapproject(x=data$long2, y=data$lat1), col=3, pch=20, cex=.1)。您也许应该编辑您的问题。
【解决方案2】:

回复更新后的问题

对我的原始答案的此更新不依赖于列名的任何内容,但它确实依赖于在 (latitude, longitude) 对中排序的列。由于?mapproject 表示它期望坐标对以相反的顺序,我在函数调用中切换它们。

我也不假设有很多列。

library(dplyr)
library(stringr)
library(maps)
library(mapproj)

lat_lons <- data.frame(lat1 = c(30, 31, 32), 
                       lon1 = c(70, 71, 70), 
                       lat2 = c(32, 32, 31), 
                       lon2 = c(70, 75, 75), 
                       lat3 = c(33, 33, 33),
                       lon3 = c(72, 73, 74))

ints <- seq(1, ncol(lat_lons) - 1)

map_lines <- function(int) {

  lines(mapproject(x = df[, int + 1], y = df[, int]), col = 3, pch = 20, cex = .1)

}

map('world', proj = 'orth', fill = TRUE, col = "#f2f2f2", border = 0, orient = c(90, 0, 0))

sapply(ints, map_lines)

回复原始问题

我认为这是你想要的:

library(dplyr)
library(stringr)
library(maps)
library(mapproj)

lat_lons <- data.frame(lat1 = c(30, 31, 32), 
                       lon1 = c(70, 71, 70), 
                       lat2 = c(32, 32, 31), 
                       lon2 = c(70, 75, 75), 
                       lat3 = c(33, 33, 33),
                       lon3 = c(72, 73, 74))

ints <- names(lat_lons) %>% str_extract("[0-9]") %>% unique()

map_lines <- function(int) {

  df <- select(lat_lons, matches(int))
  lines(mapproject(x = df[, 2], y = df[, 1]), col = 3, pch = 20, cex = .1)

}

map('world', proj = 'orth', fill = TRUE, col = "#f2f2f2", border = 0, orient = c(90, 0, 0))

sapply(ints, map_lines)

首先,从数据框列名中获取唯一的整数集。然后,对于每个整数,应用一个函数来选择与该整数匹配的 2 列并调用 lines。

函数map_lines 不返回任何内容,而是将线条添加到活动图中作为副作用。不是一种理想的做事方式(我宁愿使用ggplot2 并在制作情节之前将这些部分组装成一个列表),但它确实有效。

【讨论】:

  • 这也是一个很好的答案,但我遇到了一个问题,即我有不可用的列名,即 X1、Y2、X3、Y4 ......其中 X1 和 Y2 是匹配的。你觉得这样做怎么样? sapply(seq(1,449,by=2),function(i) lines(mapproject(x = (df[,i]), y = (df[,(i+1)])), col = 3, pch = 20, cex = .1))
  • 当然!我更新了我的答案以删除关于列名包含什么的假设,但仍然没有假设有多少列。
猜你喜欢
  • 1970-01-01
  • 2021-05-14
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-05-17
  • 2022-07-06
  • 1970-01-01
  • 2021-07-27
相关资源
最近更新 更多