【问题标题】:How to run a for loop (or other approach) on two column range如何在两列范围上运行 for 循环(或其他方法)
【发布时间】:2019-02-01 19:54:07
【问题描述】:

我有一个用于网络抓取的 for 循环。例如,假设它正在收集历史股票数据。

start <- 1533103200
end <- 1549004400

company <- c("fb","amzn","f")

for (i in company){
    print(paste('https://finance.yahoo.com/quote/',i, '/history?period1=',start,'&period2=',maxDate,'&interval=1d&filter=history&frequency=1d',sep=""))
}

开始和结束是日期代码。现在,我有一个包含开始和结束日期代码(100 天间隔)的 data.frame,我还想进入打印链接列表,这意味着我想要以下 data.frame 的 3 x nrow 而不是三个链接。在此示例中,它将是 6 个链接...

start <- c(1533193200,1541833200)
end <- c(1541746800,1549004400)
dates <- as.data.frame(cbind(start,end))

该列表是动态的且很长,因此我可能不得不将 for 循环嵌入到另一个 for 循环中,但我没有太多经验为此目的使用两个变量。任何帮助都会很棒!

预期的结果是......

[1] "https://finance.yahoo.com/quote/fb/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/amzn/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/f/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/fb/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/amzn/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/f/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"

...而不是第一个循环的结果...

[1] "https://finance.yahoo.com/quote/fb/history?period1=1533103200&period2=1548918000&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/amzn/history?period1=1533103200&period2=1548918000&interval=1d&filter=history&frequency=1d"
[1] "https://finance.yahoo.com/quote/f/history?period1=1533103200&period2=1548918000&interval=1d&filter=history&frequency=1d"

【问题讨论】:

    标签: r for-loop matrix


    【解决方案1】:

    我稍微简化了您的 data.frame 构造:

    df <- data.frame(
      start = c(1533193200, 1541833200),
      end = c(1541746800, 1549004400)
    )
    

    然后我会在 data.frame 中为每个公司分配新列:

    companies <- c("fb", "amzn", "f")
    df[, companies] <- ""
    

    现在您可以遍历新列并用链接填充它们:

    for (i in companies) {
      df[, i] <- paste0(
        'https://finance.yahoo.com/quote/',
        i, '/history?period1=',
        df$start,
        '&period2=',
        df$maxDate,
        '&interval=1d&filter=history&frequency=1d')
    }
    

    你会得到一个很好的干净data.frame,每个公司的链接都在一个单独的列中:

    > df
           start        end
    1 1533193200 1541746800
    2 1541833200 1549004400
    
    
    fb
    1 https://finance.yahoo.com/quote/fb/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d
    2 https://finance.yahoo.com/quote/fb/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d
                                                                                                                  amzn
    1 https://finance.yahoo.com/quote/amzn/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d
    2 https://finance.yahoo.com/quote/amzn/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d
                                                                                                                  f
    1 https://finance.yahoo.com/quote/f/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d
    2 https://finance.yahoo.com/quote/f/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d
    

    如果您更喜欢包含链接的列,而其他列用作有关链接的元信息,您可以“整理”一下:

    df_tidy <- tidyr::gather(df, company, url, -start, -end)
    
    > df_tidy$url
    [1] "https://finance.yahoo.com/quote/fb/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d"  
    [2] "https://finance.yahoo.com/quote/fb/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d"  
    [3] "https://finance.yahoo.com/quote/amzn/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d"
    [4] "https://finance.yahoo.com/quote/amzn/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d"
    [5] "https://finance.yahoo.com/quote/f/history?period1=1533193200&period2=&interval=1d&filter=history&frequency=1d"   
    [6] "https://finance.yahoo.com/quote/f/history?period1=1541833200&period2=&interval=1d&filter=history&frequency=1d"
    

    【讨论】:

    • 有趣。我喜欢这种方法,因为它比循环内的循环更易于管理。谢谢!
    【解决方案2】:

    您需要遍历公司和日期。

    start <- c(1533193200,1541833200)
    end <- c(1541746800,1549004400)
    dates <- as.data.frame(cbind(start,end))
    
    companies <- c("fb","amzn","f")
    
    string <- 'https://finance.yahoo.com/quote/%s/history?period1=%s&period2=%s&interval=1d&filter=history&frequency=1d'
    
    for (company in companies) {
      for (date in 1:nrow(dates)) {
        date <- dates[date, ]
        print(sprintf(string, company, date["start"], date["end"]))
      }
    }
    
    [1] "https://finance.yahoo.com/quote/fb/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
    [1] "https://finance.yahoo.com/quote/fb/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"
    [1] "https://finance.yahoo.com/quote/amzn/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
    [1] "https://finance.yahoo.com/quote/amzn/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"
    [1] "https://finance.yahoo.com/quote/f/history?period1=1533193200&period2=1541746800&interval=1d&filter=history&frequency=1d"
    [1] "https://finance.yahoo.com/quote/f/history?period1=1541833200&period2=1549004400&interval=1d&filter=history&frequency=1d"
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-01-05
      • 1970-01-01
      • 2018-05-22
      • 1970-01-01
      • 2019-01-12
      • 1970-01-01
      • 1970-01-01
      • 2011-06-23
      相关资源
      最近更新 更多