【问题标题】:Creating a new Data.Frame from variable values从变量值创建一个新的 Data.Frame
【发布时间】:2020-04-12 04:22:18
【问题描述】:

我目前正在处理一项需要我从 sql 数据库中查询股票列表的任务。 问题在于它是一个列表,其中每个日期交易的股票数量为 1:n。我想计算给定日期每只股票在投资组合中的份额(参见示例)并将其传递给新的数据框。换句话说,日期 x 出现了 2 次(股票 A 一次,股票 B 一次),然后将日期 x 与新值一起出现一次。

'data.frame':   1010 obs. of  5 variables:
 $ ID         : int  1 2 3 4 5 6 7 8 9 10 ...
 $ Date      : Date, format: "2019-11-22" "2019-11-21" "2019-11-20" "2019-11-19" ...
 $ Close: num  52 51 50.1 50.2 50.2 ...
 $ Volume     : num  5415 6196 3800 4784 6189 ...
 $ Stock_ID  : Factor w/ 2 levels "1","2": 1 1 1 1 1 1 1 1 1 1 ...

RawInput<-data.frame(Date=c("2017-22-11","2017-22-12","2017-22-13","2017-22-11","2017-22-12","2017-22-13","2017-22-11"), Close=c(50,55,56,10,11,12,200),Volume=c(100,110,150,60,70,80,30),Stock_ID=c(1,1,1,2,2,2,3))
RawInput$Stock_ID<-as.factor(RawInput$Stock_ID)

*在本例中不能将日期转换为日期变量

我想要一个新的数据框来生成每天交易的价值、每只股票的权重以及每天的每日回报,同时保持股票数量可变。

我希望我正确翻译了这个问题,以便我能得到帮助。

谢谢!

【问题讨论】:

  • 通过dput(head(df, 20)) 提供您的数据片段将使人们更容易帮助您。其中df 是您的data.frame。
  • Date 应该是什么2017-22-13?
  • &gt; dput(head(RawInput)) structure(list(ID = 1:6, Date = structure(c(18222, 18221, 18220, 18219, 18218, 18215), class = "Date"), Close = c(52.03, 51.03, 50.1, 50.16, 50.23, 50.68), Volume = c(5415, 6196, 3800, 4784, 6189, 7753), Stock_ID = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("1", "2"), class = "factor")), row.names = c(NA, 6L), class = "data.frame")
  • -日期当然是2017-12-23 或类似的东西

标签: r dataframe


【解决方案1】:

我认为最简单的方法是使用dplyr 包。您可能需要阅读一些文档,但 mutate 和 group_by 函数可能能够满足您的需求。此功能将允许您通过添加新列或更改现有数据来修改当前数据框。

让我们从一个可重现的数据集开始

RawInput<-data.frame(Date=c("2017-22-11","2017-22-12","2017-22-13","2017-22-11","2017-22-12","2017-22-13","2017-22-11"),
                    Close=c(50,55,56,10,11,12,200),
                    Volume=c(100,110,150,60,70,80,30),
                    Stock_ID=c(1,1,1,2,2,2,3))
RawInput$Stock_ID<-as.factor(RawInput$Stock_ID)
library(magrittr)
library(dplyr)
dat2 <- RawInput %>%
    group_by(Date, Stock_ID) %>%  #this example only has one stock type but i imagine you want to group by stock
    mutate(CloseMean=mean(Close),
           CloseSum=sum(Close),
           VolumeMean=mean(Volume),
           VolumeSum=sum(Volume)) #what ever computation you need to do with
                                  #multiple stock values for a given date goes here
dat2 %>% select(Stock_ID, Date, CloseMean, CloseSum, VolumeMean,VolumeSum) %>% distinct() #dat2 will still be the same size as dat, thus use the distinct() function to reduce it to unique values
# A tibble: 7 x 6
# Groups:   Date, Stock_ID [7]
  Stock_ID Date       CloseMean CloseSum VolumeMean VolumeSum
  <fct>    <fct>          <dbl>    <dbl>      <dbl>     <dbl>
1 1        2017-22-11        50       50        100       100
2 1        2017-22-12        55       55        110       110
3 1        2017-22-13        56       56        150       150
4 2        2017-22-11        10       10         60        60
5 2        2017-22-12        11       11         70        70
6 2        2017-22-13        12       12         80        80
7 3        2017-22-11       200      200         30        30

您提供的这个数据集实际上只有一个唯一的 Stock_ID 和 Date 组合,因此实际上没有对数据进行任何处理。但是,如果您在必要时删除 Stock_ID,您可以看到此功能的工作原理

dat2 <- RawInput %>%
        group_by(Date) %>%  
        mutate(CloseMean=mean(Close),
               CloseSum=sum(Close),
               VolumeMean=mean(Volume),
               VolumeSum=sum(Volume)) 
dat2 %>% select(Date, CloseMean, CloseSum, VolumeMean,VolumeSum) %>% distinct() 

# A tibble: 3 x 5
# Groups:   Date [3]
  Date       CloseMean CloseSum VolumeMean VolumeSum
  <fct>          <dbl>    <dbl>      <dbl>     <dbl>
1 2017-22-11      86.7      260       63.3       190
2 2017-22-12      33         66       90         180
3 2017-22-13      34         68      115         230

阅读您的第一个回复后,您必须具体说明您是如何计算重量的。还要定义你的最终结果。 我将假设重量只是总成本的百分比。最终结果是每个日期显示每只股票的重量。换句话说,日期和股票 ID 的矩阵

library(tidyr)
RawInput %>%
     group_by(Date) %>%
     mutate(weight=Close/sum(Close)) %>%
     select(Date, weight, Stock_ID) %>%
     spread(key = "Stock_ID", value = "weight", fill = 0)

# A tibble: 3 x 4
# Groups:   Date [3]
  Date         `1`    `2`   `3`
  <fct>      <dbl>  <dbl> <dbl>
1 2017-22-11 0.192 0.0385 0.769
2 2017-22-12 0.833 0.167  0    
3 2017-22-13 0.824 0.176  0   

【讨论】:

  • 实际上你最后呈现的结果是最终应该出现的结果。我想采取的障碍仍然是我想分配唯一日期(如在您的最后一张表中)一个投资组合数据框,其中包含投资组合中每只股票的权重(股票 3 的前 2 天权重为 0 )。这些权重可以转化为给定日期的加权投资组合价值。因此,如果数据库中有一个很长的列表,其中 1 天有 1 只股票,第 2 天有 3 只股票,那么我希望能够过滤该信息并保持其灵活性。非常感谢!
  • 听起来如果您正在寻找自定义摘要工具,您将需要编写一个函数。祝你工作顺利
  • 非常感谢您回答了这个问题
猜你喜欢
  • 1970-01-01
  • 2013-10-14
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-09-20
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多