【发布时间】:2014-02-12 02:56:02
【问题描述】:
我正在编写一段 R 代码并卡住了。
背景(这不是解决问题所必需的):我通过乘以独立的边际分布来计算联合概率。边际概率向量由 ProbGenerationProcess() 迭代生成。在每次迭代时,它都会输出一个向量,例如。
Iteration 1:
Color =
Blue Green
0.2 0.8
Iteration 2:
Material =
Cotton Silk
0.7 0.3
Iteration 3:
Country =
China USA
0.6 0.4
......
期望的结果:我希望得到的联合概率是每个边缘向量中每个元素的乘积。格式应如下所示。
Color Material Country Prob
Blue Cotton China 0.084 (= 0.2*0.7*0.6)
Blue Cotton USA 0.056 (= 0.2*0.7*0.4)
Blue Silk China 0.036 (= 0.2*0.3*0.6)
Blue Silk USA ..
Green Cotton China ..
Green Cotton USA ..
... ... ... ...
我的实现:这是我的代码:
joint.names = NULL # data.from store the marginal value names
joint.probs = NULL # store probabilities.
for (i in iterations) {
marginal = ProbGenerationProcess(VarUniqueToIteration) # output is numeric with names
if ( is.null(joint.names) ) {
# initialize the dataframes
joint.names = names(marginal)
joint.probs = marginal
} else {
# (my hope:) iteratively populate the joint.names and joint.probs
joint.names = expand.grid(joint.names, names(marginal))
expanded.prob = expand.grid(joint.probs, marginal)
joint.probs = expanded.prob$Var1 * expanded.prob$Var2 # Row-by-row multiplication.
}
}
输出: Joint.probs 的结果总是正确的,但是,joint.names 并不能完全按照我想要的方式工作。在前两次迭代之后,一切正常。我得到了:
joint.names =
Var1 Var2
1 Blue Cotton
2 Green Cotton
3 Blue Silk
4 Green Silk
... ...
从第三次迭代开始就有问题了:
joint.names =
Var1.Var1 Var1.Var2 Var1.Var1.1 Var1.Var2.1 Var2
1 Blue Cotton Blue Cotton China
2 Green Cotton Green Cotton China
3 Blue Silk Blue Silk USA
4 Green Silk Green Silk USA
我想我的第一个问题是:这是获得我想要的结果的最有效方法吗?如果是这样,expand.grid() 是我应该使用的函数吗,我应该如何正确初始化它?
感谢任何帮助!
【问题讨论】:
标签: r dataframe iteration distribution probability