【问题标题】:After fitting the cumulative distribution in R creating the normal distribution from fitted parameters在 R 中拟合累积分布后,根据拟合参数创建正态分布
【发布时间】:2017-11-02 20:02:21
【问题描述】:

使用 Gompertz 函数成功拟合累积数据后,我需要从拟合函数创建正态分布。

这是目前为止的代码:

      df <- data.frame(x = c(0.01,0.011482,0.013183,0.015136,0.017378,0.019953,0.022909,0.026303,0.0302,0.034674,0.039811,0.045709,0.052481,0.060256,0.069183,0.079433,0.091201,0.104713,0.120226,0.138038,0.158489,0.18197,0.20893,0.239883,0.275423,0.316228,0.363078,0.416869,0.47863,0.549541,0.630957,0.724436,0.831764,0.954993,1.096478,1.258925,1.44544,1.659587,1.905461,2.187762,2.511886,2.884031,3.311311,3.801894,4.365158,5.011872,5.754399,6.606934,7.585776,8.709636,10,11.481536,13.182567,15.135612,17.378008,19.952623,22.908677,26.30268,30.199517,34.673685,39.810717,45.708819,52.480746,60.255959,69.183097,79.432823,91.201084,104.712855,120.226443,138.038426,158.489319,181.970086,208.929613,239.883292,275.42287,316.227766,363.078055,416.869383,478.630092,549.540874,630.957344,724.43596,831.763771,954.992586,1096.478196),
                 y = c(0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.00044816,0.00127554,0.00221488,0.00324858,0.00438312,0.00559138,0.00686054,0.00817179,0.00950625,0.01085188,0.0122145,0.01362578,0.01514366,0.01684314,0.01880564,0.02109756,0.0237676,0.02683182,0.03030649,0.0342276,0.03874555,0.04418374,0.05119304,0.06076553,0.07437854,0.09380666,0.12115065,0.15836926,0.20712933,0.26822017,0.34131335,0.42465413,0.51503564,0.60810697,0.69886817,0.78237651,0.85461023,0.91287236,0.95616228,0.98569093,0.99869001,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999))

library(drc)
fm <- drm(y ~ x, data = df, fct = G.3())

options(scipen = 10) #to avoid scientific notation in x axis

plot(df$x, predict(fm),type = "l", log = "x",col = "blue",
           main = "Cumulative function distribution",xlab = "x", ylab = "y")

points(df,col = "red")

legend("topleft", inset = .05,legend = c("exp","fit")
       ,lty = c(NA,1), col = c("red", "blue"), pch = c(1,NA), lwd=1, bty = "n")


summary(fm)

这是下面的情节:

我现在的想法是以某种方式将这种累积拟合转换为正态分布。有什么想法我该怎么做?

【问题讨论】:

    标签: r math statistics curve-fitting normal-distribution


    【解决方案1】:

    我在想cumdiff(因为没有更好的术语)。 link 帮了大忙。

    编辑

     plot(df$x[-1], Mod(df$y[-length(df$y)]-df$y[-1]), log = "x", type = "b",  
          main = "Normal distribution for original data", 
          xlab = "x", ylab = "y")
    

    屈服:

    添加

    为了从fitted函数中得到高斯:

    df$y_pred<-predict(fm)
    plot(df$x[-1], Mod(df$y_pred[-length(df$y_pred)]-df$y_pred[-1]), log = "x", 
         type = "b", main="Normal distribution for fitted function", 
         xlab = "x", lab = "y")
    

    屈服:

    【讨论】:

    • 好的,这就是你从我的初始数据集而不是从拟合创建正常图的方式。我现在将深入研究,看看如何拟合 eq。
    • 非常感谢!但是有没有办法在 x 轴而不是索引上绘制实际的 x 值?我试过了,但我最终遇到了 Mod(df$y[-length(df$y)]-df$y[-1])df$x 之间长度不同的问题 ...
    • 是的,你是对的......我也在考虑这个问题。试试str(fm) 看看你能不能得到一些信息。毕竟我对drc 包不是很熟悉。现在我不能深入研究,但我保证我会尽快回复你。
    • 在哪里可以管理绘制 x 值而不是索引值?我尝试了不同的事情,但没有成功。如果没有,我会发布这个问题......
    • 我想我做到了:plot(df$x[-1], Mod(df$y[-length(df$y)]-df$y[-1]), log="x", type="b")。这给了我想要的输出。
    【解决方案2】:

    虽然您的初衷可能是非参数的,但我建议使用参数估计方法:矩量法,它广泛用于此类问题,因为您需要拟合一定的参数分布(正态分布)。思路很简单,从拟合的累积分布函数中,可以计算出均值(我的代码中E1)和方差(我的代码中SD的平方),然后问题就解决了,因为正态分布可以完全由均值和方差决定。

    df <- data.frame(x=c(0.01,0.011482,0.013183,0.015136,0.017378,0.019953,0.022909,0.026303,0.0302,0.034674,0.039811,0.045709,0.052481,0.060256,0.069183,0.079433,0.091201,0.104713,0.120226,0.138038,0.158489,0.18197,0.20893,0.239883,0.275423,0.316228,0.363078,0.416869,0.47863,0.549541,0.630957,0.724436,0.831764,0.954993,1.096478,1.258925,1.44544,1.659587,1.905461,2.187762,2.511886,2.884031,3.311311,3.801894,4.365158,5.011872,5.754399,6.606934,7.585776,8.709636,10,11.481536,13.182567,15.135612,17.378008,19.952623,22.908677,26.30268,30.199517,34.673685,39.810717,45.708819,52.480746,60.255959,69.183097,79.432823,91.201084,104.712855,120.226443,138.038426,158.489319,181.970086,208.929613,239.883292,275.42287,316.227766,363.078055,416.869383,478.630092,549.540874,630.957344,724.43596,831.763771,954.992586,1096.478196),
                     y=c(0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.00044816,0.00127554,0.00221488,0.00324858,0.00438312,0.00559138,0.00686054,0.00817179,0.00950625,0.01085188,0.0122145,0.01362578,0.01514366,0.01684314,0.01880564,0.02109756,0.0237676,0.02683182,0.03030649,0.0342276,0.03874555,0.04418374,0.05119304,0.06076553,0.07437854,0.09380666,0.12115065,0.15836926,0.20712933,0.26822017,0.34131335,0.42465413,0.51503564,0.60810697,0.69886817,0.78237651,0.85461023,0.91287236,0.95616228,0.98569093,0.99869001,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999))
    
    library(drc)
    fm <- drm(y ~ x, data = df, fct = G.3())
    
    options(scipen = 10) #to avoid scientific notation in x axis
    
    plot(df$x, predict(fm),type="l", log = "x",col="blue", main="Cumulative distribution function",xlab="x", ylab="y")
    
    points(df,col="red")
    
    E1 <- sum((df$x[-1] + df$x[-length(df$x)]) / 2 * diff(predict(fm)))
    E2 <- sum((df$x[-1] + df$x[-length(df$x)]) ^ 2 / 4 * diff(predict(fm)))
    SD <- sqrt(E2 - E1 ^ 2)
    points(df$x, pnorm((df$x - E1) / SD), col = "green")
    
    legend("topleft", inset = .05,legend= c("exp","fit","method of moment")
           ,lty = c(NA,1), col = c("red", "blue", "green"), pch = c(1,NA), lwd=1, bty="n")
    
    
    summary(fm)
    

    以及估计结果:

    ## > E1 (mean of fitted normal distribution)
    ## [1] 65.78474
    ## > E2 (second moment of fitted normal distribution)
    ##[1] 5792.767
    ## > SD (standard deviation of fitted normal distribution)
    ## [1] 38.27707
    ## > SD ^ 2 (variance of fitted normal distribution)
    ## [1] 1465.134
    

    编辑:更新了从drc 拟合的 cdf 计算矩的方法。下面定义的函数moment 使用连续 r.v 的矩公式计算矩估计。 E(X ^ k) = k * \int x ^ {k - 1} (1 - cdf(x)) dx。这些是我可以从拟合的 cdf 中得到的最佳估计。当x 接近于零时,拟合不是很好,因为我在 cmets 中讨论过原始数据集中的原因。

    df <- data.frame(x=c(0.01,0.011482,0.013183,0.015136,0.017378,0.019953,0.022909,0.026303,0.0302,0.034674,0.039811,0.045709,0.052481,0.060256,0.069183,0.079433,0.091201,0.104713,0.120226,0.138038,0.158489,0.18197,0.20893,0.239883,0.275423,0.316228,0.363078,0.416869,0.47863,0.549541,0.630957,0.724436,0.831764,0.954993,1.096478,1.258925,1.44544,1.659587,1.905461,2.187762,2.511886,2.884031,3.311311,3.801894,4.365158,5.011872,5.754399,6.606934,7.585776,8.709636,10,11.481536,13.182567,15.135612,17.378008,19.952623,22.908677,26.30268,30.199517,34.673685,39.810717,45.708819,52.480746,60.255959,69.183097,79.432823,91.201084,104.712855,120.226443,138.038426,158.489319,181.970086,208.929613,239.883292,275.42287,316.227766,363.078055,416.869383,478.630092,549.540874,630.957344,724.43596,831.763771,954.992586,1096.478196),
                     y=c(0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.00044816,0.00127554,0.00221488,0.00324858,0.00438312,0.00559138,0.00686054,0.00817179,0.00950625,0.01085188,0.0122145,0.01362578,0.01514366,0.01684314,0.01880564,0.02109756,0.0237676,0.02683182,0.03030649,0.0342276,0.03874555,0.04418374,0.05119304,0.06076553,0.07437854,0.09380666,0.12115065,0.15836926,0.20712933,0.26822017,0.34131335,0.42465413,0.51503564,0.60810697,0.69886817,0.78237651,0.85461023,0.91287236,0.95616228,0.98569093,0.99869001,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999,0.99999999))
    
    library(drc)
    fm <- drm(y ~ x, data = df, fct = G.3())
    
    moment <- function(k){
        f <- function(x){
            x ^ (k - 1) * pmax(0, 1 - predict(fm, data.frame(x = x)))
        }
        k * integrate(f, lower = min(df$x), upper = max(df$x))$value
    }
    
    E1 <- moment(1)
    E2 <- moment(2)
    SD <- sqrt(E2 - E1 ^ 2)
    

    【讨论】:

    • 感谢您提出这个想法。我将深入研究这一点,看看我该如何处理瞬间的方法。我有点担心它不太适合图表的第一部分(直到 x = 80)。你知道为什么吗?
    • 所以,我尝试从您的时刻绘制正态分布,而起点的不合适会导致看起来很奇怪的正态分布(因为它在第一阶段没有达到 0)。这是我使用的代码:y2 &lt;- dnorm(df$x,mean = E1,sd = SD) plot(df$x,y2,type = "b") 这是plot
    • @numb 我发现问题是因为 SD 估计过大。这是因为原始的df$x 在 x 小的时候是密集的,而在 x 很大的时候非常稀疏,这导致了这个问题。我正在寻找获得更好估计的方法。
    • @numb 我估计参数E1E2SD 是我能从拟合模型中做的最好的。你提到的问题有所缓解,但仍然存在。看来如果你用drs拟合你的数据,用正态分布重新拟合拟合的数据,这样的问题有点难免,除非你做一些手动调整。
    • @numb 好的,经过一番详细的调查,我想我找到了问题的根源,好像它在你的原始数据中。让我解释一下,我们都知道正态分布是对称的,你的原始数据的中心(y = 0.5)有点靠近x = 60,你可以发现x130附近,y = 0.95 ,其对称点是x 围绕-10,其中y = 0.05。以上只是一些粗略的估计,但它表明您的原始数据在远离中心的点上不是那么对称,这会导致问题。
    猜你喜欢
    • 2013-09-21
    • 1970-01-01
    • 2014-08-31
    • 2017-02-19
    • 2014-07-01
    • 2019-09-06
    • 1970-01-01
    • 2013-03-06
    • 2012-04-29
    相关资源
    最近更新 更多