【问题标题】:Correlation between dichotomous variable and continuous variable二分变量与连续变量的相关性
【发布时间】:2021-11-25 19:28:16
【问题描述】:

我有两个数据框一个(Lots),其结构如下:

Lot Group Lot Number Booking Stage Date
1 216000.00 HPRESM 2020-08-28
2 890000.01 PART 2013-04-17

另外一项测量如下:

Mid Date Measurement 1 Measurement 2
1901827 2020-08-28 44.5 23.22
2981632 2013-04-17 49.0 34.5

两个数据框中的日期列具有唯一的日期,并且它们在两个数据框中相同,因为它们具有相同的长度。

我要做的是计算作为连续变量的测量列与 1(好批次)或 2(坏批次,即二分变量)的批次组之间的相关性。测量变量有很多NaNs 超过 50%。我的问题是,当我读到它用于计算这两种类型的变量之间的相关性时,我试图计算 Point-Biserial 相关性,但我得到 nan 的统计数据和 1 的 p 值。

columns = measurement.select_dtypes(exclude = ["object", "datetime"]).columns
for col in columns:
    stat, p  = ss.pointbiserialr(lots["LosGruppe"], measurement[col])
    print(f"Variable: {col}, Correlation: {stat}, P-Value: {p}")

Output:
Variable: Mes 1, Correlation: nan, P-Value: 1.0
Variable: Mes 2, Correlation: nan, P-Value: 1.0
Variable: Mes 3, Correlation: nan, P-Value: 1.0
Variable: Mes 4, Correlation: nan, P-Value: 1.0
Variable: Mes 5, Correlation: nan, P-Value: 1.0

对于此问题的解决方案或原因,您有什么建议以及这些变量之间合适的关联方法是什么?

【问题讨论】:

  • measurement.dropna()[col] 可能有帮助吗?

标签: python pandas statistics correlation


【解决方案1】:

点双列相关是一种很好的方法,但您遇到的问题是缺失值。你需要先用dropna()删除这些

【讨论】:

    猜你喜欢
    • 2020-10-24
    • 2018-01-17
    • 2017-06-28
    • 1970-01-01
    • 2021-12-30
    • 1970-01-01
    • 2017-11-25
    • 2016-11-25
    • 1970-01-01
    相关资源
    最近更新 更多