【发布时间】:2021-11-25 19:28:16
【问题描述】:
我有两个数据框一个(Lots),其结构如下:
| Lot Group | Lot Number | Booking Stage | Date |
|---|---|---|---|
| 1 | 216000.00 | HPRESM | 2020-08-28 |
| 2 | 890000.01 | PART | 2013-04-17 |
另外一项测量如下:
| Mid | Date | Measurement 1 | Measurement 2 |
|---|---|---|---|
| 1901827 | 2020-08-28 | 44.5 | 23.22 |
| 2981632 | 2013-04-17 | 49.0 | 34.5 |
两个数据框中的日期列具有唯一的日期,并且它们在两个数据框中相同,因为它们具有相同的长度。
我要做的是计算作为连续变量的测量列与 1(好批次)或 2(坏批次,即二分变量)的批次组之间的相关性。测量变量有很多NaNs 超过 50%。我的问题是,当我读到它用于计算这两种类型的变量之间的相关性时,我试图计算 Point-Biserial 相关性,但我得到 nan 的统计数据和 1 的 p 值。
columns = measurement.select_dtypes(exclude = ["object", "datetime"]).columns
for col in columns:
stat, p = ss.pointbiserialr(lots["LosGruppe"], measurement[col])
print(f"Variable: {col}, Correlation: {stat}, P-Value: {p}")
Output:
Variable: Mes 1, Correlation: nan, P-Value: 1.0
Variable: Mes 2, Correlation: nan, P-Value: 1.0
Variable: Mes 3, Correlation: nan, P-Value: 1.0
Variable: Mes 4, Correlation: nan, P-Value: 1.0
Variable: Mes 5, Correlation: nan, P-Value: 1.0
对于此问题的解决方案或原因,您有什么建议以及这些变量之间合适的关联方法是什么?
【问题讨论】:
-
measurement.dropna()[col]可能有帮助吗?
标签: python pandas statistics correlation