【问题标题】:Pandas groupby output is not displaying null valuesPandas groupby 输出不显示空值
【发布时间】:2019-02-11 03:35:17
【问题描述】:

我正在尝试根据两列绘制出出现的值。效果很好,感谢 post 中的 Marcus。但是,我还希望它为没有计数的事件显示 0(其中评级字段为空)。它目前忽略空值。

当前输出为:

如您所见,Critical 没有出现,因此它们没有显示。如果数据框中没有出现这些环境/评级,我需要它显示 0。

我想要的输出是:

基本上,我希望评级(例如关键和其他 P3)始终显示,这样即使没有关键或其他条目,它也会在该环境中显示为 0。

这是当前代码:
csvfile = pd.read_csv("rawstats.csv", encoding = "ISO-8859-1", usecols=['Environment/s Affected', 'Rating'])
df = pd.DataFrame(csvfile)
df.groupby(['Environment/s Affected', (df['Rating'].isin(['1', '2']))]).size().rename(index={True: 'Critical', False: 'Others P3+'}, level=1).to_csv('summary.csv')

样本数据:
Rating,Environment/s Affected 3,Env1 3,Env1 3,Env1 3,Env2 3,Env2 3,Env2 3,Env2 3,Env3 3,Env3 3,Env3 3,Env3 3,Env3 3,Env4 3,Env4 3,Env4 3,Env4 3,Env4 3,Env4 4,Test5 4,Test5 4,Test5 4,Test5 4,Test5 4,Test5 4,Test5 ,Env1 ,Env1 ,Env3 ,Env4 ,Env1

谢谢!

【问题讨论】:

  • 最好给我们看一些样本数据
  • 我不确定,但如果在执行 groupby 之前对数据框执行 .fillna(0),您可能会得到想要的结果
  • 试过 .fillna(0) 但没有区别,除非我能以某种方式扩展 .rename 条件以包含 0。
  • 好的,我现在更好地理解了这个问题。抱歉,我不知道如何有效地帮助您,但一个可能不太优雅的想法是生成一个组合数据框(仅与 Environment 分组,每行有 2 行(Other 和 Critical),然后您可以使用 merge with输出,在合并时用 0 填充缺失值).. 抱歉,如果没有帮助

标签: python python-3.x pandas dataframe


【解决方案1】:

groupby 不会显示 NaN 值,您需要先将它们替换为虚拟值:

In [11]: df = pd.DataFrame([[1, 2], [3, 4], [pd.np.nan, 6]], columns=["A", "B"])

In [12]: df
Out[12]:
     A  B
0  1.0  2
1  3.0  4
2  NaN  6

In [13]: df.groupby("A").mean()  # no nulls
Out[13]:
     B
A
1.0  2
3.0  4

例如,您可以使用 -1:

In [14]: df.replace({"A": {np.nan: -1}}).groupby("A").mean()
Out[14]:
      B
A
-1.0  6
 1.0  2
 3.0  4

In [15]: df.replace({"A": {np.nan: -1}}).groupby("A").mean().reset_index().replace({"A": {-1: np.nan}})
Out[15]:
     A  B
0  NaN  6
1  1.0  2
2  3.0  4

【讨论】:

    【解决方案2】:

    您需要reindex by MultiIndex by MultiIndex by MultiIndex.from_product 的第一级唯一值的所有组合:

    s = (df.groupby(['Environment/s Affected', 
                     (df['Rating'].isin(['1', '2']))]).size()
           .rename(index={True: 'Critical', False: 'Others P3+'}, level=1))
    print (s)
    Environment/s Affected  Rating    
    Env1                    Others P3+    6
    Env2                    Others P3+    4
    Env3                    Others P3+    6
    Env4                    Others P3+    7
    Test5                   Others P3+    7
    dtype: int64
    
    mux = pd.MultiIndex.from_product([df['Environment/s Affected'].unique(),
                                     ['Others P3+', 'Critical']],
                                     names=['Environment/s Affected','Rating'])
    print (mux)
    MultiIndex(levels=[['Env1', 'Env2', 'Env3', 'Env4', 'Test5'], ['Critical', 'Others P3+']],
               codes=[[0, 0, 1, 1, 2, 2, 3, 3, 4, 4], [1, 0, 1, 0, 1, 0, 1, 0, 1, 0]],
               names=['Environment/s Affected', 'Rating'])
    
    df1 = s.reindex(mux, fill_value=0).reset_index(name='counts')
    print (df1)
      Environment/s Affected      Rating  counts
    0                   Env1  Others P3+       6
    1                   Env1    Critical       0
    2                   Env2  Others P3+       4
    3                   Env2    Critical       0
    4                   Env3  Others P3+       6
    5                   Env3    Critical       0
    6                   Env4  Others P3+       7
    7                   Env4    Critical       0
    8                  Test5  Others P3+       7
    9                  Test5    Critical       0
    

    如果需要Critical 在最后一行添加sort_index:

    df1 = (s.reindex(mux, fill_value=0)
            .sort_index(level=[1,0], ascending=[False, True])
            .reset_index(name='counts'))
    print (df1)
      Environment/s Affected      Rating  counts
    0                   Env1  Others P3+       6
    1                   Env2  Others P3+       4
    2                   Env3  Others P3+       6
    3                   Env4  Others P3+       7
    4                  Test5  Others P3+       7
    5                   Env1    Critical       0
    6                   Env2    Critical       0
    7                   Env3    Critical       0
    8                   Env4    Critical       0
    9                  Test5    Critical       0
    

    【讨论】:

      猜你喜欢
      • 2023-01-12
      • 1970-01-01
      • 1970-01-01
      • 2017-04-15
      • 2014-05-07
      • 2023-01-05
      • 1970-01-01
      • 2021-02-05
      • 2019-07-05
      相关资源
      最近更新 更多