【问题标题】:How can I count the unique values from all columns and display them in a separate dataframe w.r.t. their unique name?如何计算所有列的唯一值并将它们显示在单独的数据框中 w.r.t.他们独特的名字?
【发布时间】:2020-09-21 11:34:59
【问题描述】:
| 1st Most Common Value | 2nd Most Common Value | 3rd Most Common Value | 4th Most Common Value | 5th Most Common Value |
|-----------------------|-----------------------|-----------------------|-----------------------|-----------------------|
| Grocery Store         | Pub                   | Coffee Shop           | Clothing Store        | Park                  |
| Pub                   | Grocery Store         | Clothing Store        | Park                  | Coffee Shop           |
| Hotel                 | Theatre               | Bookstore             | Plaza                 | Park                  |
| Supermarket           | Coffee Shop           | Pub                   | Park                  | Cafe                  |
| Pub                   | Supermarket           | Coffee Shop           | Cafe                  | Park                  |

数据框的名称是 df0。如您所见,所有列中有许多重复的值。所以我想创建一个数据框,其中包含所有列中的所有唯一值及其频率。由于我想创建它的条形图,有人可以帮忙处理代码吗?

输出应该如下:

| Venues         | Count |
|----------------|-------|
| Bookstore      | 1     |
| Cafe           | 2     |
| Coffee Shop    | 4     |
| Clothing Store | 2     |
| Grocery Store  | 2     |
| Hotel          | 1     |
| Park           | 5     |
| Plaza          | 1     |
| Pub            | 4     |
| Supermarket    | 2     |
| Theatre        | 1     |

【问题讨论】:

  • 您的预期输出是什么?如果您可以不将数据粘贴为图像,那也很好
  • 首先运行 fd0.describe()
  • 所以基本上每列都需要.value_counts()?
  • @NYCCoder 我已经修改了我的代码,请检查并告诉我。谢谢。
  • @CeliusStingher 我已经修改了我的代码,请检查并告诉我。谢谢。

标签: python pandas numpy dataframe data-science


【解决方案1】:

编辑:我在原始答案中领先于自己(也感谢 OP 添加编辑/预期输出)。你要this post,我觉得最简单的答案:

new_df = pd.DataFrame(df0.stack().value_counts())

如果您不关心值来自哪一列,而您只想要它们的计数,请使用 value_counts()(正如 @Celius Stingher 在 cmets 中所说),遵循 this post。

如果您确实想报告每一列的每个值的频率,您可以为每一列使用value_counts(),但您最终可能会得到不均匀的条目(要回到DataFrame,您可以做一些有点像join)。

我创建了一个小函数来计算 df 中值的出现次数,并返回一个新值:

import pandas as pd
import numpy as np

def counted_entries(df, array):
    output = pd.DataFrame(columns=df.columns, index=array)
    for i in array:
        output.loc[i] = (df==i).sum()
    return output

这适用于填充随机动物值名称的df。您只需通过获取其值的set 来传递df 中的唯一条目:

columns = ['Column ' + str(i+1) for i in range(10)]
index = ['Row ' + str(i+1) for i in range(5)]

df = pd.DataFrame(np.random.choice(['pig','cow','sheep','horse','dog'],size=(5,10)), columns=columns, index=index)

unique_vals = list(set(df.stack())) #this is all the possible entries in the df

df2 = counted_entries(df, unique_vals)

df 之前:

      Column 1 Column 2 Column 3 Column 4  ... Column 7 Column 8 Column 9 Column 10
Row 1      pig      pig      cow      cow  ...      cow      pig      dog       pig
Row 2    sheep      cow      pig    sheep  ...      dog      pig      pig       cow
Row 3      cow      cow      cow    sheep  ...    horse      dog    sheep     sheep
Row 4    sheep      cow    sheep      cow  ...      cow    horse      pig       pig
Row 5      dog      pig    sheep    sheep  ...    sheep    sheep    horse     horse

counted_entries()的输出

       Column 1  Column 2  Column 3  ...  Column 8  Column 9  Column 10
pig           1         2         1  ...         2         2          2
horse         0         0         0  ...         1         1          1
sheep         2         0         2  ...         1         1          1
dog           1         0         0  ...         1         1          0
cow           1         3         2  ...         0         0          1

【讨论】:

  • 我认为 pandas 有足够的功能,因此不需要定义自定义功能,但答案都很好! +1 :)
【解决方案2】:

感谢您的编辑,也许这就是您正在寻找的内容,使用 value_counts 获取完整数据框,然后聚合输出:

df0 = pd.DataFrame({'1st':['Grocery','Pub','Hotel','Supermarket','Pub'],
                    '2nd':['Pub','Grocery','Theatre','Coffee','Supermarket'],
                    '3rd':['Coffee','Clothing','Supermarket','Pub','Coffee'],
                    '4th':['Clothing','Park','Plaza','Park','Cafe'],
                    '5th':['Park','Coffee','Park','Cafe','Park']})

df1 = df0.apply(pd.Series.value_counts)
df1['Count'] = df1.sum(axis=1)
df1 = df1.reset_index().rename(columns={'index':'Venues'}).drop(columns=list(df0))
print(df1)

输出:

        Venues  Count
5         Park    5.0
2       Coffee    4.0
7          Pub    4.0
8  Supermarket    3.0
0         Cafe    2.0
1     Clothing    2.0
3      Grocery    2.0
4        Hotel    1.0
6        Plaza    1.0
9      Theatre    1.0

【讨论】:

    【解决方案3】:

    您也可以这样做:

    df = pd.read_csv('test.csv', sep=',')
    list_of_list = df.values.tolist()
    t_list = sum(list_of_list, [])
    df = pd.DataFrame(t_list)
    df.columns = ['Columns']
    df = df.groupby(by=['Columns'], as_index=False).size().to_frame().reset_index().rename(columns={0: 'Count'})
    print(df)
    
               Columns  Count
    0        Bookstore      1
    1             Cafe      2
    2   Clothing Store      2
    3      Coffee Shop      4
    4    Grocery Store      2
    5            Hotel      1
    6             Park      5
    7            Plaza      1
    8              Pub      4
    9      Supermarket      2
    10         Theatre      1
    

    【讨论】:

      猜你喜欢
      • 2018-12-08
      • 1970-01-01
      • 2021-10-24
      • 2023-01-16
      • 2020-01-16
      • 1970-01-01
      • 2017-11-13
      • 1970-01-01
      • 2015-09-27
      相关资源
      最近更新 更多