【问题标题】:What's the equivalent of Panda's value_counts() in PySpark?PySpark 中 Panda 的 value_counts() 相当于什么?
【发布时间】:2018-12-06 09:34:50
【问题描述】:

我有以下 python/pandas 命令:

df.groupby('Column_Name').agg(lambda x: x.value_counts().max()

我在哪里获取 DataFrameGroupBy 对象中所有列的值计数。

如何在 PySpark 中执行此操作?

【问题讨论】:

  • 我要求的任务非常简单。我想按数据框获取组中所有列的 vale 计数(最高不同计数)。这很容易在 Pandas 中使用 value_counts() 方法完成。
  • 这是我的 DF:>>> schemaTrans.show() +----+----+------+-----+----+ ----+ |COL1|COL2| COL3| COL4|COL5|身份证| +----+----+------+-----+----+----+ | 123| 456|ABC123| XYZ| 525|ID01| | 123| 456|ABC123| XYZ| 634|ID01| | 123| 456|ABC123| XYZ| 802|ID01| | 456| 123| BC01|K_L_M| 213|ID01| | 456| 123| BC01|K_L_M| 401|ID01| | 456| 123| BC01|P_Q_M| 213|ID01| | 123| 456|XYZ012| ABC| 117|ID02| | 123| 456|XYZ012|安倍| 117|ID02| | 456| 123| QPR12|S_T_U| 204|ID02| | 456| 123| QPR12|S_T_X| 415|ID02| +----+----+------+-----+----+----+
  • from pyspark.sql.functions import count exprs = {x: "count" for x in schemaTrans.columns} schemaTrans.groupBy("ID").agg(exprs).show(5) + ----+---------+------------+------------+--------- +-----------+------------+ ID|count(ID)|count(COL4)|count(COL2)|count(COL3)|count(COL1 )|计数(COL5)| +----+---------+-----------+------------+---------- -+-----------+------------+ |ID01| 6| 6| 6| 6| 6| 6| |ID02| 4| 4| 4| 4| 4| 4| +----+---------+-----------+------------+---------- -+-----------+---------
  • exprs = [countDistinct(x) for x in schemaTrans.columns] schemaTrans.groupBy("ID").agg(*exprs).show(5) | ID|(DISTINCT COL1)|(DISTINCT COL2)|(DISTINCT COL3)|(DISTINCT COL4)|(DISTINCT COL5)|(DISTINCT ID)| +----+----------------+---------------+------------ ---+---------------+---------------+----------|ID01| 2 | 2 | 2 | 3 | 5 | 1 | |ID02| 2 | 2 | 2 | 4 | 3 | 1 | +----+----------------+---------------+------------ ---+---------------+---------------+---------
  • 请不要将这些添加为 cmets。 Edit你的问题并把它放在那里。另请阅读how do I format my code blocks。另请查看how to make good reproducible apache spark dataframe examples。

标签: dataframe count pyspark pandas-groupby


【解决方案1】:
from pyspark.sql import SparkSession
from pyspark.sql.functions import count, desc
spark = SparkSession.builder.appName('whatever_name').getOrCreate()
spark_sc = spark.read.option('header', True).csv(your_file)    
value_counts=spark_sc.select('Column_Name').groupBy('Column_Name').agg(count('Column_Name').alias('counts')).orderBy(desc('counts'))
value_counts.show()

但是spark在单机上比pandas value_counts()慢很多

【讨论】:

    【解决方案2】:

    试试这个:

    spark_df.groupBy('column_name').count().show()
    

    【讨论】:

      【解决方案3】:

      当你想控制订单时试试这个:

      data.groupBy('col_name').count().orderBy('count', ascending=False).show()
      

      【讨论】:

        【解决方案4】:

        或多或少是一样的:

        spark_df.groupBy('column_name').count().orderBy('count')
        

        在 groupBy 中,您可以有多个由 , 分隔的列

        例如groupBy('column_1', 'column_2')

        【讨论】:

        • 您好谭锦,感谢您的回复!我没有得到相同的结果。我一直在做以下事情:(Action-1): from pyspark.sql.functions import count exprs = {x: "count" for x in df.columns} df.groupBy("ID").agg(exprs)。 show(5),这行得通,但我得到了每组的所有记录数。那不是我想要的。 (Action-2) from pyspark.sql.functions import countDistinct exprs = [countDistinct(x) for x in df.columns] df.groupBy("ID").agg(*exprs).show(5) 这打破了!!它的错误如下: ERROR client.TransportResponseHandler:
        • 缺少的.show() 需要添加到该行的末尾才能真正看到结果,这可能会让初学者感到困惑。
        • 要匹配 Pandas 中的行为,您希望按降序返回计数:spark_df.groupBy('column_name').count().orderBy(col('count').desc()).show()
        猜你喜欢
        • 2019-03-06
        • 2023-03-25
        • 2018-09-07
        • 1970-01-01
        • 1970-01-01
        • 2018-12-15
        • 2020-04-26
        • 2019-09-12
        • 2020-02-28
        相关资源
        最近更新 更多