【问题标题】:Pyspark remove columns with 10 null valuesPyspark 删除具有 10 个空值的列
【发布时间】:2019-09-27 21:14:31
【问题描述】:

我是 PySpark 的新手。

我已经阅读了镶木地板文件。我只想保留至少有 10 个值的列

我使用 describe 来获取每列的非空记录数

我现在如何提取值少于 10 个的列名,然后在写入新文件之前删除这些列

df = spark.read.parquet(file)

col_count = df.describe().filter($"summary" == "count")

【问题讨论】:

标签: pyspark parquet


【解决方案1】:

您可以将其转换为字典,然后根据它们的值过滤掉键(列名)(计数StringType(),需要转换为int 在 Python 代码中):

# here is what you have so far which is a dataframe
col_count = df.describe().filter('summary == "count"')

# exclude the 1st column(`summary`) from the dataframe and save it to a dictionary
colCountDict = col_count.select(col_count.columns[1:]).first().asDict()

# find column names (k) with int(v) < 10
bad_cols = [ k for k,v in colCountDict.items() if int(v) < 10 ]

# drop bad columns
df_new = df.drop(*bad_cols)

一些注意事项:

  • 如果无法直接从 df.describe() 或 df.summary() 等获取信息,请使用 @pault 的方法。

  • 您需要 drop() 而不是 select() 列,因为 describe()/summary() 仅包含 numeric和 string 列,selecting 来自 df.describe() 处理的列表中的列将丢失 TimestampType()、ArrayType() 等列

【讨论】:

  • n/p,提醒一下,如果您还想检查和删除 Date、TimeStamp 列,这将无济于事,因为 df.describe 或 df.summary 不会计算这些列。周末愉快:)
  • 感谢@jxc。 parquet 文件具有作为列表中结构的 Row 对象的属性。因此,我能够使用 groupby 和 count 来查找坏列,然后使用 where 子句将它们过滤掉。而且,该解决方案处理所有数据类型。提供的示例有助于达成通用解决方案
猜你喜欢
  • 2021-06-27
  • 2018-12-21
  • 1970-01-01
  • 1970-01-01
  • 2017-10-25
  • 2021-11-19
  • 2017-11-26
  • 1970-01-01
  • 2012-02-25
相关资源
最近更新 更多