【问题标题】:PySpark dataframe: filter records with four or more non-null columnsPySpark 数据框:过滤具有四个或更多非空列的记录
【发布时间】:2016-04-03 23:08:31
【问题描述】:

我有许多 PySpark 数据框,其中两列中的数据是必需的,而其他列是可选的。必填列包含日期和记录 ID;最有价值的数据位于可选列中。我正在尝试捕获可选列中元素之间的联系。

数据框,预过滤器:

id     col1    col2    col3    date
123            xyz             20160401
234    abc     pqr             20160401
345    def     hij     klm     20160401
456                            20160401

过滤后,数据框将如下所示:

id     col1    col2    col3    date
234    abc     pqr             20160401
345    def     hij     klm     20160401

具有多个非空列值的记录很有趣,因为它们描述了关系。

我注意到 PySpark 有一个 .filter 方法。文档中的示例通常显示过滤列,例如df_filtered = df.filter(df.some_col > some_value)。我正在尝试编写一个过滤器来捕获具有四个或更多非空列的所有记录,用于 arbitrary 数据帧,即不得明确说明列名。

在 PySpark 中是否有一种简单的方法可以做到这一点?

更新

虽然.dropna(thresh=4) 似乎正是我正在寻找的东西,但由于某种原因它不起作用。例如

df.collect()

[Row(id=123, col1=None,         col2=None, col3=3754907743, date='20160403'),
 Row(id=124, col1=7911019393,   col2=None, col3=1456473867, date='20160403'),
 Row(id=125, col1=None,         col2=None, col3=2049622472, date='20160403'),
 Row(id=126, col1=4345043212,   col2=None, col3=3168577324, date='20160403'),
 Row(id=127, col1=None,         col2=None, col3=3185277065, date='20160403'),
 Row(id=128, col1=1336048242,   col2=None, col3=1322345860, date='20160403')]

不管thresh是多少,它总是返回原始数据框中的所有记录:

df_filtered = df.dropna(thresh=[any number])

df_filtered.collect()

[Row(id=123, col1=None,         col2=None, col3=3754907743, date='20160403'),
 Row(id=124, col1=7911019393,   col2=None, col3=1456473867, date='20160403'),
 Row(id=125, col1=None,         col2=None, col3=2049622472, date='20160403'),
 Row(id=126, col1=4345043212,   col2=None, col3=3168577324, date='20160403'),
 Row(id=127, col1=None,         col2=None, col3=3185277065, date='20160403'),
 Row(id=128, col1=1336048242,   col2=None, col3=1322345860, date='20160403')]

我正在运行 Spark 版本 1.5.0-cdh5.5.2。

【问题讨论】:

标签: python pyspark


【解决方案1】:

从docs,您正在寻找dropna:

dropna(how='any', thresh=None, subset=None)

返回一个新的 DataFrame,省略空值的行。 DataFrame.dropna() 和 DataFrameNaFunctions.drop() 是每个的别名 其他。-

参数:

how – ‘any’ or ‘all’. If ‘any’, drop a row if it contains any nulls.
       If ‘all’, drop a row only if all its values are null.

thresh – int, default None If specified, drop rows that have less than
         thresh non-null values. This overwrites the how parameter.

subset – optional list of column names to consider.

所以,要回答你的问题,你可以试试df.dropna(thresh=4)。

【讨论】:

  • 谢谢@AkshatMahajan。 dropna(thresh=...) 似乎正是我正在寻找的。由于某种奇怪的原因,它不起作用。
  • @AlexWoolford 我一直在尝试复制您正在经历的事情,我同意 - dropna 确实有点出人意料。我发现它确实适用于不直观的 thresh 值 - 你尝试了哪些测试数字?
  • 看来thresh 表示至少有多少列必须是非空的。如果我有三行,其中恰好有两个非空列,那么摆脱它的唯一方法是设置thresh=3。
  • 我尝试了1、2、3、4和5。它们都返回了原始数据帧。
  • @AlexWoolford:所以我尝试使用您在上面从collect() 发布的快照重新创建您的实际数据框,并在尝试转换列表时不断被告知“推断后无法确定某些类型”行到数据框。问题可能仅仅是数据框中的类型定义不明确吗?
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-03-07
  • 2017-06-28
  • 2012-06-07
相关资源
最近更新 更多