【问题标题】:Pandas - Remove rows based on combinations of NaN valuesPandas - 根据 NaN 值的组合删除行
【发布时间】:2014-11-11 19:44:08
【问题描述】:

我有一个看起来像这样的数据框:

NUM   A      B        C      D        E        F
p1    NaN    -1.183   NaN    NaN      NaN      1.829711
p5    NaN    NaN      NaN    NaN      1.267   -1.552721
p9    1.138  NaN      NaN    -1.179   NaN      1.227306

在 F 列和至少一个其他 A-E 列中始终存在一个非 NaN 值。

我想创建一个子表,其中仅包含那些在列中包含非 NaN 值的某些组合的行。有许多这些所需的组合,包括双胞胎和三胞胎。以下是我想要提取的三种此类组合的示例:

  1. A 列和 B 列中包含非 NaN 值的行
  2. 在 C 和 D 中包含非 NaN 值的行
  3. 在 A & B & C 中包含非 NaN 值的行

我已经从question 中了解了 np.isfinite 和 pd.notnull 命令,但我不知道如何将它们应用于列组合。

此外,一旦我有一个用于删除与我所需组合之一不匹配的行的命令列表,我不知道如何告诉 Pandas 仅在它们与任何所需组合不匹配时才删除行。

【问题讨论】:

  • 如果您添加代码为其他人重现数据框,那就太好了。这样,我们可以直接回答问题,而不是花时间尝试设置示例数据框。
  • 我认为 Taha 的答案有我正在寻找的部分。酷。

标签: python pandas combinations dataframe


【解决方案1】:

很多时候,作为选择数据帧子集的一部分,我们需要对布尔数组(numpy 数组或 pandas 系列)进行逻辑运算。对此使用 'and'、'or'、'not' 运算符将不起作用。

In [79]: df[pd.notnull(df['A']) and pd.notnull(df['F'])]

ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().

在 Python 中,当使用 'and'、'or' 和 'not' 运算符时,非布尔变量通常被认为是 True,除非它们表示像 []、int(0)、float(0) 这样的“空”对象, None 等。因此,在 Pandas 中使用这些相同的运算符进行数组布尔运算会令人困惑。有些人会期望他们简单地评估为True

相反,我们应该为此使用&、| 和~。

In [69]: df[pd.notnull(df['A']) & pd.notnull(df['F'])]
Out[69]:
  NUM      A   B   C      D   E         F
2  p9  1.138 NaN NaN -1.179 NaN  1.227306

另一种更短但不太灵活的方法是使用any()、all() 或empty。

In [78]: df[pd.notnull(df[['A', 'F']]).all(axis=1)]
Out[78]:
  NUM      A   B   C      D   E         F
2  p9  1.138 NaN NaN -1.179 NaN  1.227306

你可以阅读更多关于这个here

【讨论】:

  • 这很有帮助。我可以使用它为每个兴趣组合创建一个新的数据框,然后将它们全部连接在一起。
【解决方案2】:

您可以使用apply 和lambda 函数来选择非Nan 值。您可以使用Numpy.isNan(..) 验证它是否是 Nan 值。

data="""NUM   A      B        C      D        E        F
p1    NaN    -1.183   NaN    NaN      NaN      1.829711
p5    NaN    NaN      NaN    NaN      1.267   -1.552721
p9    1.138  NaN      NaN    -1.179   NaN      1.227306"""

import pandas as pd
from io import StringIO

df= pd.read_csv(StringIO(data.decode('UTF-8')),delim_whitespace=True )
print df



# Rows which contain non-NaN values in columns A & B
df["A_B"]= df.apply(lambda x: x['A'] if np.isnan(x['B']) else x['B'] if np.isnan(x['A']) else 0, axis=1)

# Rows which contain non-NaN values in C & D
df["C_D"]= df.apply(lambda x: x['C'] if np.isnan(x['D']) else x['D'] if np.isnan(x['C']) else 0, axis=1)

# Rows which contain non-NaN values in A & B & C
df["A_B_C"]= df.apply(lambda x: x['C'] if np.isnan(x['A_B']) else x['A_B'] if np.isnan(x['C']) else 0, axis=1)
print df

# Rows which contain non-NaN values in A & B & C
df["A_B_C_D"]= df.apply(lambda x: x['A_B'] if np.isnan(x['C_D']) else x['C_D'] if np.isnan(x['A_B']) else 0, axis=1)
print df

输出:

  NUM      A      B   C      D      E         F    A_B    C_D  A_B_C
0  p1    NaN -1.183 NaN    NaN    NaN  1.829711 -1.183    NaN -1.183
1  p5    NaN    NaN NaN    NaN  1.267 -1.552721    NaN    NaN    NaN
2  p9  1.138    NaN NaN -1.179    NaN  1.227306  1.138 -1.179  1.138

如果你不需要经过条件案例,你可以查看其他帖子中解释的其他方式。

【讨论】:

    【解决方案3】:

    假设您的数据框名为df。你可以像这样使用布尔掩码。

    # Specify column combinations that you want to pull 
    combo1 = ['A', 'B'] 
    
    # Select rows in the data frame that have non-NaN values in the combination
    # of columns specified above
    
    notmissing = ((df.loc[:, combo1].notnull()))
    df = df.loc[notmissing, :] 
    

    【讨论】:

    • 这产生了一条错误消息:“ValueError:无法将大小为 2 的序列复制到维度为 716 的数组轴”。我的数据框有 716 行...
    猜你喜欢
    • 1970-01-01
    • 2014-09-28
    • 2017-07-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-08-12
    相关资源
    最近更新 更多