【问题标题】:What is the most idiomatic way to index an object with a boolean array in pandas?在熊猫中用布尔数组索引对象的最惯用方法是什么?
【发布时间】:2013-05-12 07:26:49
【问题描述】:

我特别在谈论 Pandas 0.11 版,因为我正忙着用 .loc 或 .iloc 替换我对 .ix 的使用。我喜欢这样一个事实,即区分 .loc 和 .iloc 可以传达我是打算按标签还是整数位置进行索引。我看到任何一个都会接受一个布尔数组,但我想保持它们的使用纯粹以清楚地传达我的意图。

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    我目前正在使用[],即__getitem__(),例如

    df = pd.DataFrame(dict(a=range(5)))
    df[df.a%2==0]
    

    【讨论】:

    • 不要使用dict构造函数,使用{'a': range(5)}看起来更好
    • dict 文字也恰好要快得多。
    • 很高兴知道速度差异。我更喜欢 dict 构造函数的外观,而且我不必在关键字参数周围加上引号,这使得输入更容易。但是,如果速度差异很大,那么也许我会切换到 dict 文字。
    【解决方案2】:

    在 11.0 中,这三种方法都有效,suggested in the docs 的方式就是使用df[mask]。然而,这不是在位置上完成的,而是纯粹使用标签,所以在我看来loc 最能描述实际发生的事情。

    更新:我在 github 上询问过这个问题,结论是 df.iloc[msk] 将在 pandas @ 中给出 NotImplementedError(如果整数索引掩码)或 ValueError(如果非整数索引) 987654328@.

    In [1]: df = pd.DataFrame(range(5), list('ABCDE'), columns=['a'])
    
    In [2]: mask = (df.a%2 == 0)
    
    In [3]: mask
    Out[3]:
    A     True
    B    False
    C     True
    D    False
    E     True
    Name: a, dtype: bool
    
    In [4]: df[mask]
    Out[4]:
       a
    A  0
    C  2
    E  4
    
    In [5]: df.loc[mask]
    Out[5]:
       a
    A  0
    C  2
    E  4
    
    In [6]: df.iloc[mask]  # Due to this question, this will give a ValueError (in 11.1)
    Out[6]:
       a
    A  0
    C  2
    E  4
    

    也许值得注意的是,如果你给掩码整数索引它会抛出一个错误:

    mask.index = range(5)
    df.iloc[mask]  # or any of the others
    IndexingError: Unalignable boolean Series key provided
    

    这表明 iloc 并未实际实现,它使用标签,因此当我们尝试此操作时,为什么 11.1 会抛出 NotImplementedError。

    【讨论】:

    • 谢谢,我没有考虑整数索引的 .iloc 行为。老实说,在这种情况下,我实际上忘记了掩码具有 .index 属性,我认为它纯粹是一个布尔 numpy 数组。我同意,考虑到 .index 属性实际上首先用于对齐,.loc 可能是最好的选择。
    猜你喜欢
    • 2021-05-17
    • 2021-03-19
    • 1970-01-01
    • 1970-01-01
    • 2017-05-06
    • 1970-01-01
    • 2020-06-27
    • 2018-01-13
    • 2017-02-27
    相关资源
    最近更新 更多