【问题标题】:doing calculations in pandas dataframe based on trailing row根据尾随行在熊猫数据框中进行计算
【发布时间】:2013-07-31 18:28:23
【问题描述】:

是否可以根据不同列中的尾随行在 pandas 数据框中进行计算?像这样。

frame = pd.DataFrame({'a' : [True, False, True, False],
                  'b' : [25, 22, 55, 35]})

我希望输出是这样的:

A     B     C
True  25    
False 22   44
True  55   55
False 35   70

当 A 列中的尾随行为 False 时,C 列与 B 列相同,而当列中尾随行 时,C 列是 B 列 * 2 A 是真的吗?

【问题讨论】:

  • C 列中的第一个条目是空白吗?

标签: python python-2.7 pandas


【解决方案1】:

您可以使用where 系列方法:

In [11]: frame['b'].where(frame['a'], 2 * frame['b'])
Out[11]:
0    25
1    44
2    55
3    70
Name: b, dtype: int64

In [12]: frame['c'] = frame['b'].where(frame['a'], 2 * frame['b'])

您也可以使用apply(但这通常会更慢):

In [21]: frame.apply(lambda x: 2 * x['b'] if x['a'] else x['b'], axis=1

由于您使用的是“尾随行”,因此您需要使用shift:

In [31]: frame['a'].shift()
Out[31]:
0      NaN
1     True
2    False
3     True
Name: a, dtype: object

In [32]: frame['a'].shift().fillna(False)  # actually this is not needed, but perhaps clearer
Out[32]:
0    False
1     True
2    False
3     True
Name: a, dtype: object

然后使用 where 反过来:

In [33]: c = (2 * frame['b']).where(frame['a'].shift().fillna(False), frame['b'])

In [34]: c
Out[34]:
0    25
1    44
2    55
3    70
Name: b, dtype: int64

并更改第一行(例如,更改为 NaN,in pandas we use NaN for missing data)

In [35]: c = c.astype(np.float)  # needs to accept NaN

In [36]: c.iloc[0] = np.nan

In [36]: frame['c'] = c

In [37]: frame
Out[37]:
       a   b   c
0   True  25 NaN
1  False  22  44
2   True  55  55
3  False  35  70

【讨论】:

  • 我想根据“a”列中的尾行进行计算。因此,第二行列“c”中的值取决于列“b”第二行中的值和列“a”中的第一行。我不知道如何或是否可以获取尾行的值。​​
  • 上一行第 2 行是第 3 行的上一行
  • 感谢 shift 是我一直在寻找的。现在我可以在 Pandas 书中找到示例
猜你喜欢
  • 2023-01-10
  • 2017-12-29
  • 2018-01-22
  • 2019-06-30
  • 1970-01-01
  • 2022-01-25
  • 1970-01-01
  • 2014-12-29
  • 2014-04-23
相关资源
最近更新 更多