【问题标题】:What's the best way to filter a pandas series and then set the values of the part selected?过滤熊猫系列然后设置所选部分的值的最佳方法是什么?
【发布时间】:2015-03-12 16:05:20
【问题描述】:

例如我有一个这样的系列:

AFalse    0.220522 
BTrue    -1.050370
CFalse   -1.202922
DTrue     0.950305
EFalse    0.003110
FTrue     1.115483
GFalse    0.767281
HTrue    -1.376692
IFalse    1.729867
JTrue     2.574027
dtype: float64

我想只过滤掉具有“真”的行并将值设置为无。从速度的角度来看,最好的方法是什么?我将运行这个操作几百万次。谢谢。

【问题讨论】:

  • True 和 False 行总是交替的吗?
  • 这个系列是如何构建的?如果您可以在构建系列期间将A,B,Cs 与True、Falses 分开,则可以更快地选择True 行。
  • 这里的系列只是一个例子。实际系列是一个限价订单簿快照,其名称为 'bid.1.price'、'bid.2.price' ...、'ask.1.price'、'ask.2.price' .. .,并且我希望能够根据快速选择方案更改此类条目的值。

标签: select pandas filtering series setvalue


【解决方案1】:

使用矢量化操作有助于提高速度。从你的Series 开始

df = s.reset_index()
mask = df['index'].str.contains('True')
df.loc[mask, 'a'] = None
df.set_index('index')['a']

返回

index
AFalse    0.220522
BTrue          NaN
CFalse   -1.202922
DTrue          NaN
EFalse    0.003110
FTrue          NaN
GFalse    0.767281
HTrue          NaN
IFalse    1.729867
JTrue          NaN
Name: a, dtype: float64

【讨论】:

  • 该方法有效,但是该操作创建了一个新的df,这将极大地影响我的内存使用。当执行大量此操作时,这会影响代码性能。有没有办法过滤掉系列的视图并设置它的值,而不是创建系列的副本?
  • 这是我最好的:idxs = filter(lambda x: x.endswith('True'),s.index)s.loc[idxs] = None
【解决方案2】:

初始化一个随机序列:

s = pd.Series(np.random.rand(4), index=idx, name='test')

使用列表推导在索引上创建掩码。请注意,“True”可以是索引值内的任何位置。然后使用 .loc 设置您已经完成的值。

idx = ['True' in i for i in s.index]
s.loc[idx] = None  # or np.NaN

>>> s
Out[50]: 
ATrue          None
BFalse    0.9134165
CTrue          None
DFalse    0.2530273
Name: test, dtype: object

【讨论】:

  • 和我自己的很相似,但我相信这就是我要找的。​​span>
猜你喜欢
  • 1970-01-01
  • 2021-10-10
  • 2019-10-26
  • 1970-01-01
  • 1970-01-01
  • 2022-06-15
  • 2015-11-23
  • 2021-05-23
  • 1970-01-01
相关资源
最近更新 更多