【发布时间】:2014-04-04 18:00:12
【问题描述】:
我想根据 Pandas DataFrame 中的索引条件分配值。
class test():
def __init__(self):
self.l = 1396633637830123000
self.dfa = pd.DataFrame(np.arange(20).reshape(10,2), columns = ['A', 'B'], index = arange(self.l,self.l+10))
self.dfb = pd.DataFrame([[self.l+1,self.l+3], [self.l+6,self.l+9]], columns = ['beg', 'end'])
def update(self):
self.dfa['true'] = False
self.dfa['idx'] = np.nan
for i, beg, end in zip(self.dfb.index, self.dfb['beg'], self.dfb['end']):
self.dfa.ix[beg:end]['true'] = True
self.dfa.ix[beg:end]['idx'] = i
def do(self):
self.update()
print self.dfa
t = test()
t.do()
结果:
A B true idx
1396633637830123000 0 1 False NaN
1396633637830123001 2 3 True NaN
1396633637830123002 4 5 True NaN
1396633637830123003 6 7 True NaN
1396633637830123004 8 9 False NaN
1396633637830123005 10 11 False NaN
1396633637830123006 12 13 True NaN
1396633637830123007 14 15 True NaN
1396633637830123008 16 17 True NaN
1396633637830123009 18 19 True NaN
true 列已正确分配,而 idx 列未正确分配。此外,这似乎取决于列的初始化方式,因为如果我这样做:
def update(self):
self.dfa['true'] = False
self.dfa['idx'] = False
true 列也未正确分配。
我做错了什么?
附言预期结果是:
A B true idx
1396633637830123000 0 1 False NaN
1396633637830123001 2 3 True 0
1396633637830123002 4 5 True 0
1396633637830123003 6 7 True 0
1396633637830123004 8 9 False NaN
1396633637830123005 10 11 False NaN
1396633637830123006 12 13 True 1
1396633637830123007 14 15 True 1
1396633637830123008 16 17 True 1
1396633637830123009 18 19 True 1
编辑:我尝试使用 loc 和 iloc 进行分配,但它似乎不起作用: 位置:
self.dfa.loc[beg:end]['true'] = True
self.dfa.loc[beg:end]['idx'] = i
iloc:
self.dfa.loc[self.dfa.index.get_loc(beg):self.dfa.index.get_loc(end)]['true'] = True
self.dfa.loc[self.dfa.index.get_loc(beg):self.dfa.index.get_loc(end)]['idx'] = i
【问题讨论】:
-
您是链式索引,请参见此处:pandas.pydata.org/pandas-docs/stable/…,不适用于多类型框架。试试
df.loc[row_indexer,col_indexer] = value -
是的,我看过,但我不明白如何解决它。如果 dfb 使用标签索引值,如何获取 row_indexer、col_indexer?找到它:self.dfa.index.get_loc(beg)
-
另外,如果我使用
pd.set_option('mode.chained_assignment','warn'),我不会收到任何警告