【问题标题】:Assign value to subset of rows in Pandas dataframe为 Pandas 数据框中的行子集赋值
【发布时间】:2014-04-04 18:00:12
【问题描述】:

我想根据 Pandas DataFrame 中的索引条件分配值。

class test():
    def __init__(self):
        self.l = 1396633637830123000
        self.dfa = pd.DataFrame(np.arange(20).reshape(10,2), columns = ['A', 'B'], index = arange(self.l,self.l+10))
        self.dfb = pd.DataFrame([[self.l+1,self.l+3], [self.l+6,self.l+9]], columns = ['beg', 'end'])

    def update(self):
        self.dfa['true'] = False
        self.dfa['idx'] = np.nan
        for i, beg, end in zip(self.dfb.index, self.dfb['beg'], self.dfb['end']):
            self.dfa.ix[beg:end]['true'] = True
            self.dfa.ix[beg:end]['idx'] = i

    def do(self):
        self.update()
        print self.dfa

t = test()
t.do()

结果:

                      A   B   true  idx
1396633637830123000   0   1  False  NaN
1396633637830123001   2   3   True  NaN
1396633637830123002   4   5   True  NaN
1396633637830123003   6   7   True  NaN
1396633637830123004   8   9  False  NaN
1396633637830123005  10  11  False  NaN
1396633637830123006  12  13   True  NaN
1396633637830123007  14  15   True  NaN
1396633637830123008  16  17   True  NaN
1396633637830123009  18  19   True  NaN

true 列已正确分配,而 idx 列未正确分配。此外,这似乎取决于列的初始化方式,因为如果我这样做:

    def update(self):
        self.dfa['true'] = False
        self.dfa['idx'] = False

true 列也未正确分配。

我做错了什么?

附言预期结果是:

                      A   B   true  idx
1396633637830123000   0   1  False  NaN
1396633637830123001   2   3   True  0
1396633637830123002   4   5   True  0
1396633637830123003   6   7   True  0
1396633637830123004   8   9  False  NaN
1396633637830123005  10  11  False  NaN
1396633637830123006  12  13   True  1
1396633637830123007  14  15   True  1
1396633637830123008  16  17   True  1
1396633637830123009  18  19   True  1

编辑:我尝试使用 loc 和 iloc 进行分配,但它似乎不起作用: 位置:

self.dfa.loc[beg:end]['true'] = True
self.dfa.loc[beg:end]['idx'] = i

iloc:

self.dfa.loc[self.dfa.index.get_loc(beg):self.dfa.index.get_loc(end)]['true'] = True
self.dfa.loc[self.dfa.index.get_loc(beg):self.dfa.index.get_loc(end)]['idx'] = i

【问题讨论】:

  • 您是链式索引,请参见此处:pandas.pydata.org/pandas-docs/stable/…,不适用于多类型框架。试试df.loc[row_indexer,col_indexer] = value
  • 是的,我看过,但我不明白如何解决它。如果 dfb 使用标签索引值,如何获取 row_indexer、col_indexer?找到它:self.dfa.index.get_loc(beg)
  • 另外,如果我使用pd.set_option('mode.chained_assignment','warn'),我不会收到任何警告

标签: python pandas


【解决方案1】:

您正在链索引,请参阅here。警告并非保证会发生。

你应该这样做。不需要真正跟踪 b 中的索引,顺便说一句。

In [44]: dfa = pd.DataFrame(np.arange(20).reshape(10,2), columns = ['A', 'B'], index = np.arange(l,l+10))

In [45]: dfb = pd.DataFrame([[l+1,l+3], [l+6,l+9]], columns = ['beg', 'end'])

In [46]: dfa['in_b'] = False

In [47]: for i, s in dfb.iterrows():
   ....:     dfa.loc[s['beg']:s['end'],'in_b'] = True
   ....:     

或者如果你有非整数数据类型

In [36]: for i, s in dfb.iterrows():
             dfa.loc[(dfa.index>=s['beg']) & (dfa.index<=s['end']),'in_b'] = True


In [48]: dfa
Out[48]: 
                      A   B  in_b
1396633637830123000   0   1  False
1396633637830123001   2   3  True
1396633637830123002   4   5  True
1396633637830123003   6   7  True
1396633637830123004   8   9  False
1396633637830123005  10  11  False
1396633637830123006  12  13  True
1396633637830123007  14  15  True
1396633637830123008  16  17  True
1396633637830123009  18  19  True

[10 rows x 3 columns

如果 b 很大,这可能不是那么高效。

顺便说一句,这些看起来像纳秒级。通过转换它们可以更友好。

In [49]: pd.to_datetime(dfa.index)
Out[49]: 
<class 'pandas.tseries.index.DatetimeIndex'>
[2014-04-04 17:47:17.830123, ..., 2014-04-04 17:47:17.830123009]
Length: 10, Freq: None, Timezone: None

【讨论】:

  • 谢谢杰夫!这有效,除了 dfb 也有浮动列的情况。在这种情况下 iterrows() 将返回浮点数组,自动将整数值转换为浮点数,然后这些不能用于索引。
  • 在这种情况下,我使用以下方法解决了它:for i, s in self.dfb[['beg', 'end']].iterrows():
  • 最后,如果我必须简单地进行查找,为什么要转换为日期时间?
  • 自纪元以来,我发现 datetimes 比 ns 更直观,但您可能会发现其他情况:)
  • 好的。我只是觉得对日期时间进行索引很麻烦,但这可能是因为我对 pandas 的了解不够。感谢您的帮助!
猜你喜欢
  • 1970-01-01
  • 2012-08-31
  • 2019-01-03
  • 1970-01-01
  • 1970-01-01
  • 2017-08-02
  • 2020-12-09
  • 2016-01-19
  • 1970-01-01
相关资源
最近更新 更多