【问题标题】:Indexing using value from another DataFrame使用来自另一个 DataFrame 的值进行索引
【发布时间】:2021-07-31 06:21:13
【问题描述】:

熊猫版本 1.1.3.

我正在处理 pandas 中的两个数据框(格式为 csv),一个包含我正在分析的数据,另一个包含标签。两者都包含一个带有标识号的列。我已将数据框 2 的行索引设置为标识号。我正在使用一个 nlp 处理库来分析数据帧 1 中的数据,它返回一个布尔值。我想在数据帧 2 中索引一个值,使用来自数据帧 1 的标识号以一种花哨的索引方法。

df1 看起来像这样:

  ID report

1 A1 'bla bla'
2 A1 'blo blo'
3 A1 'blo bla'
4 A2 'bla bla'
5 A2 'blo blo'
6 A3 'foo blo'

df2 看起来像这样:

   label hypothesis
ID
A1 0     0
A2 1     0 
A3 1     0
A4 0     0
A5 1     0
A6 0     0

这是我的代码:

for i, row in df1.iterrows():
    index = row['ID']
    display(index)
    report = row['report']
    doc = nlp(report)
    if boole_is_True(doc):
        df2[[index, 'hypothesis']] = 1

这是结果:

'A1'

'A1'

'A1'

'A2'

bla bla
---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
<ipython-input-22-f9ba77c1e7ea> in <module>
     17         print(report)
---> 18         df2[[index, 'hypothesis']] = 1

~\anaconda3\lib\site-packages\pandas\core\frame.py in __getitem__(self, key)
   2906             if is_iterator(key):
   2907                 key = list(key)
-> 2908             indexer = self.loc._get_listlike_indexer(key, axis=1, raise_missing=True)[1]
   2909 
   2910         # take() does not accept boolean indexers

~\anaconda3\lib\site-packages\pandas\core\indexing.py in _get_listlike_indexer(self, key, axis, raise_missing)
   1252             keyarr, indexer, new_indexer = ax._reindex_non_unique(keyarr)
   1253 
-> 1254         self._validate_read_indexer(keyarr, indexer, axis, raise_missing=raise_missing)
   1255         return keyarr, indexer
   1256 

~\anaconda3\lib\site-packages\pandas\core\indexing.py in _validate_read_indexer(self, key, indexer, axis, raise_missing)
   1302             if raise_missing:
   1303                 not_found = list(set(key) - set(ax))
-> 1304                 raise KeyError(f"{not_found} not in index")
   1305 
   1306             # we skip the warning on Categorical

KeyError: "['A2'] not in index"

如何去掉方括号?为什么会有括号?这个问题有什么更好的标题,以便其他人更容易找到?

提前致谢!我的实际数据包含医疗报告,所以我害怕无法向您展示我的实际数据和代码。

【问题讨论】:

    标签: python pandas indexing


    【解决方案1】:

    使用时:

    df2[[index, 'hypothesis']] = 1
    

    pandas 在列索引中搜索传递的index 值,但无法在您的列名中找到A2。在您的情况下,您试图在df2 的行索引中找到这些值,因此您需要编写如下:

    df2.loc[index, 'hypothesis'] = 1
    

    .loc[] 接受行和列索引值。

    【讨论】:

      猜你喜欢
      • 2016-08-15
      • 2019-11-28
      • 1970-01-01
      • 2016-02-17
      • 1970-01-01
      • 1970-01-01
      • 2021-12-17
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多