【问题标题】:pandas - change value in column based on another columnpandas - 根据另一列更改列中的值
【发布时间】:2018-04-15 05:15:12
【问题描述】:

假设我有一个数据框all_data,例如:

Id  Zone        Neighb
1   NaN         IDOTRR
2   RL          Veenker
3   NaN         IDOTRR
4   RM          Crawfor
5   NaN         Mitchel

我想在“Zone”列中输入缺失值,这样“Neighb”是“IDOTRR”我将“Zone”设置为“RM”,而“Neighb”是“Mitchel”我设置了“RL” '。

all_data.loc[all_data.MSZoning.isnull() 
             & all_data.Neighborhood == "IDOTRR", "MSZoning"] = "RM"
all_data.loc[all_data.MSZoning.isnull() 
             & all_data.Neighborhood == "Mitchel", "MSZoning"] = "RL"

我明白了:

TypeError: 无效类型比较

C:\Users\pprun\Anaconda3\lib\site-packages\pandas\core\ops.py:798: FutureWarning:元素比较失败;返回标量 相反,但将来会执行元素比较
结果 = getattr(x, name)(y)

我确定这应该很简单,但我已经搞砸了太久了。请帮忙。

【问题讨论】:

  • 尝试将测试括在括号中:(all_data.MSZoning.isnull() ) & (all_data.Neighborhood == "IDOTRR")
  • @积累,谢谢

标签: python pandas dataframe


【解决方案1】:

使用 np.select 即

df['Zone'] = np.select([df['Neighb'] == 'IDOTRR',df['Neighb'] == 'Mitchel'],['RM','RL'],df['Zone'])
Id 区域邻居 0 1 RM IDOTRR 1 2 RL 文克尔 2 3 RM IDOTRR 3 4 RM 克劳福 4 5 RL 米切尔

在你的情况下,你可以使用

# Boolean mask of condition 1 
m1 = (all_data.MSZoning.isnull()) & (all_data.Neighborhood == "IDOTRR")
# Boolean mask of condition 2
m2 = (all_data.MSZoning.isnull()) & (all_data.Neighborhood == "Mitchel")

np.select([m1,m2],['RM','RL'],all_data["MSZoning"])

【讨论】:

  • 非常感谢,效果很好。但是,不能用.loc 来达到同样的目的吗?
  • .loc 使事情变得复杂,所以我使用了 np.select ,它更快更高效。如果您有更多条件,请告诉我。我会更新
【解决方案2】:

在 Python 中,& 优先于 ==

http://www.annedawson.net/Python_Precedence.htm

所以当你执行all_data.MSZoning.isnull() & all_data.Neighborhood == "Mitchel" 时,它会被解释为(all_data.MSZoning.isnull() & all_data.Neighborhood) == "Mitchel",而现在Python 会尝试将AND 与一个str 系列结合起来,并查看它是否等于单个str "Mitchel"。解决方案是将测试括在括号中:(all_data.MSZoning.isnull()) & (all_data.Neighborhood == "Mitchel")。有时如果我有很多选择器,我会将它们分配给变量,然后AND 它们,例如:

null_zoning = all_data.MSZoning.isnull()
Mitchel_neighb = all_data.Neighborhood == "Mitchel"
all_data.loc[null_zoning & Mitchel_neighb, "MSZoning"] = "RL"

这不仅解决了操作顺序问题,还意味着all_data.loc[null_zoning & Mitchel_neighb, "MSZoning"] = "RL"适合一行。

【讨论】:

    【解决方案3】:
    df.Zone=df.Zone.fillna(df.Neighb.replace({'IDOTRR':'RM','Mitchel':'RL'}))
    df
    Out[784]: 
       Id Zone   Neighb
    0   1   RM   IDOTRR
    1   2   RL  Veenker
    2   3   RM   IDOTRR
    3   4   RM  Crawfor
    4   5   RL  Mitchel
    

    【讨论】:

    • 那是开箱即用的想法,并且做了 OP 正在寻找的 +1。
    猜你喜欢
    • 2012-10-15
    • 2019-04-04
    • 2020-08-20
    • 2022-11-14
    • 2020-10-29
    • 1970-01-01
    • 2021-09-05
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多