【问题标题】:Run same function on multiples columns and return a new column在多个列上运行相同的函数并返回一个新列
【发布时间】:2021-12-30 02:58:21
【问题描述】:

我在数据框中有几列具有绿色/黄色/红色值:

示例:

Date Index1 Index2
20-Dec-21 Green Yellow
21-Dec-21 Red Yellow

我想在此数据框中再添加一列,首先根据逻辑为每一列分配一个分数:如果绿色,则得分 = 1,如果黄色,则为 0.5,如果红色,则为 0,然后将这些单独的分数相加以产生最终分数.例如。对于第 1 行,得分 = 1+0.5 = 1.5,对于第 2 行得分 = 0+0.5 =0.5,依此类推。

func 本身很容易写:

def color_to_score(x):
    if (x=='Green'):
        return 1
    elif (x=='Yellow'):
        return 0.5
    else: return 0

但我正在努力将其应用于每一列,然后将结果分数添加到各列中,以一种优雅的方式生成一个新分数。 我显然可以这样做:

df['Index1score'] = df['Index1'].apply(color_to_score)

为每个相关列生成一个分数列,然后添加它们,但这是非常不雅且不可扩展的。寻求帮助。

【问题讨论】:

  • 您为什么认为这是“不优雅且不可扩展”?我不同意这两种说法。您有一个语句从现有数据创建一个全新的列。
  • 你可以使用replace;例如df['Index1score'] = df['Index1'].replace({'Green': 1, 'Yellow': 0.5, 'Red': 0}, regex=True)
  • 所以对于最终的(综合)得分,我可以这样说:``` df['score'] = (df['Index1'].apply(color_to_score) + df['Index1 '].apply(color_to_score)) + ..... ``` 但是一旦列数超过某个否,这个解决方案就不是很优雅。所以可能我在这里缺少一个技巧来指定所有列名。

标签: python pandas dataframe


【解决方案1】:

这是使用replace()的替代方法:

replace_dict = {'Green':1,'Yellow':.5,'\w':0}
df.assign(new_col = df[['col1','col2']].replace(replace_dict,regex=True).sum(axis=1))

另外,您可以使用pd.to_numeric() 并设置errors = 'coerce' 将所有非数值转换为NaN,而不是使用\w 来替换所有其他单词

replace_dict = {'Green':1,'Yellow':.5}
df.assign(new_col = pd.to_numeric(df[['col1','col2']].replace(replace_dict).stack(),errors='coerce').unstack().sum(axis=1))

输出:

        Date   col1    col2  new_col
0  20-Dec-21  Green  Yellow      1.5
1  21-Dec-21    Red  Yellow      0.5

【讨论】:

    【解决方案2】:
    1. 您需要提供 Axis=1 才能将函数应用于每一行。 2.x 在您的函数中将是一行(不是单元格)。
    2. 您可以将x 行转换为列表。
    3. 计算每个值在列表中出现的次数,然后乘以它的值。
    4. 求和,并输出结果。

    【讨论】:

      【解决方案3】:

      选择 Python 而不是 pandas,让您的生活更轻松。

      score_dict = {'Green': 1, 'Yellow': 0.5, 'Red': 0}
      
      df = pd.DataFrame(data)
      df["total_score"] = 0.0
      
      for index, row in df.iterrows():
         df.at[index, "total_score"] = score_dict[row["Index1"]] + score_dict[row["Index2"]]
      
      print(df)
      

      【讨论】:

        【解决方案4】:

        想出了这个解决方案。

        scores = []
        for index in range(len(df.index)):
            scoreTotal = 0
            for column in df.columns:
                color = df[column][index]
                scoreTotal += color_to_score(color)
        
            scores.append(scoreTotal)
        
        df["Score"] = scores
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2015-02-18
          • 2019-05-15
          • 2022-11-20
          • 1970-01-01
          • 2021-01-24
          • 1970-01-01
          • 2012-11-22
          • 1970-01-01
          相关资源
          最近更新 更多