【问题标题】:Pandas: how do you map a dictionary of dictionaries to 2 columns?Pandas:如何将字典映射到 2 列?
【发布时间】:2021-08-31 11:26:15
【问题描述】:

我有以下字典:

rates = {'USD': 
              {'2019': 1,
               '2020': 2,
               '2021': 3},
         'CAD':
              {'2019': 4,
               '2020': 5,
               '2021': 6}
         }

我有以下虚拟数据框:

   Item Currency Year Rate
0  1    USD      2019 
1  2    USD      2020
2  3    CAD      2021
3  4    CAD      2019
4  5    GBP      2020

我现在想通过映射正确的速率来填充列Rate,其中rate = f(currency,year)。我正在尝试:

def map_rate(data, rates):

    for index, row in data.iterrows():

        currency = str(row['Currency'])

        if currency in list(rates.keys()):

            year = str(row['Year'])
            rate = rates[currency][year]

        else:
            rate = 1

    return rate

我像下面这样使用上面的:

df['Rate'] = map_rate(test, rates)

但是,这只是返回第一速率,例如值 1,而不是适当的费率:

    Item Currency Year  Rate
0   1    USD      2019  1
1   2    USD      2020  1
2   3    CAD      2021  1
3   4    CAD      2019  1
4   5    GBP      2020  1

预期结果是:

    Item Currency Year  Rate
0   1    USD      2019  1
1   2    USD      2020  2
2   3    CAD      2021  6
3   4    CAD      2019  4
4   5    GBP      2020  1

我的错误在哪里?

【问题讨论】:

  • 旁注:您可以直接检查字典是否存在键:if currency in rates 而不是if currency in list(rates.keys())。后者形成一个列表并丢失~O(1) 查找时间。

标签: python pandas dataframe dictionary


【解决方案1】:

为费率创建另一个数据框

rates_df = pd.DataFrame(rates).unstack().reset_index()
rates_df.columns = ['Currency', 'Year', 'Rates']
rates_df['Year'] = rates_df['Year'].astype(int)

然后合并

df.merge(rates_df, on=['Currency', 'Year'], how='left').fillna(1)

费率数据框

  Currency  Year  Rates
0      USD  2019      1
1      USD  2020      2
2      USD  2021      3
3      CAD  2019      4
4      CAD  2020      5
5      CAD  2021      6

输出

   Item Currency  Year  Rates
0     1      USD  2019    1.0
1     2      USD  2020    2.0
2     3      CAD  2021    6.0
3     4      CAD  2019    4.0
4     5      GBP  2020    1.0

【讨论】:

    【解决方案2】:

    这可以通过内置的 Pandas 方法df.apply() 轻松完成。这是一个比其他发布的答案更详细的示例。

    代码:

    def get_rate(row):
      if row['Currency'] in rates.keys():
        return rates[row['Currency']][row['Year']]
      else:
        return 1
    
    df['Rate'] = df.apply(get_rate,axis=1)
    
    print(df)
    

    【讨论】:

      【解决方案3】:

      这是一种方法,使用 stack 从费率创建多索引系列,您可以使用 df 中的值 reindex 来获得每行所需的费率。

      df['rate'] = (
          pd.DataFrame(rates)
            .stack()
            .reindex(pd.MultiIndex.from_frame(df[['Year','Currency']].astype(str)), 
                     fill_value=1)
           .to_numpy()
      )
      print(df)
         Item Currency  Year  rate
      0     1      USD  2019     1
      1     2      USD  2020     2
      2     3      CAD  2021     6
      3     4      CAD  2019     4
      4     5      GBP  2020     1
      

      【讨论】:

        【解决方案4】:

        使用.apply

        例如:

        df['Rate'] = df.apply(lambda x: rates[x['Currency']][x['Year']], axis=1)
        # OR
        df['Rate'] = df.apply(lambda x: rates.get(x['Currency'], dict()).get(x['Year'], 1), axis=1)
        print(df)
        

        输出:

          Item Currency  Year  Rate
        0    1      USD  2019     1
        1    2      USD  2020     2
        2    3      CAD  2021     6
        3    4      CAD  2019     4
        4    5      GBP  2020     1
        

        【讨论】:

        • 谢谢。为什么您认为我的解决方案不起作用?
        • 我不知道为什么它不工作....可能是因为if currency in list(rates.keys()): 失败
        • 即使删除条件,我仍然会得到 1 的值。似乎它没有正确循环不同的货币名称。
        • @Zizzipupp 这是因为您的函数返回一个标量(在这种情况下为 1,因为它是循环中获得的最后一个值),而您应该返回一个值列表,每次迭代都有一个值跨度>
        • 您的数据可能不正确。检查它是否有前导空格?
        猜你喜欢
        • 2016-09-02
        • 2021-11-25
        • 1970-01-01
        • 2023-02-21
        • 2020-09-01
        • 2021-12-27
        • 1970-01-01
        • 1970-01-01
        • 2010-12-31
        相关资源
        最近更新 更多