【问题标题】:How to groupby and map by two columns pandas dataframe如何按两列熊猫数据框分组和映射
【发布时间】:2018-06-03 10:40:57
【问题描述】:

我在 python 上使用 pandas 数据框时遇到问题,我正在尝试使机器学习模型在表面上进行预测。我在火车数据框中有表面列,但在测试数据框中没有。所以,我会根据火车的表面创建一些特征。

train['error_cat1'] = abs(train.groupby(train['cat1'])['surface'].transform('mean')  - train.surface.mean())

在这里,我通过“cat”功能设置了 grouby 的值,平均值为 suface 。酷

现在我也必须将它添加到测试中。因此,将使用此方法将每个 groupby 的训练值映射到测试行。

mp = {k: g['error_cat1'].tolist()[0] for k,g in train.groupby('cat1')}
test['error_cat1'] = test['cat1'].map(mp)

所以,到目前为止没有问题。现在,我将在 groupby 中使用两列。

train['error_cat1_cat2'] = abs(train.groupby(train[['cat1','cat2']])['surface'].transform('mean')  - train.surface.mean())

但我不知道如何将其映射到测试数据框。请你帮我处理这个问题或者给我一些其他的方法让我可以做到。

谢谢

例如我的火车是

+------+------+-------+
| Cat1 | Cat2 | surface |
+------+------+-------+
| 1    | 3    | 10    |
+------+------+-------+
| 2    | 2    | 12    |
+------+------+-------+
| 3    | 1    | 12    |
+------+------+-------+
| 1    | 3    | 5     |
+------+------+-------+
| 2    | 2    | 10    |
+------+------+-------+
| 3    | 2    | 13    |
+------+------+-------+

我的测试是

+------+------+
| Cat1 | Cat2 |
+------+------+
| 1    | 2    |
+------+------+
| 2    | 1    |
+------+------+
| 3    | 1    |
+------+------+
| 1    | 3    |
+------+------+
| 2    | 3    |
+------+------+
| 3    | 1    |
+------+------+

现在我将在 cat1 和 cat2 上进行分组平均曲面,例如 (cat1,cat2)=(1,3) 上的平均曲面为 (10+5)/2 = 7.5

现在,我必须进行测试并将这个值映射到 (cat1,cat2)=(1,3) 行。

我希望你得到我。

【问题讨论】:

  • 您可以使用示例数据创建简单的代码,以便每个人都可以运行它并创建解决方案。

标签: python pandas pandas-groupby sklearn-pandas


【解决方案1】:

你可以使用

  • groupby().means()计算手段
  • reset_index() 将索引Cat1Cat2 再次转换为列
  • merge(how='left', ) 连接两个数据框,例如数据库中的表(LEFT JOIN in SQL)。

.

headers = ['Cat1', 'Cat2', 'surface']

train_data = [
    [1, 3, 10],
    [2, 2, 12],
    [3, 1, 12],
    [1, 3, 5],
    [2, 2, 10],
    [3, 2, 13],
]

test_data = [
    [1, 2],
    [2, 1],
    [3, 1],
    [1, 3],
    [2, 3],
    [3, 1],
]
import pandas as pd

train = pd.DataFrame(train_data, columns=headers)
test = pd.DataFrame(test_data, columns=headers[:-1])

print('--- train ---')
print(train)

print('--- test ---')
print(test)

print('--- means ---')
means = train.groupby(['Cat1', 'Cat2']).mean()
print(means)

print('--- means (dataframe) ---')
means = means.reset_index(level=['Cat1', 'Cat2'])
print(means)

print('--- result ----')
result = pd.merge(df2, means, on=['Cat1', 'Cat2'], how='left')
print(result)

print('--- result (fillna)---')
result = result.fillna(0)
print(result)

结果:

--- train ---
   Cat1  Cat2  surface
0     1     3       10
1     2     2       12
2     3     1       12
3     1     3        5
4     2     2       10
5     3     2       13
--- test ---
   Cat1  Cat2
0     1     2
1     2     1
2     3     1
3     1     3
4     2     3
5     3     1
--- means ---
           surface
Cat1 Cat2         
1    3         7.5
2    2        11.0
3    1        12.0
     2        13.0
--- means (dataframe) ---
   Cat1  Cat2  surface
0     1     3      7.5
1     2     2     11.0
2     3     1     12.0
3     3     2     13.0
--- result ----
   Cat1  Cat2  surface
0     1     2      NaN
1     2     1      NaN
2     3     1     12.0
3     1     3      7.5
4     2     3      NaN
5     3     1     12.0
--- result (fillna)---
   Cat1  Cat2  surface
0     1     2      0.0
1     2     1      0.0
2     3     1     12.0
3     1     3      7.5
4     2     3      0.0
5     3     1     12.0

【讨论】:

    猜你喜欢
    • 2017-05-13
    • 1970-01-01
    • 1970-01-01
    • 2019-11-15
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-10-17
    • 2018-07-19
    相关资源
    最近更新 更多