【问题标题】:Create a new pandas columns from multiple columns从多列创建新的熊猫列
【发布时间】:2018-06-14 11:48:23
【问题描述】:

这是数据框

    MatchId EventCodeId EventCode   Team1   Team2   Team1_Goals Team2_Goals xG_Team1    xG_Team2    CurrentPlaytime
0   865314  1029    Goal Home   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  457040
1   865314  1029    Goal Home   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  1405394
2   865314  2053    Goal Away   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  1898705
3   865314  2053    Goal Away   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  4388278
4   865314  1029    Goal Home   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  4507898
5   865314  1030    Cancel Goal Home    Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  4517728
6   865314  1029    Goal Home   Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  4956346
7   865314  1030    Cancel Goal Home    Northampton Crawley Town    2   2   2.067663207769023   0.8130662505484256  4960633
8   865316  2053    Goal Away   Coventry    Bradford    0   0   1.0847662440468118  1.2526705617472387  447858
9   865316  2054    Cancel Goal Away    Coventry    Bradford    0   0   1.0847662440468118  1.2526705617472387  456361

新列将按如下方式创建:

for EventCodeId = 1029 and EventCode = Goal Home
new_col1 = CurrentPlaytime/3*10**4

for EventCodeId = 2053 and ventCode = Goal Away
new_col2 = CurrentPlaytime/3*10**4

对于所有其他 EventCodeId 和 EventCode new_co1 和 new_col2 将占用 0.

这是我开始但无法继续前进的方式。请帮忙

new_col1 = []
new_col2 = []
def timeslot(EventCodeId, EventCode, CurrentPlaytime):
    if x == 1029 and y == 'Goal Home':
        new.Col1.append(z/(3*10**4))
    elif x == 2053 and y == 'Goal Away':
        new_col2.append(z/(3*10**4))
    else:
        new_col1.append(0)
        new_col2.append(0)
    return new_col1
    return new_col2



df1['new_col1', 'new_col2'] = df1.apply(lambda x,y,z: timeslot(x['EventCodeId'], y['EventCode'], z['CurrentPlaytime']), axis=1)  

TypeError: ("<lambda>() missing 2 required positional arguments: 'y' and 'z'", 'occurred at index 0')

【问题讨论】:

    标签: python pandas function dataframe lambda


    【解决方案1】:

    您不需要显式循环。尽可能使用矢量化操作。

    使用numpy.where:

    s = df1['CurrentPlaytime']/3*10**4
    
    mask1 = (df1['EventCodeId'] == 1029) & (df1['EventCode'] == 'Goal')
    mask2 = (df1['EventCodeId'] == 2053) & (df1['EventCode'] == 'Away')
    
    df1['new_col1'] = np.where(mask1, s, 0)
    df1['new_col2'] = np.where(mask2, s, 0)
    

    【讨论】:

    • 不错的解决方案:)
    • @jpp 感谢您抽出宝贵时间查看我的问题。您的解决方案看起来如此简单和优雅,但我收到以下错误。 TypeError: 不支持的操作数类型 /: 'str' 和 'int'
    • @A.Abs,将相关序列转换为数字,例如df1['EventCodeId'] = pd.to_numeric(df1['EventCodeId'], errors='coerce')
    • @jpp,对于每场比赛,无论是基于 MatchId 还是 xG_Team1 与 xG_Team2(连接行),我如何从 new_col1 创建 Home_Goal 和从 new_col2 创建 Away_Goal 列表?
    • @A.Abs,我建议您以separate question 的身份提问。
    猜你喜欢
    • 2019-12-23
    • 1970-01-01
    • 2023-01-14
    • 2020-12-27
    • 1970-01-01
    • 1970-01-01
    • 2019-07-02
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多