【问题标题】:Replacing strings in a a column of similar categories, map to new column in python替换相似类别列中的字符串,映射到python中的新列
【发布时间】:2018-09-10 11:12:37
【问题描述】:

我有一个现有的数据框 (coffee_directions_df),如下所示

coffee_directions_df

Utterance                         Frequency   

Directions to Starbucks           1045
Directions to Tullys              1034
Give me directions to Tullys      986
Directions to Seattles Best       875
Show me directions to Dunkin      812
Directions to Daily Dozen         789
Show me directions to Starbucks   754
Give me directions to Dunkin      612
Navigate me to Seattles Best      498
Display navigation to Starbucks   376
Direct me to Starbucks            201

DF 显示人们发出的话语和话语的频率。

即“去星巴克的路线”被说出了 1045 次。

我试图弄清楚如何将coffee_directions_df.Utterance 列中的类似词(例如“Starbucks”、“Tullys”、“Seattles Best”)替换为一个字符串,例如“Coffee”。我看过类似的答案,建议使用字典,如下所示,但我还没有成功。

{'Utterance':['Starbucks','Tullys','Seattles Best'],
      'Combi_Utterance':['Coffee','Coffee','Coffee','Coffee']}

{'Utterance':['Dunkin','Daily Dozen'],
      'Combi_Utterance':['Donut','Donut']}

{'Utterance':['Give me','Show me','Navigate me','Direct me'],
      'Combi_Utterance':['V_me','V_me','V_me','V_me']}

想要的输出如下:

coffee_directions_df

Utterance                         Frequency  Combi_Utterance
Directions to Starbucks           1045       Directions to Coffee
Directions to Tullys              1034       Directions to Coffee
Give me directions to Tullys      986        V_me to Coffee
Directions to Seattles Best       875        Directions to Coffee
Show me directions to Dunkin      812        V_me to Donut
Directions to Daily Dozen         789        Directions to Donut
Show me directions to Starbucks   754        V_me to Coffee
Give me directions to Dunkin      612        V_me to Donut
Navigate me to Seattles Best      498        V_me to Coffee
Display navigation to Starbucks   376        Display navigation to Coffee
Direct me to Starbucks            201        V_me to Coffee

最终,我希望能够使用我必须生成最终输出的代码。

df = (df.set_index('Frequency')['Utterance']
        .str.split(expand=True)
        .stack()
        .reset_index(name='Words')
        .groupby('Words', as_index=False)['Frequency'].sum()
        )

print (df)
         Words  Frequency
0   Directions       6907
1         V_me       3863
2        Donut       2213
3       Coffee       5769
4        Other        376

谢谢!!

【问题讨论】:

    标签: python pandas dataframe statistics apply


    【解决方案1】:

    下面是一种方法。根据您之前的问题,我选择使用collections.Counter 而不是pandas 作为您的计数逻辑。

    所需的输入是映射字典rep_dict 的形式。我们将此应用于df['Utterance'] 系列中的字符串子字符串。

    from collections import Counter
    import pandas as pd
    
    df = pd.DataFrame([['Directions to Starbucks', 1045],
                       ['Show me directions to Starbucks', 754],
                       ['Give me directions to Starbucks', 612],
                       ['Navigate me to Starbucks', 498],
                       ['Display navigation to Starbucks', 376],
                       ['Direct me to Starbucks', 201],
                       ['Navigate to Starbucks', 180]],
                      columns=['Utterance', 'Frequency'])
    
    # define dictionary of mappings
    rep_dict = {'Starbucks': 'Coffee', 'Tullys': 'Coffee', 'Seattles Best': 'Coffee'}
    
    # apply substring mapping
    df['Utterance'] = df['Utterance'].replace(rep_dict, regex=True).str.lower()
    
    # previous logic below
    c = Counter()
    
    for row in df.itertuples():
        for i in row[1].split():
            c[i] += row[2]
    
    res = pd.DataFrame.from_dict(c, orient='index')\
                      .rename(columns={0: 'Count'})\
                      .sort_values('Count', ascending=False)
    
    def add_combinations(df, lst):
        for i in lst:
            words = '_'.join(i)
            df.loc[words] = df.loc[df.index.isin(i), 'Count'].sum()
        return df.sort_values('Count', ascending=False)
    
    lst = [('give', 'show', 'navigate', 'direct')]
    
    res = add_combinations(res, lst)
    

    结果

                               Count
    to                          3666
    coffee                      3666
    directions                  2411
    give_show_navigate_direct   2245
    me                          2065
    show                         754
    navigate                     678
    give                         612
    display                      376
    navigation                   376
    direct                       201
    

    【讨论】:

    • 嗨,如果你有时间,请看我的下一个问题。永远感谢您的帮助! (按照你的步骤,但我也在尝试做其他事情)。谢谢! stackoverflow.com/questions/49656991/…
    • @user_seaweed,如果它有效,请接受这个答案(左边的绿色勾号)。
    • 谢谢,抱歉,这还是个新手。基本上是在寻找一种新的方法来使用更大的数据框来做到这一点。 (理想情况下,希望将 rep_dict 导入为 txt 文件,而不是将其全部写在我的 shell/终端中)。谢谢!
    猜你喜欢
    • 1970-01-01
    • 2017-06-03
    • 1970-01-01
    • 2021-09-23
    • 2012-01-15
    • 2018-03-25
    • 1970-01-01
    • 2023-03-26
    相关资源
    最近更新 更多