你可以使用:
out = (
df['Comments'].str.split().explode().to_frame('Word').join(df['Label']).assign(value=1) \
.pivot_table('value', 'Word', 'Label', aggfunc='count', fill_value=0) \
.assign(Count=lambda x: x.sum(axis=1))
)
输出:
>>> out
Label Negative Positive Suggestions Count
Word
I 1 1 0 2
Teammates 0 1 0 1
We 0 0 1 1
boss 1 0 0 1
hate 1 0 0 1
higher 0 0 1 1
love 0 1 0 1
my 1 1 0 2
need 0 0 1 1
pay 0 0 1 1
详情:
第 1 步。 将每个评论分解为单词并分配一个虚拟值。
out = df['Comments'].str.split().explode().to_frame('Word').join(df['Label']).assign(value=1)
print(out)
# Output:
Word Label value
0 I Positive 1
0 love Positive 1
0 my Positive 1
0 Teammates Positive 1
1 We Suggestions 1
1 need Suggestions 1
1 higher Suggestions 1
1 pay Suggestions 1
2 I Negative 1
2 hate Negative 1
2 my Negative 1
2 boss Negative 1
第 2 步。 旋转您的数据框。
out = out.pivot_table('value', 'Word', 'Label', aggfunc='count', fill_value=0)
print(out)
# Output:
Label Negative Positive Suggestions
Word
I 1 1 0
Teammates 0 1 0
We 0 0 1
boss 1 0 0
hate 1 0 0
higher 0 0 1
love 0 1 0
my 1 1 0
need 0 0 1
pay 0 0 1
第 3 步。:创建 count 列。
out = out.assign(Count=lambda x: x.sum(axis=1))
print(out)
# Output:
Label Negative Positive Suggestions Count
Word
I 1 1 0 2
Teammates 0 1 0 1
We 0 0 1 1
boss 1 0 0 1
hate 1 0 0 1
higher 0 0 1 1
love 0 1 0 1
my 1 1 0 2
need 0 0 1 1
pay 0 0 1 1