【发布时间】:2019-06-27 20:23:45
【问题描述】:
这是我的数据集
id descriptions
0 kartu debit 20 10 indomaretcipete r
1 tarikan atm 20 10
2 tarikan atm 19 10
3 biaya adm
4 trsf 18 10 wsid 23881 indah lestari
这就是我所做的
def cleaning(text):
stops = {'10', '18','19', '20', '23881'}
text = [word for word in text if not word in stops]
text = " ".join(text)
return(text)
df['description_clean'] = df['description'].apply(cleaning)
这就是我得到的
id descriptions
0 kartu debit indomaretcipete r
1 tarikan atm
2 tarikan atm
3 biaya adm
4 trsf wsid indah lestari
这个效果不好,我一直在添加新的数字来改进停用词,一次怎么办?
【问题讨论】:
标签: python regex pandas dataframe