【问题标题】:How do I drop all numbers as data cleansing on pandas effectively?如何有效地删除所有数字作为 pandas 上的数据清理?
【发布时间】:2019-06-27 20:23:45
【问题描述】:

这是我的数据集

id                                             descriptions
0                       kartu debit 20 10 indomaretcipete r
1                                         tarikan atm 20 10
2                                         tarikan atm 19 10
3                                                 biaya adm
4                       trsf 18 10 wsid 23881 indah lestari

这就是我所做的

def cleaning(text):
    stops = {'10', '18','19', '20', '23881'}
    text = [word for word in text if not word in stops]
    text = " ".join(text)
return(text)

df['description_clean'] = df['description'].apply(cleaning)

这就是我得到的

  id                                              descriptions
  0                             kartu debit indomaretcipete r
  1                                               tarikan atm
  2                                               tarikan atm
  3                                                 biaya adm
  4                                   trsf wsid indah lestari

这个效果不好,我一直在添加新的数字来改进停用词,一次怎么办?

【问题讨论】:

    标签: python regex pandas dataframe


    【解决方案1】:

    IIUC,您需要从数据框中删除数字,使用如下:

    df_new=df.replace('\d+ ','',regex=True)
    print(df_new)
    
       id                   descriptions
    0   0  kartu debit indomaretcipete r
    1   1                 tarikan atm 10
    2   2                 tarikan atm 10
    3   3                      biaya adm
    4   4        trsf wsid indah lestari
    

    对于一个系列:df['descriptions']=df['descriptions'].replace('\d+ ','',regex=True)

    注意:根据您的示例,我在正则表达式中的 d+ 之后添加了一个空格,如果您愿意,可以不使用它。

    【讨论】:

      【解决方案2】:

      使用str.extractall 和groupby.agg:

      df['descriptions'] = (df['descriptions'].str.extractall('([a-zA_Z]+)')
                                              .groupby(level=0).agg({0:' '.join}))
      

      或者:

      df['descriptions'] = (df['descriptions'].str.replace('\d+','')
                                              .str.replace('  ',''))
      

      或者:

      df['descriptions'] = [' '.join(re.findall('[a-zA-Z]+',s)) for s in df['descriptions']]
      

      print(df)
         id                   descriptions
      0   0  kartu debit indomaretcipete r
      1   1                    tarikan atm
      2   2                    tarikan atm
      3   3                      biaya adm
      4   4        trsf wsid indah lestari
      

      【讨论】:

      • 目前最好的答案,但是太慢了
      • @NabihBawazir 大熊猫中的字符串操作通常很慢。检查其他解决方案是否加快了进程。
      【解决方案3】:

      你需要:

      def replace_numbers(s):
          return re.sub(r'\d*', '', s)
      
      
      df['description'] = df['description'].apply(replace_numbers)
      

      【讨论】:

        猜你喜欢
        • 2021-06-18
        • 2017-01-03
        • 1970-01-01
        • 2012-05-29
        • 2011-01-15
        • 1970-01-01
        • 1970-01-01
        • 2019-11-05
        • 2022-06-13
        相关资源
        最近更新 更多