【问题标题】:Python pandas data frame clean with dictionary of regular expressions使用正则表达式字典清理 Python pandas 数据框
【发布时间】:2020-09-10 13:18:56
【问题描述】:

我想使用表示允许的数据输入格式的正则表达式字典来清理 pandas 数据框。

我正在尝试遍历输入数据框,以便根据给定列的允许数据输入格式检查每一行。

如果条目不符合列允许的格式,我想用 NaN 替换它(请参阅下面的所需输出)。

我当前的代码给了我一条错误消息:“DataFrame”对象没有属性“col”。

我的 MWE 有两个有代表性的正则表达式,但对于我的实际数据集,我有大约 40 个。

感谢您的帮助!

# Packages 
import pandas as pd
import re
import numpy as np 


# Input data frame 
data = {'score': [71,72,55,'a'],
        'bet': [0.260,0.380,'0.8dd',0.260]
        }
df1 = pd.DataFrame(data, columns = ['score', 'bet'])


# Input dictionary 
dict1 = {'score':'^\d+$', 
         'bet': '^\d[\.]\d+$'}


# Cleaning function  
def cleaner(df, dict):
    for col in df.columns:
        if col in dict: 
            for row in df.col:
                if re.match(dict[col], str(row)):
                    row = row 
                else:
                    row = np.nan
    return(df)


cleaned_df = cleaner(df1, dict1)


# ERROR MESSAGE 
# 'DataFrame' object has no attribute 'col'


# Desired output 
goal_data = {'score': [71,72,55, np.nan],
        'bet': [0.260,0.380, np.nan, 0.260]
        }
goal_df = pd.DataFrame(goal_data, columns = ['score', 'bet'])

【问题讨论】:

    标签: python regex pandas data-science data-cleaning


    【解决方案1】:

    您的 if 语句中的清理功能有问题。 尝试运行以下清洁功能代替您的功能。

    # Cleaning function  
    def cleaner(df, dict):
        for col in df.columns:
            if col in dict.keys(): 
                for row in df.index:
                    if type(re.match(dict[col], str(df[col][row]))) is re.Match:
                        df[col][row] = df[col][row] 
                        print(df[col][row])
                    else:
                        df[col][row] = np.nan
        return(df)
    print(cleaner(df1, dict1))
    cleaned_df = cleaner(df1, dict1)
    

    【讨论】:

      【解决方案2】:

      试试np.where(if condition, yes,else alternative)

      import pandas as pd
      import numpy as np
      df1['score']=np.where(df1.score.str.match('^\d+$'),df1['score'],np.nan)
      df1['bet']=np.where(df1.bet.str.match('^\d[\.]\d+$'),df1['bet'],np.nan)
      
      score   bet
      0   71  0.26
      1   72  0.38
      2   55  NaN
      3   NaN 0.26
      

      【讨论】:

        猜你喜欢
        • 2021-02-04
        • 2018-06-04
        • 2020-07-03
        • 2021-04-09
        • 2019-08-22
        • 2010-10-31
        • 1970-01-01
        相关资源
        最近更新 更多