【发布时间】:2020-09-10 13:18:56
【问题描述】:
我想使用表示允许的数据输入格式的正则表达式字典来清理 pandas 数据框。
我正在尝试遍历输入数据框,以便根据给定列的允许数据输入格式检查每一行。
如果条目不符合列允许的格式,我想用 NaN 替换它(请参阅下面的所需输出)。
我当前的代码给了我一条错误消息:“DataFrame”对象没有属性“col”。
我的 MWE 有两个有代表性的正则表达式,但对于我的实际数据集,我有大约 40 个。
感谢您的帮助!
# Packages
import pandas as pd
import re
import numpy as np
# Input data frame
data = {'score': [71,72,55,'a'],
'bet': [0.260,0.380,'0.8dd',0.260]
}
df1 = pd.DataFrame(data, columns = ['score', 'bet'])
# Input dictionary
dict1 = {'score':'^\d+$',
'bet': '^\d[\.]\d+$'}
# Cleaning function
def cleaner(df, dict):
for col in df.columns:
if col in dict:
for row in df.col:
if re.match(dict[col], str(row)):
row = row
else:
row = np.nan
return(df)
cleaned_df = cleaner(df1, dict1)
# ERROR MESSAGE
# 'DataFrame' object has no attribute 'col'
# Desired output
goal_data = {'score': [71,72,55, np.nan],
'bet': [0.260,0.380, np.nan, 0.260]
}
goal_df = pd.DataFrame(goal_data, columns = ['score', 'bet'])
【问题讨论】:
标签: python regex pandas data-science data-cleaning