【发布时间】:2021-05-18 11:07:38
【问题描述】:
我有一个 7300 万行数据集,我需要过滤掉与几个条件中的任何一个匹配的行。我一直在使用布尔索引进行此操作,但这需要很长时间(约 30 分钟),我想知道是否可以使其更快(例如花式索引、np.where、np.compress?)
我的代码:
clean_df = df[~(df.project_name.isin(p_to_drop) |
df.workspace_name.isin(ws_to_drop) |
df.campaign_name.str.contains(regex_string,regex=True) |
df.campaign_name.isin(small_launches))]
正则表达式字符串是
regex_string = '(?i)^.*ARCHIVE.*$|^.*birthday.*$|^.*bundle.*$|^.*Competition followups.*$|^.*consent.*$|^.*DOI.*$|\
^.*experiment.*$|^.*hello.*$|^.*new subscribers.*$|^.*not purchased.*$|^.*parent.*$|\
^.*re engagement.*$|^.*reengagement.*$|^.*re-engagement.*$|^.*resend.*$|^.*Resend of.*$|\
^.*reward.*$|^.*survey.*$|^.*test.*$|^.*thank.*$|^.*welcome.*$'
其他三个条件是少于 50 项的字符串列表。
【问题讨论】:
-
请在您的问题中提及
regex_string的价值,谢谢。 -
我已经添加了它 - 它非常广泛
-
让我修好然后回到这里。
-
不要将正则表达式过滤器应用于大数据框,使用其他 3 个条件首先制作临时 df,然后在那个小 temp_df 上使用正则表达式条件,因为正则表达式操作成本很高。
-
@travelsandbooks,这里
.*不是必需的,您可以简化正则表达式模式,这可以提高速度/性能。尝试制作如下列表:words = ['ARCHIVE', 'birthday', 'bundle', 'Competition followups', 'consent', 'DOI', 'experiment', 'hello', 'new subscribers', 'not purchased', 'parent', 're engagement', 'reengagement', 're-engagement', 'resend', 'Resend of', 'reward', 'survey', 'test', 'thank', 'welcome']然后制作regex_string = r'(?i)\b(' + '|'.join(words) +r')\b'然后尝试运行代码一次,但又一次,因为我没有那么多大数据,所以无法测试它。
标签: python pandas string-operations boolean-indexing