【问题标题】:How to speed up pandas boolean indexing with multiple string conditions如何使用多个字符串条件加速熊猫布尔索引
【发布时间】:2021-05-18 11:07:38
【问题描述】:

我有一个 7300 万行数据集,我需要过滤掉与几个条件中的任何一个匹配的行。我一直在使用布尔索引进行此操作,但这需要很长时间(约 30 分钟),我想知道是否可以使其更快(例如花式索引、np.where、np.compress?)

我的代码:

clean_df = df[~(df.project_name.isin(p_to_drop) | 
                df.workspace_name.isin(ws_to_drop) | 
                df.campaign_name.str.contains(regex_string,regex=True) | 
                df.campaign_name.isin(small_launches))]

正则表达式字符串是

regex_string = '(?i)^.*ARCHIVE.*$|^.*birthday.*$|^.*bundle.*$|^.*Competition followups.*$|^.*consent.*$|^.*DOI.*$|\
                    ^.*experiment.*$|^.*hello.*$|^.*new subscribers.*$|^.*not purchased.*$|^.*parent.*$|\
                    ^.*re engagement.*$|^.*reengagement.*$|^.*re-engagement.*$|^.*resend.*$|^.*Resend of.*$|\
                    ^.*reward.*$|^.*survey.*$|^.*test.*$|^.*thank.*$|^.*welcome.*$'

其他三个条件是少于 50 项的字符串列表。

【问题讨论】:

  • 请在您的问题中提及regex_string 的价值,谢谢。
  • 我已经添加了它 - 它非常广泛
  • 让我修好然后回到这里。
  • 不要将正则表达式过滤器应用于大数据框,使用其他 3 个条件首先制作临时 df,然后在那个小 temp_df 上使用正则表达式条件,因为正则表达式操作成本很高。
  • @travelsandbooks,这里.* 不是必需的,您可以简化正则表达式模式,这可以提高速度/性能。尝试制作如下列表:words = ['ARCHIVE', 'birthday', 'bundle', 'Competition followups', 'consent', 'DOI', 'experiment', 'hello', 'new subscribers', 'not purchased', 'parent', 're engagement', 'reengagement', 're-engagement', 'resend', 'Resend of', 'reward', 'survey', 'test', 'thank', 'welcome'] 然后制作 regex_string = r'(?i)\b(' + '|'.join(words) +r')\b' 然后尝试运行代码一次,但又一次,因为我没有那么多大数据,所以无法测试它。

标签: python pandas string-operations boolean-indexing


【解决方案1】:

如果您有这么多行,我认为先一步一步删除记录会更快。正则表达式通常很慢,因此您可以将其用作最后一步,数据框要小得多。

例如:

clean_df = df.copy()
clean_df = clean_df.loc[~(df.project_name.isin(p_to_drop)]
clean_df = clean_df.loc[~df.workspace_name.isin(ws_to_drop)]
clean_df = clean_df.loc[~df.campaign_name.isin(small_launches)]
clean_df = clean_df.loc[~df.campaign_name.str.contains(regex_string,regex=True)]

【讨论】:

    【解决方案2】:

    我曾认为链接我的条件是一个好主意,但让它们连续的答案帮助我重新思考:每次我运行布尔索引操作时,我都在缩小数据集 - 因此下一次操作更便宜。

    按照建议,我已将它们分开,并将删除最多行的操作放在顶部,这样接下来的操作会更快。我把正则表达式放在最后 - 因为它很昂贵,所以在尽可能小的 df 上做它是有意义的。

    希望这对某人有所帮助! TIL 链接您的操作看起来不错,但效率不高:)

    【讨论】:

    • 感谢分享,能否请您告诉我我的正则表达式(以 cmets 给出)是否适合您?
    • 嗨,我用过它,但我不确定它是否有很大的不同。它也没有抓住一切 - 例如。当单词位于字符串的开头或结尾时。
    猜你喜欢
    • 1970-01-01
    • 2021-05-17
    • 2018-01-13
    • 1970-01-01
    • 1970-01-01
    • 2017-09-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多