【发布时间】:2021-12-06 15:32:34
【问题描述】:
我有:
-
大约 40k 双/三字词的位置列表。
['San Francisco CA', 'Oakland CA', 'San Diego CA',...] -
具有数百万行的 Pandas DataFrame。
| string_column | string_column_location_removed |
|---|---|
| Burger King Oakland CA | Burger King |
| Walmart Walnut Creek CA | Walmart |
我目前正在遍历位置列表,如果该位置存在于 string_column 中,则创建一个新列 string_column_location_removed 并删除该位置。
这是我的尝试,虽然有效,但速度很慢。有关如何加快速度的任何想法?
我尝试从 this 和 this 中获取想法,但不确定如何使用 Pandas 数据框真正推断这一点。
from random import choice
from string import ascii_lowercase, digits
import pandas
#making random list here
chars = ascii_lowercase + digits
locations_lookup_list = [''.join(choice(chars) for _ in range(10)) for _ in range(40000)]
locations_lookup_list.append('Walnut Creek CA')
locations_lookup_list.append('Oakland CA')
strings_for_df = ["Burger King Oakland CA", "Walmart Walnut Creek CA",
"Random Other Thing Here", "Another random other thing here", "Really Appreciate the help on this", "Thank you so Much!"] * 250000
df = pd.DataFrame(strings_for_df)
def location_remove(txnString):
for locationString in locations_lookup_list:
if re.search(f'\\b{locationString}\\b', txnString):
return re.sub(f'\\b{locationString}\\b','', txnString)
else:
continue
df['string_column_location_removed'] = df['string_column'].apply(lambda x: location_remove(x))
【问题讨论】:
-
This 是如果你有一个单词/短语列表来匹配整个单词的方式。为什么不简单地将它与
Series.str.replace一起使用?您不需要运行两次正则表达式,if re.search(f'\\b{locationString}\\b', txnString):检查是否多余。
标签: python regex pandas performance replace