【发布时间】:2018-02-07 01:01:23
【问题描述】:
List top_brands 包含品牌列表,例如
top_brands = ['Coca Cola', 'Apple', 'Victoria\'s Secret', ....]
items 是pandas.DataFrame,其结构如下所示。如果缺少brand_name,我的任务是从item_title 中填写brand_name
row item_title brand_name
1 | Apple 6S | Apple
2 | New Victoria\'s Secret | missing <-- need to fill with Victoria\'s Secret
3 | Used Samsung TV | missing <--need fill with Samsung
4 | Used bike | missing <--No need to do anything because there is no brand_name in the title
....
我的代码如下。问题在于 对于包含 200 万条记录的数据框来说太慢了。有什么方法可以使用 pandas 或 numpy 来处理任务?
def get_brand_name(row):
if row['brand_name'] != 'missing':
return row['brand_name']
item_title = row['item_title']
for brand in top_brands:
brand_start = brand + ' '
brand_in_between = ' ' + brand + ' '
brand_end = ' ' + brand
if ((brand_in_between in item_title) or item_title.endswith(brand_end) or item_title.startswith(brand_start)):
print(brand)
return brand
return 'missing' ### end of get_brand_name
items['brand_name'] = items.apply(lambda x: get_brand_name(x), axis=1)
【问题讨论】:
-
几个问题:我们可以假设没有重叠的品牌名称,例如“苹果”和“苹果公司”。在您的示例中,brand_name = VS,您是如何获得 VS 缩写的?
-
场景中没有VS缩写。我不应该把它放在那里。并且有重叠的品牌名称,但为了性能,我们可以假设没有重叠