【发布时间】:2020-04-23 13:33:21
【问题描述】:
我有一个分类的系列。
目前我正在使用以下代码映射到字符串。
import pandas as pd
import numpy as np
test = np.random.rand(int(5e6))
test[0] = np.nan
test_cut = pd.cut(test,(-np.inf,0.2,0.4,np.inf))
test_str = test_cut.astype('str')
test_str[test_str.isna()] = 'missing'
这个 astype('str') 操作很慢,有没有办法加快速度?
根据下面的链接,我了解到 apply 比 astype 快。我尝试了以下方法。
test_str = test_cut.apply(str)
#AttributeError: 'Categorical' object has no attribute 'apply'
test_str = test_cut.map(str)
# still categorical type
test_str = test_cut.values.astype(str)
# AttributeError: 'Categorical' object has no attribute 'values'
Converting a series of ints to strings - Why is apply much faster than astype?
我不关心类别的确切字符串表示,只关心组被保留并转换为字符串。
作为替代方案,有没有办法在 test_cut 分类“Missing”(或其他)中定义一个新类别,并将“test”中的“missing”案例设置为该类别?
# some code to create 'MISSING' category
test_cat[test_str.isna()] = 'MISSING'
【问题讨论】:
标签: python pandas optimization categories