【发布时间】:2020-06-14 05:53:29
【问题描述】:
我有一个包含 13 个特征和 1000 万行的数据库。我想应用 k-mean 来消除任何异常。我的想法是应用k-mean,创建一个包含数据点和聚类质心之间距离的新列,以及一个包含平均距离的新列,如果距离大于平均距离,我删除整行。但是我写的代码好像不行。
数据集示例: https://drive.google.com/open?id=1iB1qjnWQyvoKuN_Pa8Xk4BySzXVTwtUk
df = pd.read_csv('Final After Simple Filtering.csv',index_col=None,low_memory=True)
# Dropping columns with low feature importance
del df['AmbTemp_DegC']
del df['NacelleOrientation_Deg']
del df['MeasuredYawError']
#applying kmeans
#applying kmeans
kmeans = KMeans( n_clusters=8)
clusters= kmeans.fit_predict(df)
centroids = kmeans.cluster_centers_
distance1 = kmeans.fit_transform(df)
distance2 = distance1.mean()
df['distances']=distance1-distance2
df = df[df['distances'] >=0]
del df['distances']
df.to_csv('/content//drive/My Drive/K TEST.csv', index=False)
错误:
KeyError Traceback (most recent call last)
/usr/local/lib/python3.6/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
2896 try:
-> 2897 return self._engine.get_loc(key)
2898 except KeyError:
pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()
pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()
pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
KeyError: 'distances'
During handling of the above exception, another exception occurred:
KeyError Traceback (most recent call last)
9 frames
pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()
pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()
pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
KeyError: 'distances'
During handling of the above exception, another exception occurred:
ValueError Traceback (most recent call last)
/usr/local/lib/python3.6/dist-packages/pandas/core/internals/blocks.py in __init__(self, values, placement, ndim)
126 raise ValueError(
127 "Wrong number of items passed {val}, placement implies "
--> 128 "{mgr}".format(val=len(self.values), mgr=len(self.mgr_locs))
129 )
130
ValueError: Wrong number of items passed 8, placement implies 1
谢谢
【问题讨论】:
-
您能告诉我们您遇到了什么错误吗?
-
@Ehrendil 我已在主帖中发布了错误。
-
我建议也发布一个数据框样本
-
我已将我的数据集样本添加到主帖
标签: python pandas dataframe machine-learning jupyter-notebook