【问题标题】:Run a Principal Component Analysis (PCA) on the dataset to reduce the number of features (components) from 64 to 2对数据集运行主成分分析 (PCA),以将特征(成分)的数量从 64 个减少到 2 个
【发布时间】:2018-12-22 04:26:18
【问题描述】:

我正在尝试将组件减少到 2 个而不是 64 个,但我不断收到此错误: “长度不匹配:预期轴有 64 个元素,新值有 4 个元素” 为什么我在数据集上运行的 PCA 没有将数字更改为 2?

这就是我所拥有的:

import matplotlib.pyplot as plt
from sklearn import datasets
from sklearn.cluster import KMeans
import sklearn.metrics as sm

import pandas as pd
import numpy as np


import scipy
from sklearn import decomposition

digits = datasets.load_digits()      #load the digits dataset instead of the iris dataset


x = pd.DataFrame(digits.data)     #was(iris.data)
x.columns = ['Sepal_L', 'Sepal_W', 'Sepal_L', 'Sepal_W']

plt.cla()
pca = decomposition.PCA(n_components=2)
pca.fit(x)
x = pca.transform(x)


y = pd.DataFrame(digits.target)
y.columns = ['Targets']

# this line actually builds the machine learning model and runs the algorithm
# on the dataset
model = KMeans(n_clusters = 10)    #Run k-means on this datatset to cluster the data into 10 classes
model.fit(x)

#print(model.labels_)


colormap = np.array(['red', 'blue', 'yellow', 'black'])

# Plot the Models Classifications
plt.subplot(1, 2, 2)
plt.scatter(x.Petal_L, x.Petal_W, c=colormap[model.labels_], s=40)
plt.title('K Means Classification')

plt.show()

【问题讨论】:

  • @sacul 哦,这有道理,你知道我应该如何修复我的列以代替用于数字数据集吗?
  • 在下面查看我的答案。

标签: python pandas scipy scikit-learn


【解决方案1】:

实际上不是 PCA 有问题,而只是重命名了您的列:digits 数据集有 64 列,您正在尝试根据 @987654324 中 4 列的列名来命名列@数据集。

由于数字数据集(像素)的性质,没有真正适合列的命名方案。所以不要重命名它们。

digits = datasets.load_digits()      

x = pd.DataFrame(digits.data)     

pca = decomposition.PCA(n_components=2)
pca.fit(x)
x = pca.transform(x)

# Here is the result of your PCA (2 components)
>>> x
array([[ -1.25946636,  21.27488332],
       [  7.95761139, -20.76869904],
       [  6.99192268,  -9.9559863 ],
       ..., 
       [ 10.80128366,  -6.96025224],
       [ -4.87210049,  12.42395326],
       [ -0.34438966,   6.36554934]])

如果这就是你想要的(我从你的代码中收集的),那么你可以绘制第一台电脑和第二台电脑的对比图

plt.scatter(x[:,0], x[:,1], s=40)
plt.show()

【讨论】:

  • 谢谢。我今天刚被介绍到数据集,所以它对我来说真的很新鲜。
  • 虽然,在执行 plt.scatter(x[:,0], x[:,1], c=colormap[model.labels_], s=40) 之后,我不断收到错误“索引8 超出轴 0 的范围,尺寸为 8"
  • 很确定这是您的颜色图的问题,但我不确定您要实现什么 TBH
猜你喜欢
  • 2017-04-09
  • 2020-06-05
  • 1970-01-01
  • 2012-11-05
  • 2021-05-10
  • 1970-01-01
  • 2012-11-10
  • 2013-04-21
  • 2012-12-24
相关资源
最近更新 更多