【发布时间】:2016-02-06 17:42:16
【问题描述】:
我有一个作业,我应该使用 scikit、numpy 和 pylab 来执行以下操作:
"以下所有内容都应使用来自 training_data.csv 文件的数据 假如。 training_data 为您提供一组标记的整数对, 代表两个运动队的分数,标签给出 运动。
编写以下函数:
plot_scores() 应该绘制数据的散点图。
predict(dataset) 应该产生一个训练有素的 Estimator 来猜测这项运动 这导致了一个给定的分数(来自我们保留的数据集,这将 作为 1000 x 2 np 数组的输入)。您可以使用 scikit 中的任何算法。
一个称为“预处理”的可选附加功能将处理数据集 在我们通过它来预测之前。 "
这是我到目前为止所做的:
import numpy as np
import scipy as sp
import pylab as pl
from random import shuffle
def plot_scores():
k=open('training_data.csv')
lst=[]
for triple in k:
temp=triple.split(',')
lst.append([int(temp[0]), int(temp[1]), int(temp[2][:1])])
array=np.array(lst)
pl.scatter(array[:,0], array[:,1])
pl.show()
def preprocess(dataset):
k=open('training_data.csv')
lst=[]
for triple in k:
temp=triple.split(',')
lst.append([int(temp[0]), int(temp[1]), int(temp[2][:1])])
shuffle(lst)
return lst
在预处理中,我对数据进行了混洗,因为我应该使用其中的一些来训练和测试,但原始数据根本不是随机的。我的问题是,我应该如何在预测(数据集)中“产生训练有素的估计器”?这应该是一个返回另一个函数的函数吗?哪种算法最适合根据如下所示的数据集进行分类:
【问题讨论】:
标签: algorithm python-2.7 numpy machine-learning scikit-learn