【发布时间】:2019-02-26 02:59:02
【问题描述】:
我正在通过我找到的免费在线资源进行一些数据挖掘自学。基本上,我得到了一个 csv 文件,其中包含一堆名称、电影名称以及每个人对它的评价。我正在尝试使用余弦度量从中获取 K-Nearest Neighbor,但我无法让输出看起来不那么糟糕。这是我到目前为止的代码:
from pandas import DataFrame
import pandas as pd
import numpy as np
from sklearn.neighbors import NearestNeighbors as nn
df = pd.read_csv("https://docs.google.com/spreadsheets/d/1MSBm3M6YmaLf0aiJCvkvrPsIJB2pPuBwse5ylnzEHRI/pub?gid=639849687&single=true&output=csv",index_col='Unnamed: 0')
df = df.fillna(0)
nn([df], metric = 'cosine')
做起来很简单!除了我的输出看起来像这样:
NearestNeighbors(algorithm='auto', leaf_size=30, metric='cosine',
metric_params=None, n_jobs=1,
n_neighbors=[ Patrick C Heather Bryan
Patrick T Thomas aaron \
Alien NaN NaN 2.0 NaN 5.0
4.0
Avatar 4.0 5.0 5.0 4.0 2.0 NaN
Blade Runner 5.0 NaN NaN N...
You Got Mail NaN 2.0 2.0 1.0 2.0 NaN 2.0
[25 rows x 25 columns]],
p=2, radius=1.0)
它很乱,甚至没有显示所有数据。我尝试将其转换为数组,但出现错误消息“'ABCMeta' 对象不支持索引”
我对 Python 还很陌生,我可以做一些基本的事情,但我不是专家。我希望有人可以帮助我朝着帮助清理这个方向前进。
谢谢。
【问题讨论】:
标签: python python-3.x pandas scikit-learn sklearn-pandas