【发布时间】:2020-01-06 14:45:10
【问题描述】:
我将数据存储在 pandas DataFrame 中 我使用以下代码移动到一个 numpy 数组
# used to be train_X = np.array(train_df.iloc[1:,3:].values.tolist())
# but was split for me to find he source of change
pylist = train_df.iloc[1:,3:].values.tolist()
print(pylist[0])
train_X = np.array(pylist)
print(train_X[0])
第一个打印返回:
[0.0, 0.0, 0.0, 0.0, 1.0, 504.0, 0.0, 2.0, 8.0, 0.0, 0.0, 0.0, 0.0, 2.0, 8.0, 0.0, 189.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 85143.0, 57219.0, 62511.267857142804, 2649.26669430866]
在我将它移动到 Numpy 数组之后的第二个打印返回这个
[0.00000000e+00 0.00000000e+00 0.00000000e+00 0.00000000e+00
1.00000000e+00 5.04000000e+02 0.00000000e+00 2.00000000e+00
8.00000000e+00 0.00000000e+00 0.00000000e+00 0.00000000e+00
0.00000000e+00 2.00000000e+00 8.00000000e+00 0.00000000e+00
1.89000000e+02 0.00000000e+00 0.00000000e+00 0.00000000e+00
0.00000000e+00 0.00000000e+00 0.00000000e+00 0.00000000e+00
0.00000000e+00 0.00000000e+00 8.51430000e+04 5.72190000e+04
6.25112679e+04 2.64926669e+03]
为什么会这样?以及如何阻止它
【问题讨论】:
-
数据是一样的,不用担心。 NumPy 只是在数组中的某些值超过某个阈值时将整个数组的表示转换为指数表示法。