【问题标题】:How to change format of dataset using tensorflow?如何使用 tensorflow 更改数据集的格式?
【发布时间】:2018-04-02 23:38:37
【问题描述】:

有没有使用 tensorflow 库处理数据集的有效方法?我想操作它们,比如删除列/行、编辑它们等。

【问题讨论】:

  • 你的意思是numpy数组吗?使用熊猫有什么问题?
  • 实际上我是一个完全新手,我不知道如何处理在 python 中读取数据集,我在问是否有一种快速有效的方法可以通过 tensorflow 来做到这一点,但是当我检查熊猫应该像@Maxim 提到的那样做我的工作
  • 请阅读文档

标签: machine-learning tensorflow dataset data-manipulation


【解决方案1】:

此示例代码应该可以帮助您开始导入 CSV、进行一些基本操作以及在 tensorflow 中运行 NN 学习。您可以获取data here。

import numpy as np
import pandas as pd
import tensorflow as tf

# Read and manipulate data from CSV
df = pd.read_csv('df.csv')
df = df.dropna(how='any')
df = df.drop('DayOfWeek', axis=1)
df.Customers = df.Customers / 1000.0
df.CompetitionDistance = df.CompetitionDistance / 1000.0
df.Sales = df.Sales / 1000.0

# Parameters
features = 2
hidden = 3
learning_rate = 0.2

# Prepare input and output arrays
train_x = np.array(df[['CompetitionDistance', 'Customers']])
train_y = np.array(df[['Sales']]).reshape([-1])

# Build a simple TF graph
x = tf.placeholder(tf.float32, shape=[None, features], name='x')
y = tf.placeholder(tf.float32, shape=[None], name='y')
W = tf.get_variable(name='W', shape=[features, hidden])
b = tf.get_variable(name='b', shape=[hidden], initializer=tf.zeros_initializer)
z = tf.matmul(x, W) + b
predict = tf.reduce_sum(z, axis=1)
loss = tf.reduce_mean(tf.square(y - predict))
optimizer = tf.train.AdamOptimizer(learning_rate).minimize(loss)

# Run the training
with tf.Session() as session:
  session.run(tf.global_variables_initializer())
  for i in xrange(101):
    _, loss_value = session.run([optimizer, loss], 
                                feed_dict={x: train_x, y: train_y})
    if i % 10 == 0:
      print "epoch=%03i, loss=%.5f" % (i, loss_value)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-03-19
    • 1970-01-01
    • 2015-08-28
    • 2019-06-19
    • 1970-01-01
    相关资源
    最近更新 更多