【问题标题】:LSTM input produces nans in non-text classification problemLSTM 输入在非文本分类问题中产生 nans
【发布时间】:2019-08-25 15:43:12
【问题描述】:

对非文本数据的 LSTM 模型进行训练以分类两个类别。 我对每个产品(N=730)有 225 个时间点,包括目标在内的 167 个特征。只预测最后一个时间点。 我在预测中使用目标作为特征:这是我准备输入的方式:

def split_sequences(sequences, n_steps, n_steps_out):
    X, y = list(), list()
    for i in range(n_steps_out):
        # gather input and output parts of the pattern
        y.append(sequences[n_steps + i:n_steps + i + 1, -1][0])
        #targ = sequences[n_steps + i:n_steps + i + 1, -1][0]
        #y.append(int(targ)) if ((targ==0) | (targ==1)) else y.append(2)
    X.append(sequences[:n_steps, :])
    return np.asarray(X).reshape(n_steps, sequences.shape[1]), np.asarray(y).reshape(n_steps_out)

#del X_train_minmax, X_test_minmax
min_max_scaler = preprocessing.MinMaxScaler(feature_range=(0, 1))

#X_train_minmax = min_max_scaler.fit_transform(X_train.iloc[:, 0:166])
#X_test_minmax  = min_max_scaler.fit_transform(X_test.iloc[:, 0:166])
X_train_minmax = min_max_scaler.fit_transform(X_train) ##all features included
X_test_minmax  = min_max_scaler.fit_transform(X_test)

print(X_train_minmax.shape)
print(X_test_minmax.shape) 

seq_samples = 631
seq_samples2 = 99
time_steps = 225
periods_to_predict = 1
periods_to_train = time_steps - periods_to_predict ##here may be a problem
# 
features = 167
X_train_reshaped = X_train_minmax.reshape(seq_samples,time_steps,features)
X_test_reshaped  = X_test_minmax.reshape(seq_samples2, time_steps,features)

data_train = [split_sequences(x, periods_to_train , periods_to_predict) for x in X_train_reshaped] ##and here i shoud check the function
data_test = [split_sequences(x, periods_to_train , periods_to_predict) for x in X_test_reshaped]
X_train, y_train, X_test, y_test = [], [], [], []

for x in data_train:
    X_train.append(x[0])
    y_train.append(x[1])
for x in data_test:
    X_test.append(x[0])
    y_test.append(x[1])

X_train = np.asarray(X_train)
y_train = np.asarray(y_train)
X_test = np.asarray(X_test)
y_test = np.asarray(y_test)

我对输入数据尝试了以下形状

print(X_train.shape) #(631, 224, 167)
print(X_test.shape) #(99, 224, 167)
print(y_train.shape) #(631, 1) 
print(np.unique(y_train)) #[0. 1.]
y_train_cat=to_categorical(y_train) 
print(y_train_cat.shape)  #(631, 2) 

分类模型和二元模型在预测中都会产生 nans,训练显然是错误的。很明显我错过了一些东西(我怀疑训练期间的问题=224,即最后一层的225-1或units=2)。我尝试了不同的形状和组合,但失败了,将不胜感激任何线索。

model=Sequential([
            LSTM(units=100,  
            input_shape=(periods_to_train,features), kernel_initializer='he_uniform',
            activation ='linear', kernel_constraint=maxnorm(3.), return_sequences=False),
            Dropout(rate=0.5),
            Dense(units=100,                   kernel_initializer='he_uniform', 
            activation='linear', kernel_constraint=maxnorm(3)),
            Dropout(rate=0.5),
            Dense(units=100,                  kernel_initializer='he_uniform',
            activation='linear', kernel_constraint=maxnorm(3)),
            Dropout(rate=0.5),
            Dense(units=1, kernel_initializer='he_uniform', activation='sigmoid')])

       # Compile model
optimizer = Adamax(lr=0.001, decay=0.1)
model.compile(loss='binary_crossentropy', optimizer=optimizer, metrics=['accuracy'])

configure(gpu_ind=True)
model.fit(X_train, y_train, validation_split=0.1, batch_size=100, epochs=8, shuffle=True)

_________________________________________________________________
_________________________________________________________________
Layer (type)                 Output Shape              Param #   
=================================================================
lstm_1 (LSTM)                (None, 100)               107200    
_________________________________________________________________
dropout_1 (Dropout)          (None, 100)               0         
_________________________________________________________________
dense_1 (Dense)              (None, 100)               10100     
_________________________________________________________________
dropout_2 (Dropout)          (None, 100)               0         
_________________________________________________________________
dense_2 (Dense)              (None, 100)               10100     
_________________________________________________________________
dropout_3 (Dropout)          (None, 100)               0         
_________________________________________________________________
dense_3 (Dense)              (None, 1)                 101       
=================================================================
Total params: 127,501
Trainable params: 127,501
Non-trainable params: 0
_________________________________________________________________ 

这是我预测的数组, y_hat_val = model.predict(X_test)

       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan],
       [nan], 

感谢您的帮助!

【问题讨论】:

    标签: python input data-structures lstm keras-layer


    【解决方案1】:

    运行模拟后,我发现对于这种形状的矩阵,不会导致 nans 的最大可能 time_steps (m) 为 m=163。

    修正后,模型会产生有意义的预测。 另一个需要关注的问题是输入训练集的准备。 如果使用return_sequences 参数,训练集 应该包括实际的 N 个 time_steps 而不是 N-1,如示例中所示。 下面是如何转换火车组

    X_train_minmax = min_max_scaler.fit_transform(X_train) ##all features included
    X_train_reshaped = X_train_minmax.reshape(seq_samples,time_steps,features)
    
    

    【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2018-01-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-11-13
    • 1970-01-01
    • 1970-01-01
    • 2019-09-27
    相关资源
    最近更新 更多