【问题标题】:Is there a way by which we can train RNN without using one hot encoders?有没有一种方法可以在不使用热编码器的情况下训练 RNN?
【发布时间】:2019-12-11 14:08:09
【问题描述】:

我正在尝试为我的日志分析项目开发一个顺序 RNN。

输入是一个日志序列,比如 [1,2,3,4,5,6,1,5,2,7,8,2,1]

目前我正在使用 keras 库中的 to_categorical 函数,该函数将序列转换为 one-hot 编码。

def to_categorical(y, num_classes=None, dtype='float32'):
    """Converts a class vector (integers) to binary class matrix.

    E.g. for use with categorical_crossentropy.

    # Arguments
        y: class vector to be converted into a matrix
            (integers from 0 to num_classes).
        num_classes: total number of classes.
        dtype: The data type expected by the input, as a string
            (`float32`, `float64`, `int32`...)

    # Returns
        A binary matrix representation of the input. The classes axis
        is placed last.

    # Example

    ```python
    # Consider an array of 5 labels out of a set of 3 classes {0, 1, 2}:
    > labels
    array([0, 2, 1, 2, 0])
    # `to_categorical` converts this into a matrix with as many
    # columns as there are classes. The number of rows
    # stays the same.
    > to_categorical(labels)
    array([[ 1.,  0.,  0.],
           [ 0.,  0.,  1.],
           [ 0.,  1.,  0.],
           [ 0.,  0.,  1.],
           [ 1.,  0.,  0.]], dtype=float32)
    ```
    """

    y = np.array(y, dtype='int')
    input_shape = y.shape
    if input_shape and input_shape[-1] == 1 and len(input_shape) > 1:
        input_shape = tuple(input_shape[:-1])
    y = y.ravel()
    if not num_classes:
        num_classes = np.max(y) + 1
    n = y.shape[0]
    categorical = np.zeros((n, num_classes), dtype=dtype)
    categorical[np.arange(n), y] = 1
    output_shape = input_shape + (num_classes,)
    categorical = np.reshape(categorical, output_shape)
    return categorical

我面临的问题是,可能有一些日志可能不属于经过训练的数据,比如说 [9,10,11]

如果我有 2000 个日志键和 275 个唯一日志的序列。

总会有看不见的日志,但如果我想保存这个模型并在新数据上重用它,它可能无法将其转换为相同的分类格式,因为我最初只有 275 个唯一的日志类RNN 但现在我有 275+ 3 个新课程。

我们如何解决这个问题?

【问题讨论】:

  • 在不了解全部词汇的情况下,谷歌键盘预测如何工作?它还包括我们自定义词汇表中的单词。他们不使用 RNN 来预测序列吗?如果他们看到新的单词/类,他们不需要重新训练他们的模型吗?

标签: python machine-learning keras recurrent-neural-network


【解决方案1】:

关于@Dainel 对类一致性的回答,您可以将训练序列中未出现的任何值替换为np.nan 并使用pd.get_dummies,如下所示。

train_seq = np.array([1,2,3,4,5])
test_seq = np.array([1,2,3,4,5,6,7,8,9,10], dtype=np.float32)

test_seq[~np.isin(test_seq, train_seq)] = np.nan

df = pd.get_dummies(test_seq, dummy_na=True)
print(df)

它为看不见的数据生成一个单独的类。

   1.0  2.0  3.0  4.0  5.0  NaN
0    1    0    0    0    0    0
1    0    1    0    0    0    0
2    0    0    1    0    0    0
3    0    0    0    1    0    0
4    0    0    0    0    1    0
5    0    0    0    0    0    1
6    0    0    0    0    0    1
7    0    0    0    0    0    1
8    0    0    0    0    0    1
9    0    0    0    0    0    1

【讨论】:

  • 在不了解全部词汇的情况下,谷歌键盘预测如何工作?它还包括我们自定义词汇表中的单词。他们不使用 RNN 来预测序列吗?如果他们看到新的单词/类,他们不需要重新训练他们的模型吗?
  • 这是通过使用子词之类的方法来处理的。那就是你在训练语料库中学习最常见的 ngram。 (例如,lower 表示为low +er)。有了足够大的语料库,您就可以在测试语料库中表示看不见的单词。
  • 个人词典中新增的单词不属于训练语料库,如何编码。如果词汇量增加,我们使用旧词汇量的一个热向量训练的模型就没有用了。每次看到一个新词,我们都不能重新训练模型。我在网上读到的是谷歌使用有限状态传感器来生成文本,但我需要阅读更多关于它的内容
  • 这是一个非常有趣的研究课题。我认为this 论文将是一个很好的起点。这是我在之前的评论中提到的。另一种方法是让字符级翻译器在单词级翻译器中处理稀有词
  • 最接近解决方案的是使用 lucene 有限状态转换器而不是使用 RNN。它速度快,内存效率高。
【解决方案2】:

您必须具有类一致性,否则您的模型将无法正常工作。

如果数字在数值上有意义,您可以使用数字而不是 one-hot。但是既然你说它们是类,它们可能没有意义。

您可以尝试将几个训练类分离为未知类,并将它们分组为一个单一的单热编码。然后所有新类都将收到相同的编码。

但不能保证该模型会给您带来良好的结果。

【讨论】:

    猜你喜欢
    • 2019-11-16
    • 1970-01-01
    • 2020-07-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-11-22
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多