【问题标题】:Why are dot products backwards in Stanford's cs231n SVM?为什么斯坦福的 cs231n SVM 中的点积是倒退的?
【发布时间】:2019-04-13 14:47:40
【问题描述】:

我正在观看斯坦福大学 cs231n 的 Youtube 视频,并尝试将作业作为练习来完成。在执行 SVM 时,我遇到了以下代码:

def svm_loss_naive(W, X, y, reg):
  """
  Structured SVM loss function, naive implementation (with loops).

  Inputs have dimension D, there are C classes, and we operate on minibatches
  of N examples.

  Inputs:
  - W: A numpy array of shape (D, C) containing weights.
  - X: A numpy array of shape (N, D) containing a minibatch of data.
  - y: A numpy array of shape (N,) containing training labels; y[i] = c means
    that X[i] has label c, where 0 <= c < C.
  - reg: (float) regularization strength

  Returns a tuple of:
  - loss as single float
  - gradient with respect to weights W; an array of same shape as W
  """
  dW = np.zeros(W.shape) # initialize the gradient as zero

  # compute the loss and the gradient
  num_classes = W.shape[1]
  num_train = X.shape[0]
  loss = 0.0
  for i in range(num_train):
    scores = X[i].dot(W) # This line
    correct_class_score = scores[y[i]]
    for j in range(num_classes):
      if j == y[i]:
        continue
      margin = scores[j] - correct_class_score + 1 # note delta = 1
      if margin > 0:
        loss += margin

这是我遇到问题的线路:

scores = X[i].dot(W) 

这是在做乘积xW,不应该是Wx吗?我的意思是W.dot(X[i])

【问题讨论】:

    标签: python math machine-learning deep-learning svm


    【解决方案1】:

    因为数组形状分别为W 和X 的(D, C) 和(N, D),所以不能直接取点积,除非先将它们都转置(对于矩阵乘法,它们必须是(C, D)·(D, N) .

    从X.T.dot(W.T) == W.dot(X) 开始,实现只是简单地反转点积的顺序,而不是对每个数组进行变换。实际上,这只是取决于如何安排输入的决定。在这种情况下,(有点武断)决定以更直观的方式排列样本和特征,而不是将点积设为x·W。

    【讨论】:

    • 很好的解释,虽然我觉得应该是Xᵀ.dot(Wᵀ) == [W.dot(X)]ᵀ。
    猜你喜欢
    • 1970-01-01
    • 2019-03-17
    • 1970-01-01
    • 2016-10-20
    • 1970-01-01
    • 2014-04-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多