Row 57282

Row ID: 57282 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 57282 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

You bring up two important points here.

1. Shifting-up can cause clipping if we shift the gradients up too much. That's why in practice we can also dynamically find the loss scale factor that doesn't cause such overflows (e.g. multiplying the max gradient by the scale and seeing if it overflows in FP16 -> https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html#scalefactor) 2. While I don't have a complete answer here, part of the reasons are the softmax function, and backpropagation. Remember the gradients are the partial derivatives with respect to the loss. And the loss we use is the cross entropy loss, that involves softmaxing the given logits.

When we start calculating the weight gradients, we first start the backpropagation chain by calculating dlogits, then dW2, etc. Notice how dlogits, because we softmax it, is always a pretty small negative number. We multiply this number with the rest of the gradients because of backpropagation and thus our weight gradients are bound to be pretty small too.

FieldValue
text You bring up two important points here. 1. Shifting-up can cause clipping if we shift the gradients up too much. That's why in practice we can also dynamically find the loss scale factor that doesn't cause such overflows (e.g. multiplying the max gradient by the scale and seeing if it overflows in FP16 -> https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html#scalefactor) 2. While I don't have a complete answer here, part of the reasons are the softmax function, an…
label r/deeplearning
dataType comment
communityName r/deeplearning
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1Xdnh0OXF1ZDB0RFFUdWwzOWFnZVNrXzZyNFh6cUFlNldlR2pOZGg0TDZuaVlPWWJJWHlXLURQb0NPVzhhaC0yTlY4WDlXMG9keFBFa3ZlRVUyNGRsOWc9PQ==
url_encoded Z0FBQUFBQm5Lak9tNnVOV3BFNFFTZjlrQjVvcFBkSEZQVWNhQTBma1FpeXdiSkI4aHEzUk1UcDlySTljVXVJaG1td3hZTnBpNzdtNEk0alc0UUZUZ1FrQmxObEJJRml6TXpGaHFFNTFnUWt0QXRxSllBdHFwaWdhdFd3d0ZFamNvRE9leEt3ZHVZVmQ0ckZfbzBRTFktWHJjSTFTTnJ6bkF2Vy1tYWF5c3JxVlhPT3RoOHdQejltallJbF9TbWxoNm9FTzNiOEdHN2h5ZXUtNkJadW5wcGZCbTB2eVhEcG1wZz09

Raw Record

{
  "text": "You bring up two important points here.\n\n1. Shifting-up can cause clipping if we shift the gradients up too much. That's why in practice we can also dynamically find the loss scale factor that doesn't cause such overflows (e.g. multiplying the max gradient by the scale and seeing if it overflows in FP16 -> https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html#scalefactor)\n2. While I don't have a complete answer here, part of the reasons are the softmax function, and backpropagation. Remember the gradients are the partial derivatives with respect to the loss. And the loss we use is the cross entropy loss, that involves softmaxing the given logits.\n\nWhen we start calculating the weight gradients, we first start the backpropagation chain by calculating dlogits, then dW2, etc. Notice how dlogits, because we softmax it, is always a pretty small negative number. We multiply this number with the rest of the gradients because of backpropagation and thus our weight gradients are bound to be pretty small too.",
  "label": "r/deeplearning",
  "dataType": "comment",
  "communityName": "r/deeplearning",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1Xdnh0OXF1ZDB0RFFUdWwzOWFnZVNrXzZyNFh6cUFlNldlR2pOZGg0TDZuaVlPWWJJWHlXLURQb0NPVzhhaC0yTlY4WDlXMG9keFBFa3ZlRVUyNGRsOWc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9tNnVOV3BFNFFTZjlrQjVvcFBkSEZQVWNhQTBma1FpeXdiSkI4aHEzUk1UcDlySTljVXVJaG1td3hZTnBpNzdtNEk0alc0UUZUZ1FrQmxObEJJRml6TXpGaHFFNTFnUWt0QXRxSllBdHFwaWdhdFd3d0ZFamNvRE9leEt3ZHVZVmQ0ckZfbzBRTFktWHJjSTFTTnJ6bkF2Vy1tYWF5c3JxVlhPT3RoOHdQejltallJbF9TbWxoNm9FTzNiOEdHN2h5ZXUtNkJadW5wcGZCbTB2eVhEcG1wZz09"
}

Entry Information