Row 6936

Row ID: 6936 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 6936 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

This is a pretty obvious.

I recently see that the last linear layer of transformer is kind of a waste of parameters.

A transformer model is a stack of many transformer layers.

These layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\_model \* d\_dim\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer.

We all know that two consecutive linear transformation is representable by one linear transformation, which is the reason why we use activation functions at all.

So why we hasn't use a super sparse linear transformation, maybe do convolution by treating the embedding dimension as sequence dimension at that particular linear transformation dimension.

FieldValue
text This is a pretty obvious. I recently see that the last linear layer of transformer is kind of a waste of parameters. A transformer model is a stack of many transformer layers. These layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\_model \* d\_dim\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer. We all know that two consecutive linear transformation is…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-14
username_encoded Z0FBQUFBQm5LakwzanpXVmtkWTZ1bkFtRlU3Y3JHZzhUUkdDN1hLTGw3R0phdnB6ZVpWWDdDbTEzTkpESDdwVFB1QTFNUWJ3RkNjdUpHSTBKTkx1YWpVUW9vRUNFSXQtS0E9PQ==
url_encoded Z0FBQUFBQm5Lak9Ha3g3bHVHclM3Ml9OUzMzZWw2MUFVTGtmbzFEMlc0OEx2cC11TU9LNTBMNG9OVVlfTUxuSmtIN1RTWk9xbmJzcVY5akNvcDA4aFR2VGdzREVZUF9KUGJBN1BfbWpSdlQ2SzI3UjhudG93eUN6OV9mcXZNVWpSVXJSUzExX0RGazRzeTdzQUJBNk5veGZjVDNsNm1BeFpVSWs5UVlpM215bzBxTVFWaHB3bGJ4cU1QbE1GejJNdy16Y0xteEdIRDd6YU5NTUVwZTYwMGgtbl9XQUwtakpLdz09

Raw Record

{
  "text": "This is a pretty obvious.\n\nI recently see that the last linear layer of transformer is kind of a waste of parameters.\n\nA transformer model is a stack of many transformer layers.\n\nThese layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\\_model \\* d\\_dim\\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer.\n\nWe all know that two consecutive linear transformation is representable by one linear transformation, which is the reason why we use activation functions at all.\n\nSo why we hasn't use a super sparse linear transformation, maybe do convolution by treating the embedding dimension as sequence dimension at that particular linear transformation dimension.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-14",
  "username_encoded": "Z0FBQUFBQm5LakwzanpXVmtkWTZ1bkFtRlU3Y3JHZzhUUkdDN1hLTGw3R0phdnB6ZVpWWDdDbTEzTkpESDdwVFB1QTFNUWJ3RkNjdUpHSTBKTkx1YWpVUW9vRUNFSXQtS0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9Ha3g3bHVHclM3Ml9OUzMzZWw2MUFVTGtmbzFEMlc0OEx2cC11TU9LNTBMNG9OVVlfTUxuSmtIN1RTWk9xbmJzcVY5akNvcDA4aFR2VGdzREVZUF9KUGJBN1BfbWpSdlQ2SzI3UjhudG93eUN6OV9mcXZNVWpSVXJSUzExX0RGazRzeTdzQUJBNk5veGZjVDNsNm1BeFpVSWs5UVlpM215bzBxTVFWaHB3bGJ4cU1QbE1GejJNdy16Y0xteEdIRDd6YU5NTUVwZTYwMGgtbl9XQUwtakpLdz09"
}

Entry Information