Row 6936
Content Data
This page contains data entry 6936 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
This is a pretty obvious.
I recently see that the last linear layer of transformer is kind of a waste of parameters.
A transformer model is a stack of many transformer layers.
These layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\_model \* d\_dim\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer.
We all know that two consecutive linear transformation is representable by one linear transformation, which is the reason why we use activation functions at all.
So why we hasn't use a super sparse linear transformation, maybe do convolution by treating the embedding dimension as sequence dimension at that particular linear transformation dimension.
| Field | Value |
|---|---|
| text | This is a pretty obvious. I recently see that the last linear layer of transformer is kind of a waste of parameters. A transformer model is a stack of many transformer layers. These layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\_model \* d\_dim\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer. We all know that two consecutive linear transformation is… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-14 |
| username_encoded | Z0FBQUFBQm5LakwzanpXVmtkWTZ1bkFtRlU3Y3JHZzhUUkdDN1hLTGw3R0phdnB6ZVpWWDdDbTEzTkpESDdwVFB1QTFNUWJ3RkNjdUpHSTBKTkx1YWpVUW9vRUNFSXQtS0E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9Ha3g3bHVHclM3Ml9OUzMzZWw2MUFVTGtmbzFEMlc0OEx2cC11TU9LNTBMNG9OVVlfTUxuSmtIN1RTWk9xbmJzcVY5akNvcDA4aFR2VGdzREVZUF9KUGJBN1BfbWpSdlQ2SzI3UjhudG93eUN6OV9mcXZNVWpSVXJSUzExX0RGazRzeTdzQUJBNk5veGZjVDNsNm1BeFpVSWs5UVlpM215bzBxTVFWaHB3bGJ4cU1QbE1GejJNdy16Y0xteEdIRDd6YU5NTUVwZTYwMGgtbl9XQUwtakpLdz09 |
Raw Record
{
"text": "This is a pretty obvious.\n\nI recently see that the last linear layer of transformer is kind of a waste of parameters.\n\nA transformer model is a stack of many transformer layers.\n\nThese layers starts with 3 QKV Linear Transformation and ends with FFN Network, which consists of two linear layers. The last one costs (d\\_model \\* d\\_dim\\_feedforward) parameter and multiplication and its output is linearly transformed again at the next layer.\n\nWe all know that two consecutive linear transformation is representable by one linear transformation, which is the reason why we use activation functions at all.\n\nSo why we hasn't use a super sparse linear transformation, maybe do convolution by treating the embedding dimension as sequence dimension at that particular linear transformation dimension.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-14",
"username_encoded": "Z0FBQUFBQm5LakwzanpXVmtkWTZ1bkFtRlU3Y3JHZzhUUkdDN1hLTGw3R0phdnB6ZVpWWDdDbTEzTkpESDdwVFB1QTFNUWJ3RkNjdUpHSTBKTkx1YWpVUW9vRUNFSXQtS0E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9Ha3g3bHVHclM3Ml9OUzMzZWw2MUFVTGtmbzFEMlc0OEx2cC11TU9LNTBMNG9OVVlfTUxuSmtIN1RTWk9xbmJzcVY5akNvcDA4aFR2VGdzREVZUF9KUGJBN1BfbWpSdlQ2SzI3UjhudG93eUN6OV9mcXZNVWpSVXJSUzExX0RGazRzeTdzQUJBNk5veGZjVDNsNm1BeFpVSWs5UVlpM215bzBxTVFWaHB3bGJ4cU1QbE1GejJNdy16Y0xteEdIRDd6YU5NTUVwZTYwMGgtbl9XQUwtakpLdz09"
}
Entry Information
- Entry ID: 6936
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000