Row 6260
Content Data
This page contains data entry 6260 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts).
This is really surprising to me. I would think a single update large enough to substantially modify an output would be quite destructive, and possibly just not do what you want given the non-monotonicity of the loss landscape. Is there a good answer to how/why this happens, and if so can anyone provide a link to research that investigates this question? Is it a feature of large models (something like NTK), a feature of the transformer architecture, or something else? Note that I'm not asking about in-context learning, but the change from a single gradient step.
| Field | Value |
|---|---|
| text | I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts). This is really surprising to me. I would think a single update large enough to substantially modify an output would be… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-08 |
| username_encoded | Z0FBQUFBQm5LakwyV0hENXFmekJOZktCa2pKWU1xckdNT3pNUFpOb2txNHpzRnRXOGRoRDhQVDNHZkw3LVp5MzNMZXk5azZraHJkRjFVOEo4NmtWUE5ZN1RCeTNERk9OaGc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HLXVySnhveFkxZmFZdkg0bTFjNkU3cVVXb3NMMTNDcVNwcW9sN1pCS19zdzJUb2hid2lIbDRvdVZieWdlcnNzd01PSTIwNFJDN09Yemhob0tUU3VpalFra0Fnd0ZzSG9BSjFUR0xUXzk5T2NkWGVxSmhIcloxNUx3RkdGaXFDZmJhMl9GNUVXSGNYbVdPLWxtWTFhc3U1ZTNoSkpEeHI4MDlHa3VjNG9hbHBmM1U3RGRXSGd1SmdYWVRHT1ZxQ2NrYlYtV1RyajFFVjFxWTR3ODRPTjYxUT09 |
Raw Record
{
"text": "I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts).\n\nThis is really surprising to me. I would think a single update large enough to substantially modify an output would be quite destructive, and possibly just not do what you want given the non-monotonicity of the loss landscape. Is there a good answer to how/why this happens, and if so can anyone provide a link to research that investigates this question? Is it a feature of large models (something like NTK), a feature of the transformer architecture, or something else? Note that I'm not asking about in-context learning, but the change from a single gradient step.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-08",
"username_encoded": "Z0FBQUFBQm5LakwyV0hENXFmekJOZktCa2pKWU1xckdNT3pNUFpOb2txNHpzRnRXOGRoRDhQVDNHZkw3LVp5MzNMZXk5azZraHJkRjFVOEo4NmtWUE5ZN1RCeTNERk9OaGc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HLXVySnhveFkxZmFZdkg0bTFjNkU3cVVXb3NMMTNDcVNwcW9sN1pCS19zdzJUb2hid2lIbDRvdVZieWdlcnNzd01PSTIwNFJDN09Yemhob0tUU3VpalFra0Fnd0ZzSG9BSjFUR0xUXzk5T2NkWGVxSmhIcloxNUx3RkdGaXFDZmJhMl9GNUVXSGNYbVdPLWxtWTFhc3U1ZTNoSkpEeHI4MDlHa3VjNG9hbHBmM1U3RGRXSGd1SmdYWVRHT1ZxQ2NrYlYtV1RyajFFVjFxWTR3ODRPTjYxUT09"
}
Entry Information
- Entry ID: 6260
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000