Row 6260

Row ID: 6260 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 6260 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts).

This is really surprising to me. I would think a single update large enough to substantially modify an output would be quite destructive, and possibly just not do what you want given the non-monotonicity of the loss landscape. Is there a good answer to how/why this happens, and if so can anyone provide a link to research that investigates this question? Is it a feature of large models (something like NTK), a feature of the transformer architecture, or something else? Note that I'm not asking about in-context learning, but the change from a single gradient step.

FieldValue
text I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts). This is really surprising to me. I would think a single update large enough to substantially modify an output would be…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-08
username_encoded Z0FBQUFBQm5LakwyV0hENXFmekJOZktCa2pKWU1xckdNT3pNUFpOb2txNHpzRnRXOGRoRDhQVDNHZkw3LVp5MzNMZXk5azZraHJkRjFVOEo4NmtWUE5ZN1RCeTNERk9OaGc9PQ==
url_encoded Z0FBQUFBQm5Lak9HLXVySnhveFkxZmFZdkg0bTFjNkU3cVVXb3NMMTNDcVNwcW9sN1pCS19zdzJUb2hid2lIbDRvdVZieWdlcnNzd01PSTIwNFJDN09Yemhob0tUU3VpalFra0Fnd0ZzSG9BSjFUR0xUXzk5T2NkWGVxSmhIcloxNUx3RkdGaXFDZmJhMl9GNUVXSGNYbVdPLWxtWTFhc3U1ZTNoSkpEeHI4MDlHa3VjNG9hbHBmM1U3RGRXSGd1SmdYWVRHT1ZxQ2NrYlYtV1RyajFFVjFxWTR3ODRPTjYxUT09

Raw Record

{
  "text": "I guess I would first like a citation of that fact, or for somebody to tell me I made it up. But folk knowledge is that a transformer trained for a single epoch can recall facts that only appear a single time in the training dataset. This implies that a single update is enough to modify the weights to produce the correct output (without catastrophically forgetting other facts).\n\nThis is really surprising to me. I would think a single update large enough to substantially modify an output would be quite destructive, and possibly just not do what you want given the non-monotonicity of the loss landscape. Is there a good answer to how/why this happens, and if so can anyone provide a link to research that investigates this question? Is it a feature of large models (something like NTK), a feature of the transformer architecture, or something else? Note that I'm not asking about in-context learning, but the change from a single gradient step.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-08",
  "username_encoded": "Z0FBQUFBQm5LakwyV0hENXFmekJOZktCa2pKWU1xckdNT3pNUFpOb2txNHpzRnRXOGRoRDhQVDNHZkw3LVp5MzNMZXk5azZraHJkRjFVOEo4NmtWUE5ZN1RCeTNERk9OaGc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9HLXVySnhveFkxZmFZdkg0bTFjNkU3cVVXb3NMMTNDcVNwcW9sN1pCS19zdzJUb2hid2lIbDRvdVZieWdlcnNzd01PSTIwNFJDN09Yemhob0tUU3VpalFra0Fnd0ZzSG9BSjFUR0xUXzk5T2NkWGVxSmhIcloxNUx3RkdGaXFDZmJhMl9GNUVXSGNYbVdPLWxtWTFhc3U1ZTNoSkpEeHI4MDlHa3VjNG9hbHBmM1U3RGRXSGd1SmdYWVRHT1ZxQ2NrYlYtV1RyajFFVjFxWTR3ODRPTjYxUT09"
}

Entry Information