Row 95527

Row ID: 95527 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 95527 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**TL;DR**: do NOT stuff more than one document in the context window while training an LM.

**Paper:** [https://arxiv.org/abs/2405.13226](https://arxiv.org/abs/2405.13226)

**Abstract:** Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences. These datasets are created by randomly concatenating documents of various lengths and then chunking them into sequences of a predetermined target length. However, this method of concatenation can lead to cross-document attention within a sequence, which is neither a desirable learning signal nor computationally efficient. Additionally, training on long sequences becomes computationally prohibitive due to the quadratic cost of attention. In this study, we introduce dataset decomposition, a novel variable sequence length training technique, to tackle these challenges. We decompose a dataset into a union of buckets, each containing sequences of the same size extracted from a unique document. During training, we use variable sequence length and batch size, sampling simultaneously from all buckets with a curriculum. In contrast to the concat-and-chunk baseline, which incurs a fixed attention cost at every step of training, our proposed method incurs a penalty proportional to the actual document lengths at each step, resulting in significant savings in training time. We train an 8k context-length 1B model at the same cost as a 2k context-length model trained with the baseline approach. Experiments on a web-scale corpus demonstrate that our approach significantly enhances performance on standard language evaluations and long-context benchmarks, reaching target accuracy 3x faster compared to the baseline. Our method not only enables efficient pretraining on long sequences but also scales effectively with dataset size. Lastly, we shed light on a critical yet less studied aspect of training large language models: the distribution and curriculum of sequence lengths, which results in a non-negligible difference in performance.

**Visual Summary:**

https://preview.redd.it/nnvi519tvj2d1.png?width=1123&format=png&auto=webp&s=334b8990f4ac2d4298e1a622d71301cd7d6beae3

FieldValue
text **TL;DR**: do NOT stuff more than one document in the context window while training an LM. **Paper:** [https://arxiv.org/abs/2405.13226](https://arxiv.org/abs/2405.13226) **Abstract:** Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences. These datasets are created by randomly concatenating documents of various lengths and then chunking them into sequences of a predetermined target length. However, this method of concatenation can lead to cro…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak11ZjViNEFjUVNySEgzU0Z3cFBDSzJkQ3M0aGVuQTNYakt2VmJPVUJNT2YwYklkdVRjX1ZlOHExUXc5cWNlYnJ2TWxyeFM4NjVhdWVHTUppN2hERXo4TnJkeGdRdmJfQzNGa2xqWFgtaGNpV2M9
url_encoded Z0FBQUFBQm5LalBBOGxFRjFXRDUzNjRMS01IaXoxbUlrU21wZ0ZPQVJDRmE4dURUTEN1UkNwX09CbEVBSzlOdGNXTVRwY1BtaXJmVFZHWHl1U0hWay1aQ0dnVE9GQUsyY25EVHlpMHdMNWNPOFhzVFYwUVR2M1VrUUFJb2M2RmxGR0M5MUJ4ZEt6RmdrQ2xsSUpPX1VFc3Q3UlZ6SUhESkUxTkJjek9BNXQ4UDV1YVZEWkxXYnh6VjQzcS0ydHdDeGVSX3puNlF2MnhlTmVpakptR3JSVHhpdnNqUm81R21MQT09

Raw Record

{
  "text": "**TL;DR**: do NOT stuff more than one document in the context window while training an LM.\n\n**Paper:** [https://arxiv.org/abs/2405.13226](https://arxiv.org/abs/2405.13226)\n\n**Abstract:** Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences. These datasets are created by randomly concatenating documents of various lengths and then chunking them into sequences of a predetermined target length. However, this method of concatenation can lead to cross-document attention within a sequence, which is neither a desirable learning signal nor computationally efficient. Additionally, training on long sequences becomes computationally prohibitive due to the quadratic cost of attention. In this study, we introduce dataset decomposition, a novel variable sequence length training technique, to tackle these challenges. We decompose a dataset into a union of buckets, each containing sequences of the same size extracted from a unique document. During training, we use variable sequence length and batch size, sampling simultaneously from all buckets with a curriculum. In contrast to the concat-and-chunk baseline, which incurs a fixed attention cost at every step of training, our proposed method incurs a penalty proportional to the actual document lengths at each step, resulting in significant savings in training time. We train an 8k context-length 1B model at the same cost as a 2k context-length model trained with the baseline approach. Experiments on a web-scale corpus demonstrate that our approach significantly enhances performance on standard language evaluations and long-context benchmarks, reaching target accuracy 3x faster compared to the baseline. Our method not only enables efficient pretraining on long sequences but also scales effectively with dataset size. Lastly, we shed light on a critical yet less studied aspect of training large language models: the distribution and curriculum of sequence lengths, which results in a non-negligible difference in performance.\n\n**Visual Summary:**\n\nhttps://preview.redd.it/nnvi519tvj2d1.png?width=1123&format=png&auto=webp&s=334b8990f4ac2d4298e1a622d71301cd7d6beae3",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak11ZjViNEFjUVNySEgzU0Z3cFBDSzJkQ3M0aGVuQTNYakt2VmJPVUJNT2YwYklkdVRjX1ZlOHExUXc5cWNlYnJ2TWxyeFM4NjVhdWVHTUppN2hERXo4TnJkeGdRdmJfQzNGa2xqWFgtaGNpV2M9",
  "url_encoded": "Z0FBQUFBQm5LalBBOGxFRjFXRDUzNjRMS01IaXoxbUlrU21wZ0ZPQVJDRmE4dURUTEN1UkNwX09CbEVBSzlOdGNXTVRwY1BtaXJmVFZHWHl1U0hWay1aQ0dnVE9GQUsyY25EVHlpMHdMNWNPOFhzVFYwUVR2M1VrUUFJb2M2RmxGR0M5MUJ4ZEt6RmdrQ2xsSUpPX1VFc3Q3UlZ6SUhESkUxTkJjek9BNXQ4UDV1YVZEWkxXYnh6VjQzcS0ydHdDeGVSX3puNlF2MnhlTmVpakptR3JSVHhpdnNqUm81R21MQT09"
}

Entry Information