Row 4965
Content Data
This page contains data entry 4965 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408)
Github: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte)
Abstract:
>Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel **byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling.** SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that **performance is significantly improved** by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that **for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.**Paper: https://arxiv.org/abs/2404.14408Github: https://github.com/kjslag/spacebyteAbstract:Tokenization is widely used in large language models because it significantly improves performance. However, tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity. To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling. SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that performance is significantly improved by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.
https://preview.redd.it/v1xo6g1gzewc1.jpg?width=1507&format=pjpg&auto=webp&s=f9d415307b60639fa67e8a54c8769fa5a6c10f04
https://preview.redd.it/edvqos1gzewc1.jpg?width=1654&format=pjpg&auto=webp&s=f91c8727017e1a1bc7b80bb77a8627ff99182607
https://preview.redd.it/fe6z6i1gzewc1.jpg?width=1181&format=pjpg&auto=webp&s=24d955f30b8ca3eaa7c527f3f40545ed493f789c
| Field | Value |
|---|---|
| text | Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408) Github: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte) Abstract: >Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disad… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-04-24 |
| username_encoded | Z0FBQUFBQm5LakwyOEJfOHpOamRFcDVvd2w3Vml2YzhOQTBaZUVIVG1UeUJtRmlZb1FqY0lPVU4yT0Z1bF9QbTluWi1Mc2JZWmRlUnhSai01aUluNkIxMXhtVXdxQlkzbVE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9GQUd5bVp0SXlPaXctLUdPNm1VYnR1UnRZSlV2N2ZIUTlFVmN3N1dFWU90SGo1VEJiRHByRU1DMzR2bTlKSnUtcjZHaWNGbFM2bFlrZy0wa29zWThKaDlKd2kyTEVEY29mbEJ4cUloMjU5Z1U0aThVd1hFTzhCTWFvcThWVTNTQ3hFbFgtc2JPQzJDUEZMbTNmel9uNGF1V3ZVX1FWak5tVVI0NVNMUkNoQnpfR3R6eThSNFQteUkzS0JkOVVLclJJLUsxNjFxV1c5ZGZkR0o2WU9OZzJGQT09 |
Raw Record
{
"text": "Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408)\n\nGithub: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte)\n\nAbstract:\n\n>Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel **byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling.** SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that **performance is significantly improved** by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that **for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.**Paper: https://arxiv.org/abs/2404.14408Github: https://github.com/kjslag/spacebyteAbstract:Tokenization is widely used in large language models because it significantly improves performance. However, tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity. To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling. SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that performance is significantly improved by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.\n\nhttps://preview.redd.it/v1xo6g1gzewc1.jpg?width=1507&format=pjpg&auto=webp&s=f9d415307b60639fa67e8a54c8769fa5a6c10f04\n\nhttps://preview.redd.it/edvqos1gzewc1.jpg?width=1654&format=pjpg&auto=webp&s=f91c8727017e1a1bc7b80bb77a8627ff99182607\n\nhttps://preview.redd.it/fe6z6i1gzewc1.jpg?width=1181&format=pjpg&auto=webp&s=24d955f30b8ca3eaa7c527f3f40545ed493f789c",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-04-24",
"username_encoded": "Z0FBQUFBQm5LakwyOEJfOHpOamRFcDVvd2w3Vml2YzhOQTBaZUVIVG1UeUJtRmlZb1FqY0lPVU4yT0Z1bF9QbTluWi1Mc2JZWmRlUnhSai01aUluNkIxMXhtVXdxQlkzbVE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9GQUd5bVp0SXlPaXctLUdPNm1VYnR1UnRZSlV2N2ZIUTlFVmN3N1dFWU90SGo1VEJiRHByRU1DMzR2bTlKSnUtcjZHaWNGbFM2bFlrZy0wa29zWThKaDlKd2kyTEVEY29mbEJ4cUloMjU5Z1U0aThVd1hFTzhCTWFvcThWVTNTQ3hFbFgtc2JPQzJDUEZMbTNmel9uNGF1V3ZVX1FWak5tVVI0NVNMUkNoQnpfR3R6eThSNFQteUkzS0JkOVVLclJJLUsxNjFxV1c5ZGZkR0o2WU9OZzJGQT09"
}
Entry Information
- Entry ID: 4965
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000