Row 4965

Row ID: 4965 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 4965 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408)

Github: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte)

Abstract:

>Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel **byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling.** SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that **performance is significantly improved** by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that **for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.**Paper: https://arxiv.org/abs/2404.14408Github: https://github.com/kjslag/spacebyteAbstract:Tokenization is widely used in large language models because it significantly improves performance. However, tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity. To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling. SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that performance is significantly improved by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.

https://preview.redd.it/v1xo6g1gzewc1.jpg?width=1507&format=pjpg&auto=webp&s=f9d415307b60639fa67e8a54c8769fa5a6c10f04

https://preview.redd.it/edvqos1gzewc1.jpg?width=1654&format=pjpg&auto=webp&s=f91c8727017e1a1bc7b80bb77a8627ff99182607

https://preview.redd.it/fe6z6i1gzewc1.jpg?width=1181&format=pjpg&auto=webp&s=24d955f30b8ca3eaa7c527f3f40545ed493f789c

FieldValue
text Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408) Github: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte) Abstract: >Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disad…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-04-24
username_encoded Z0FBQUFBQm5LakwyOEJfOHpOamRFcDVvd2w3Vml2YzhOQTBaZUVIVG1UeUJtRmlZb1FqY0lPVU4yT0Z1bF9QbTluWi1Mc2JZWmRlUnhSai01aUluNkIxMXhtVXdxQlkzbVE9PQ==
url_encoded Z0FBQUFBQm5Lak9GQUd5bVp0SXlPaXctLUdPNm1VYnR1UnRZSlV2N2ZIUTlFVmN3N1dFWU90SGo1VEJiRHByRU1DMzR2bTlKSnUtcjZHaWNGbFM2bFlrZy0wa29zWThKaDlKd2kyTEVEY29mbEJ4cUloMjU5Z1U0aThVd1hFTzhCTWFvcThWVTNTQ3hFbFgtc2JPQzJDUEZMbTNmel9uNGF1V3ZVX1FWak5tVVI0NVNMUkNoQnpfR3R6eThSNFQteUkzS0JkOVVLclJJLUsxNjFxV1c5ZGZkR0o2WU9OZzJGQT09

Raw Record

{
  "text": "Paper: [https://arxiv.org/abs/2404.14408](https://arxiv.org/abs/2404.14408)\n\nGithub: [https://github.com/kjslag/spacebyte](https://github.com/kjslag/spacebyte)\n\nAbstract:\n\n>Tokenization is widely used in large language models because it significantly improves performance. However, **tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity.** To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel **byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling.** SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that **performance is significantly improved** by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that **for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.**Paper: https://arxiv.org/abs/2404.14408Github: https://github.com/kjslag/spacebyteAbstract:Tokenization is widely used in large language models because it significantly improves performance. However, tokenization imposes several disadvantages, such as performance biases, increased adversarial vulnerability, decreased character-level modeling performance, and increased modeling complexity. To address these disadvantages without sacrificing performance, we propose SpaceByte, a novel byte-level decoder architecture that closes the performance gap between byte-level and subword autoregressive language modeling. SpaceByte consists of a byte-level Transformer model, but with extra larger transformer blocks inserted in the middle of the layers. We find that performance is significantly improved by applying these larger blocks only after certain bytes, such as space characters, which typically denote word boundaries. Our experiments show that for a fixed training and inference compute budget, SpaceByte outperforms other byte-level architectures and roughly matches the performance of tokenized Transformer architectures.\n\nhttps://preview.redd.it/v1xo6g1gzewc1.jpg?width=1507&format=pjpg&auto=webp&s=f9d415307b60639fa67e8a54c8769fa5a6c10f04\n\nhttps://preview.redd.it/edvqos1gzewc1.jpg?width=1654&format=pjpg&auto=webp&s=f91c8727017e1a1bc7b80bb77a8627ff99182607\n\nhttps://preview.redd.it/fe6z6i1gzewc1.jpg?width=1181&format=pjpg&auto=webp&s=24d955f30b8ca3eaa7c527f3f40545ed493f789c",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-04-24",
  "username_encoded": "Z0FBQUFBQm5LakwyOEJfOHpOamRFcDVvd2w3Vml2YzhOQTBaZUVIVG1UeUJtRmlZb1FqY0lPVU4yT0Z1bF9QbTluWi1Mc2JZWmRlUnhSai01aUluNkIxMXhtVXdxQlkzbVE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GQUd5bVp0SXlPaXctLUdPNm1VYnR1UnRZSlV2N2ZIUTlFVmN3N1dFWU90SGo1VEJiRHByRU1DMzR2bTlKSnUtcjZHaWNGbFM2bFlrZy0wa29zWThKaDlKd2kyTEVEY29mbEJ4cUloMjU5Z1U0aThVd1hFTzhCTWFvcThWVTNTQ3hFbFgtc2JPQzJDUEZMbTNmel9uNGF1V3ZVX1FWak5tVVI0NVNMUkNoQnpfR3R6eThSNFQteUkzS0JkOVVLclJJLUsxNjFxV1c5ZGZkR0o2WU9OZzJGQT09"
}

Entry Information