Row 4813

Row ID: 4813 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 4813 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

The researchers segmented the sequence and added special memory tokens to the input: memory states from the output of the previous segment became inputs for the next one. Thus, a whole transformer acts as a recurrent cell, and memory serves as the recurrent state of the network. This approach was called Recurrent Memory Transformer (RMT).

The authors augmented small transformer models like BERT and GPT-2 with this memory and tested them on various question-answering tasks where facts needed for answering are somewhere in the text. It was found that using recurrent memory significantly increases the length of the input sequence while maintaining satisfactory neural network performance accuracy. In their experiments, scientists were able to extend this value to 2 million tokens. According to the authors, there are no fundamental limitations for this value to increase further, as the computational complexity of RMT grows linearly with the number of tokens.

[The accuracy of the pre-trained BERT model augmented with RMT on three tasks vs the number of tokens in the input sequence. The gray numbers indicate the GPU memory consumption, and the vertical lines represent the length limits in SOTA models \(as of the end of 2023\)](https://preview.redd.it/vejiotc890wc1.png?width=1504&format=png&auto=webp&s=777583bb65249b76dc9f222abbd705d8f6606d19)

The research was published in the proceedings of the AAAI-24 conference, additional details are provided in the [preprint](https://arxiv.org/abs/2304.11062), and the code is available on [GitHub](https://github.com/booydar/recurrent-memory-transformer/tree/aaai24).

FieldValue
text The researchers segmented the sequence and added special memory tokens to the input: memory states from the output of the previous segment became inputs for the next one. Thus, a whole transformer acts as a recurrent cell, and memory serves as the recurrent state of the network. This approach was called Recurrent Memory Transformer (RMT). The authors augmented small transformer models like BERT and GPT-2 with this memory and tested them on various question-answering tasks where facts needed for…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-04-22
username_encoded Z0FBQUFBQm5LakwyMnpOWWlVaklRLS1WWmlrdmJfMTlFNzFZd3JFb282VllnMTEwblZKR0cxeHJKV3ZRU0RWbUU2MmFZaTZ6dUppLVAzRXJwMXpyb2FZbXJZZmwxVlZtYVE9PQ==
url_encoded Z0FBQUFBQm5Lak9GSVo5eDB1WUhPWEhTYnRfbTV5V3NQT0tuQlRBMFBNWW4taTFCNVA0eGVRcGlsRVNTRDF4YlFOTDRPQ0NiNGszbTlNMG84dl9MT2ZZeUU2WHhIZl93WDRSMEN6WFVfY3NJYXlNQ3NSU1ZQTWxqZGEwbGFqbHZ3NjNVNnNqVklLVHI2VXV5N3J3NW1FTkFIdVJtNnNNMzF2S2ZhVHE3TGtEem55dGF4OUdCYlhwOHpiaHk5UTJLN0x5QVh0djc2cnJ1RzN2c2hhS2YtS0Z0S2YxR2d0ekJWZz09

Raw Record

{
  "text": "The researchers segmented the sequence and added special memory tokens to the input: memory states from the output of the previous segment became inputs for the next one. Thus, a whole transformer acts as a recurrent cell, and memory serves as the recurrent state of the network. This approach was called Recurrent Memory Transformer (RMT).\n\nThe authors augmented small transformer models like BERT and GPT-2 with this memory and tested them on various question-answering tasks where facts needed for answering are somewhere in the text. It was found that using recurrent memory significantly increases the length of the input sequence while maintaining satisfactory neural network performance accuracy. In their experiments, scientists were able to extend this value to 2 million tokens. According to the authors, there are no fundamental limitations for this value to increase further, as the computational complexity of RMT grows linearly with the number of tokens.\n\n[The accuracy of the pre-trained BERT model augmented with RMT on three tasks vs the number of tokens in the input sequence. The gray numbers indicate the GPU memory consumption, and the vertical lines represent the length limits in SOTA models \\(as of the end of 2023\\)](https://preview.redd.it/vejiotc890wc1.png?width=1504&format=png&auto=webp&s=777583bb65249b76dc9f222abbd705d8f6606d19)\n\nThe research was published in the proceedings of the AAAI-24 conference, additional details are provided in the [preprint](https://arxiv.org/abs/2304.11062), and the code is available on [GitHub](https://github.com/booydar/recurrent-memory-transformer/tree/aaai24).",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-04-22",
  "username_encoded": "Z0FBQUFBQm5LakwyMnpOWWlVaklRLS1WWmlrdmJfMTlFNzFZd3JFb282VllnMTEwblZKR0cxeHJKV3ZRU0RWbUU2MmFZaTZ6dUppLVAzRXJwMXpyb2FZbXJZZmwxVlZtYVE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GSVo5eDB1WUhPWEhTYnRfbTV5V3NQT0tuQlRBMFBNWW4taTFCNVA0eGVRcGlsRVNTRDF4YlFOTDRPQ0NiNGszbTlNMG84dl9MT2ZZeUU2WHhIZl93WDRSMEN6WFVfY3NJYXlNQ3NSU1ZQTWxqZGEwbGFqbHZ3NjNVNnNqVklLVHI2VXV5N3J3NW1FTkFIdVJtNnNNMzF2S2ZhVHE3TGtEem55dGF4OUdCYlhwOHpiaHk5UTJLN0x5QVh0djc2cnJ1RzN2c2hhS2YtS0Z0S2YxR2d0ekJWZz09"
}

Entry Information