Row 3655

Row ID: 3655 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 3655 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**Paper**: [https://arxiv.org/abs/2401.17505](https://arxiv.org/abs/2401.17505)

**Abstract**:

>We study the probabilistic modeling performed by Autoregressive Large Language Models through the angle of time directionality. We empirically find a time asymmetry exhibited by such models in their ability to model natural language: a difference in the average log-perplexity when trying to predict the next token versus when trying to predict the previous one. This difference is at the same time subtle and very consistent across various modalities (language, model size, training time, ...). Theoretically, this is surprising: from an information-theoretic point of view, there should be no such difference. We provide a theoretical framework to explain how such an asymmetry can appear from sparsity and computational complexity considerations, and outline a number of perspectives opened by our results.

FieldValue
text **Paper**: [https://arxiv.org/abs/2401.17505](https://arxiv.org/abs/2401.17505) **Abstract**: >We study the probabilistic modeling performed by Autoregressive Large Language Models through the angle of time directionality. We empirically find a time asymmetry exhibited by such models in their ability to model natural language: a difference in the average log-perplexity when trying to predict the next token versus when trying to predict the previous one. This difference is at the same time…
label r/neuralnetworks
dataType post
communityName r/neuralnetworks
datetime 2024-03-18
username_encoded Z0FBQUFBQm5LakwxNWlEWFRMZWx0cnlPTlFTaTFQeC0tOGVhY0VhSmNld0h1bWhXQmU0UUNfbERjNTU1QmVEaENHZ3czSzlNVlg4NWN3UGplU1h0NjhEQXNsZlRmTDhGV0E9PQ==
url_encoded Z0FBQUFBQm5Lak9FN1ktQ3NUVW5uUjA0RXlfbjA5VjFoOFgzdmdXV0IwTlRudUZxc2R5Q0M5VGZJUkdFR3g0Q05wZXhLV3NwbDdBa0Y1MWJtWnJoREpiRl83aXlCZG5ld0d2NXpYRnpUQ1BpeUZpZVZOdWZobTZtV1VmRmxwNU51aHZCWF9HUkdJTWpYT3hFZ0stREZpZExkNTNqNzUzSkFnRmxYNDZqaXdCaW8zNURJdVJiX05JNUlnTUJsWEJuQ0FzMDFjVTREM3hjNk5JWHZnaGJ1dk54SDcwZjN3QTJiZz09

Raw Record

{
  "text": "**Paper**: [https://arxiv.org/abs/2401.17505](https://arxiv.org/abs/2401.17505)\n\n**Abstract**:\n\n>We study the probabilistic modeling performed by Autoregressive Large  Language Models through the angle of time directionality. We empirically  find a time asymmetry exhibited by such models in their ability to  model natural language: a difference in the average log-perplexity when  trying to predict the next token versus when trying to predict the  previous one. This difference is at the same time subtle and very  consistent across various modalities (language, model size, training  time, ...). Theoretically, this is surprising: from an  information-theoretic point of view, there should be no such difference.  We provide a theoretical framework to explain how such an asymmetry can  appear from sparsity and computational complexity considerations, and  outline a number of perspectives opened by our results.",
  "label": "r/neuralnetworks",
  "dataType": "post",
  "communityName": "r/neuralnetworks",
  "datetime": "2024-03-18",
  "username_encoded": "Z0FBQUFBQm5LakwxNWlEWFRMZWx0cnlPTlFTaTFQeC0tOGVhY0VhSmNld0h1bWhXQmU0UUNfbERjNTU1QmVEaENHZ3czSzlNVlg4NWN3UGplU1h0NjhEQXNsZlRmTDhGV0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9FN1ktQ3NUVW5uUjA0RXlfbjA5VjFoOFgzdmdXV0IwTlRudUZxc2R5Q0M5VGZJUkdFR3g0Q05wZXhLV3NwbDdBa0Y1MWJtWnJoREpiRl83aXlCZG5ld0d2NXpYRnpUQ1BpeUZpZVZOdWZobTZtV1VmRmxwNU51aHZCWF9HUkdJTWpYT3hFZ0stREZpZExkNTNqNzUzSkFnRmxYNDZqaXdCaW8zNURJdVJiX05JNUlnTUJsWEJuQ0FzMDFjVTREM3hjNk5JWHZnaGJ1dk54SDcwZjN3QTJiZz09"
}

Entry Information