Row 5267
Content Data
This page contains data entry 5267 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
[https://arxiv.org/abs/2404.15758](https://arxiv.org/abs/2404.15758)
# From the abstract
We show that transformers can use meaningless filler tokens (e.g., '......') in place of a chain of thought to solve two hard algorithmic tasks they could not solve when responding without intermediate tokens. However, we find empirically that learning to use filler tokens is difficult and requires specific, dense supervision to converge
| Field | Value |
|---|---|
| text | [https://arxiv.org/abs/2404.15758](https://arxiv.org/abs/2404.15758) # From the abstract We show that transformers can use meaningless filler tokens (e.g., '......') in place of a chain of thought to solve two hard algorithmic tasks they could not solve when responding without intermediate tokens. However, we find empirically that learning to use filler tokens is difficult and requires specific, dense supervision to converge |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-04-28 |
| username_encoded | Z0FBQUFBQm5LakwyV1U5OFppTmRmTmRUTEU5MGkzNGlVT3B5dVhMeS1TWFlzaUQ4S1VscVBSV0xxSUVXaEd2Q21aS1lTSXcwUldxLVJSVE8xSmRUVlNFdFlkeEJCZ3V3TE16U0I0OEw4aVBCdU5KZVNUbm1hUFk9 |
| url_encoded | Z0FBQUFBQm5Lak9GLUl6Yk9UTGpwaFpGRGJiUFh0SW96b202VjJkX2t5V1pSdW9HU2NsaFI0WVo5aEN1eU5PN2w5aTUzSUtjNzh3b2txRS13LVYyV1RDRlJmRmJVcldHNUFrWGxXMS04ZTVFYmRnWHQ4X1ZWcktfUW9yS0xTdmZOTFBRcG95MmZ3a0M2VUlzTFU3ZzQtalJzaVBON2hLOGNKZFkyNkxHaThtMjhhdkE3VFU1UzVJSW5iRzdXLWRUbURnd1R4V2xZZlJNcmhWMnFucHFjUUpJSWstWnZrZlB2QT09 |
Raw Record
{
"text": "[https://arxiv.org/abs/2404.15758](https://arxiv.org/abs/2404.15758)\n\n# From the abstract\n\n \nWe show that transformers can use meaningless filler tokens (e.g., '......') in place of a chain of thought to solve two hard algorithmic tasks they could not solve when responding without intermediate tokens. However, we find empirically that learning to use filler tokens is difficult and requires specific, dense supervision to converge",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-04-28",
"username_encoded": "Z0FBQUFBQm5LakwyV1U5OFppTmRmTmRUTEU5MGkzNGlVT3B5dVhMeS1TWFlzaUQ4S1VscVBSV0xxSUVXaEd2Q21aS1lTSXcwUldxLVJSVE8xSmRUVlNFdFlkeEJCZ3V3TE16U0I0OEw4aVBCdU5KZVNUbm1hUFk9",
"url_encoded": "Z0FBQUFBQm5Lak9GLUl6Yk9UTGpwaFpGRGJiUFh0SW96b202VjJkX2t5V1pSdW9HU2NsaFI0WVo5aEN1eU5PN2w5aTUzSUtjNzh3b2txRS13LVYyV1RDRlJmRmJVcldHNUFrWGxXMS04ZTVFYmRnWHQ4X1ZWcktfUW9yS0xTdmZOTFBRcG95MmZ3a0M2VUlzTFU3ZzQtalJzaVBON2hLOGNKZFkyNkxHaThtMjhhdkE3VFU1UzVJSW5iRzdXLWRUbURnd1R4V2xZZlJNcmhWMnFucHFjUUpJSWstWnZrZlB2QT09"
}
Entry Information
- Entry ID: 5267
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000