Row 96461
Content Data
This page contains data entry 96461 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art.
Bucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient.
They also seem to be unaware of how common it is to reset the auto-regressive attention mask at document boundaries.
| Field | Value |
|---|---|
| text | This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art. Bucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient. They also seem to be unaware of how common it is to re… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak12QXJEVzA3cXpid2hOcWd3eXd0R2xHd0FXbGwwdXc4RFRoQkxwX0hTS1dMaFlaNXdMd2R6ZkNDSV90aUNwR25KMkVnNUhLOFJsSVNGX2hwcXdONU9OU3c9PQ== |
| url_encoded | Z0FBQUFBQm5LalBCZkgzYmhuNVU3dTNmcnpVNWdTY0hSbFRKY21fVjlvNmlwRkxkWnJFZncyMXpRRmhmWjFSYWNCS0xWMWVmRjN0dVhZa1hTdVhRRzQ1UmVCdFczWkFYaXZMTS1Fc1NaUmtiR0tqcnlQQlk1MFVEeWs5cW1DZWlQMURRa0N0MVdlamE5SHJVSTNrdW5LcXdEM0Rrazl0X2liWEZSUVhSbFU2WjRfYy1ZbV9UV3dtUjdqNC1BX0pseEFIdUM2cS0teG1yY2M0R1JXc0JqM01MR040V1FOZmZHNXhkU2w3UkdZblBqNmNGQS1NS2ItUT0= |
Raw Record
{
"text": "This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art. \n\nBucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient.\n\nThey also seem to be unaware of how common it is to reset the auto-regressive attention mask at document boundaries.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak12QXJEVzA3cXpid2hOcWd3eXd0R2xHd0FXbGwwdXc4RFRoQkxwX0hTS1dMaFlaNXdMd2R6ZkNDSV90aUNwR25KMkVnNUhLOFJsSVNGX2hwcXdONU9OU3c9PQ==",
"url_encoded": "Z0FBQUFBQm5LalBCZkgzYmhuNVU3dTNmcnpVNWdTY0hSbFRKY21fVjlvNmlwRkxkWnJFZncyMXpRRmhmWjFSYWNCS0xWMWVmRjN0dVhZa1hTdVhRRzQ1UmVCdFczWkFYaXZMTS1Fc1NaUmtiR0tqcnlQQlk1MFVEeWs5cW1DZWlQMURRa0N0MVdlamE5SHJVSTNrdW5LcXdEM0Rrazl0X2liWEZSUVhSbFU2WjRfYy1ZbV9UV3dtUjdqNC1BX0pseEFIdUM2cS0teG1yY2M0R1JXc0JqM01MR040V1FOZmZHNXhkU2w3UkdZblBqNmNGQS1NS2ItUT0="
}
Entry Information
- Entry ID: 96461
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000