Row 96461

Row ID: 96461 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 96461 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art.

Bucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient.

They also seem to be unaware of how common it is to reset the auto-regressive attention mask at document boundaries.

FieldValue
text This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art. Bucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient. They also seem to be unaware of how common it is to re…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak12QXJEVzA3cXpid2hOcWd3eXd0R2xHd0FXbGwwdXc4RFRoQkxwX0hTS1dMaFlaNXdMd2R6ZkNDSV90aUNwR25KMkVnNUhLOFJsSVNGX2hwcXdONU9OU3c9PQ==
url_encoded Z0FBQUFBQm5LalBCZkgzYmhuNVU3dTNmcnpVNWdTY0hSbFRKY21fVjlvNmlwRkxkWnJFZncyMXpRRmhmWjFSYWNCS0xWMWVmRjN0dVhZa1hTdVhRRzQ1UmVCdFczWkFYaXZMTS1Fc1NaUmtiR0tqcnlQQlk1MFVEeWs5cW1DZWlQMURRa0N0MVdlamE5SHJVSTNrdW5LcXdEM0Rrazl0X2liWEZSUVhSbFU2WjRfYy1ZbV9UV3dtUjdqNC1BX0pseEFIdUM2cS0teG1yY2M0R1JXc0JqM01MR040V1FOZmZHNXhkU2w3UkdZblBqNmNGQS1NS2ItUT0=

Raw Record

{
  "text": "This would almost be funny if it weren't unfortunate – it's reinventing the wheel and calling it novel, while misunderstanding the current state of the art. \n\nBucketing by length to create variable length batches with minimal padding is how Transformer models have been trained since they were invented. Their large claim is that this is more efficient, but the shift to packing documents in LLM training was made *because* it's more efficient.\n\nThey also seem to be unaware of how common it is to reset the auto-regressive attention mask at document boundaries.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak12QXJEVzA3cXpid2hOcWd3eXd0R2xHd0FXbGwwdXc4RFRoQkxwX0hTS1dMaFlaNXdMd2R6ZkNDSV90aUNwR25KMkVnNUhLOFJsSVNGX2hwcXdONU9OU3c9PQ==",
  "url_encoded": "Z0FBQUFBQm5LalBCZkgzYmhuNVU3dTNmcnpVNWdTY0hSbFRKY21fVjlvNmlwRkxkWnJFZncyMXpRRmhmWjFSYWNCS0xWMWVmRjN0dVhZa1hTdVhRRzQ1UmVCdFczWkFYaXZMTS1Fc1NaUmtiR0tqcnlQQlk1MFVEeWs5cW1DZWlQMURRa0N0MVdlamE5SHJVSTNrdW5LcXdEM0Rrazl0X2liWEZSUVhSbFU2WjRfYy1ZbV9UV3dtUjdqNC1BX0pseEFIdUM2cS0teG1yY2M0R1JXc0JqM01MR040V1FOZmZHNXhkU2w3UkdZblBqNmNGQS1NS2ItUT0="
}

Entry Information