Row 5673

Row ID: 5673 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 5673 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

There's been a lot of discussion about benchmark contamination, where models are trained on the data they are ultimately evaluated on. For example, a [recent paper](https://twitter.com/hughbzhang/status/1785877026794356858) showed that models performed substantially better on the public GSM8K vs GSM1K, which was a benchmark recently created by Scale AI to match GSM8K on difficulty and other measures.

Because of these concerns about benchmark contamination, it is often hard to take a research lab's claims about model performance at face value. It's difficult to know whether a model gets good benchmark performance because it is generally capable or because its pre-training data was contaminated and it overfit on the benchmarks.

One solution to this problem is for benchmark creators to release their datasets in stages. For example, a benchmark creator could release 50% of their dataset upon release, and then release the remaining 50% in two stages, 25% one year later and 25% two years later. This would enable model evaluators to check for benchmark contamination by comparing performance on the subset of data released prior to the training cutoff vs. the subset released after the training cutoff. It would also give us a better understanding of how well models are actually performing.

One last point - this staged release process wouldn't be anywhere near as helpful for benchmarks created by scraping the web, as even the later-released data subsets could be found in the training data. But it should be useful for other kinds of benchmarks.

FieldValue
text There's been a lot of discussion about benchmark contamination, where models are trained on the data they are ultimately evaluated on. For example, a [recent paper](https://twitter.com/hughbzhang/status/1785877026794356858) showed that models performed substantially better on the public GSM8K vs GSM1K, which was a benchmark recently created by Scale AI to match GSM8K on difficulty and other measures. Because of these concerns about benchmark contamination, it is often hard to take a research la…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-02
username_encoded Z0FBQUFBQm5LakwyVFBJbmk2TGNSOUNydHk1Smt4RnBrZ040QXhqb3ZUejVvRUxrX0JRbU4xNmZZakhkUmt3RmxfYWdWVHc3a2lIM3g2b0pESUxMdlBQR1E1Sk1hdWlWWXc9PQ==
url_encoded Z0FBQUFBQm5Lak9HV3ZvTmxUQjJEdEJFX2ttY25UZXJNNFV5TkU1NTYxc0UyZlFRY3dvdFhJVDlDQnNwaVpQR0JUdlJpRkJmSWxrNTVSMDY4OGg0SkVjcTU3S2ZLY0NWVXUycnE3MkFLbEhoOWRKczYtUjJ5WWpEcGlwWEQyZlVCNFBqWjVrZGtqQjhPQlc5TFU1eGt1OVhjNU8wcjFwUy1jbnlKMmozcC1YQjV1b2Y3RE0tUDBQVTdESzBMT1JEdm9MWHA0cTdVck5WUnFkN1RhdWl4aVRxbDVyTkxCTGsyQT09

Raw Record

{
  "text": "There's been a lot of discussion about benchmark contamination, where models are trained on the data they are ultimately evaluated on. For example, a [recent paper](https://twitter.com/hughbzhang/status/1785877026794356858) showed that models performed substantially better on the public GSM8K vs GSM1K, which was a benchmark recently created by Scale AI to match GSM8K on difficulty and other measures.\n\nBecause of these concerns about benchmark contamination, it is often hard to take a research lab's claims about model performance at face value. It's difficult to know whether a model gets good benchmark performance because it is generally capable or because its pre-training data was contaminated and it overfit on the benchmarks.\n\nOne solution to this problem is for benchmark creators to release their datasets in stages. For example, a benchmark creator could release 50% of their dataset upon release, and then release the remaining 50% in two stages, 25% one year later and 25% two years later. This would enable model evaluators to check for benchmark contamination by comparing performance on the subset of data released prior to the training cutoff vs. the subset released after the training cutoff. It would also give us a better understanding of how well models are actually performing.\n\nOne last point - this staged release process wouldn't be anywhere near as helpful for benchmarks created by scraping the web, as even the later-released data subsets could be found in the training data. But it should be useful for other kinds of benchmarks.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-02",
  "username_encoded": "Z0FBQUFBQm5LakwyVFBJbmk2TGNSOUNydHk1Smt4RnBrZ040QXhqb3ZUejVvRUxrX0JRbU4xNmZZakhkUmt3RmxfYWdWVHc3a2lIM3g2b0pESUxMdlBQR1E1Sk1hdWlWWXc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9HV3ZvTmxUQjJEdEJFX2ttY25UZXJNNFV5TkU1NTYxc0UyZlFRY3dvdFhJVDlDQnNwaVpQR0JUdlJpRkJmSWxrNTVSMDY4OGg0SkVjcTU3S2ZLY0NWVXUycnE3MkFLbEhoOWRKczYtUjJ5WWpEcGlwWEQyZlVCNFBqWjVrZGtqQjhPQlc5TFU1eGt1OVhjNU8wcjFwUy1jbnlKMmozcC1YQjV1b2Y3RE0tUDBQVTdESzBMT1JEdm9MWHA0cTdVck5WUnFkN1RhdWl4aVRxbDVyTkxCTGsyQT09"
}

Entry Information