Row 7780

Row ID: 7780 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 7780 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Hey everyone,

I am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study.

Upon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored the accuracy of the validation set and used it for early stopping. But the correct method is to make the validation set from the training set, right? I got confused because I read that in some cases, some people use validation set and test set for the same purpose.

I was surprised that neither my coauthors nor the three reviewers spotted this error during the peer-review process. Does the oversight of neither my coauthors nor the reviewers identifying the error imply that its severity can be downplayed?

Additionally, we conducted measurements on five tests to average our results - but I'm uncertain about whether this helps mitigate the impact of the error. Your input on this matter would be greatly appreciated.

I am ready to mention this error during my presentation for transparency, but I feel a bit dumb.

EDIT: Using that validation set for early stopping or not hasn't changed the capability of our model to be robust on unseen data. So... is the error negligeable?

FieldValue
text Hey everyone, I am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study. Upon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-17
username_encoded Z0FBQUFBQm5LakwzR2EyTmFWMUs0dE9NZU4wZXBvRURFNlRnT2NWRGxNQmlURjJKazVfaHNnXzdCeFpqVF9ybVR4aWVNbEowTzBMdWY3R0hTVzc0Mzh5dnZkNE02cm1wVkwtb3NzaUg4ZWlweVE2VUpCUHlCR3M9
url_encoded Z0FBQUFBQm5Lak9IRDRDelpCR3M2WGJhWk9ZYXJaa1kxamtfeWpkeXRJZVA0UXZ4VHNKdkQyUGNDaDc2Yk1GdDFFNlRoUnQzaEJ6VF95OHU2NV95RV92eTB3YlJTcHI1MnAxbng2S3FtdFh1dTUzVlQ4bDFlWjhUSHRUSDBYcUhqel9XT0NtVG1KT2VhVEpvN18tS0o2VTQxM2l5aTRxZ3V4XzFBbFhHQV9VR1oxeThneGEzYlBsZjk1dnNxVnBTeGFRblRncV9LZC00SzdQLW01SVBTcjNxSU9zS1RPc0hqZz09

Raw Record

{
  "text": "Hey everyone,\n\nI am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study.\n\nUpon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored the accuracy of the validation set and used it for early stopping. But the correct method is to make the validation set from the training set, right? I got confused because I read that in some cases, some people use validation set and test set for the same purpose.\n\nI was surprised that neither my coauthors nor the three reviewers spotted this error during the peer-review process. Does the oversight of neither my coauthors nor the reviewers identifying the error imply that its severity can be downplayed?\n\nAdditionally, we conducted measurements on five tests to average our results - but I'm uncertain about whether this helps mitigate the impact of the error. Your input on this matter would be greatly appreciated.\n\nI am ready to mention this error during my presentation for transparency, but I feel a bit dumb.\n\nEDIT: Using that validation set for early stopping or not hasn't changed the capability of our model to be robust on unseen data. So... is the error negligeable?",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-17",
  "username_encoded": "Z0FBQUFBQm5LakwzR2EyTmFWMUs0dE9NZU4wZXBvRURFNlRnT2NWRGxNQmlURjJKazVfaHNnXzdCeFpqVF9ybVR4aWVNbEowTzBMdWY3R0hTVzc0Mzh5dnZkNE02cm1wVkwtb3NzaUg4ZWlweVE2VUpCUHlCR3M9",
  "url_encoded": "Z0FBQUFBQm5Lak9IRDRDelpCR3M2WGJhWk9ZYXJaa1kxamtfeWpkeXRJZVA0UXZ4VHNKdkQyUGNDaDc2Yk1GdDFFNlRoUnQzaEJ6VF95OHU2NV95RV92eTB3YlJTcHI1MnAxbng2S3FtdFh1dTUzVlQ4bDFlWjhUSHRUSDBYcUhqel9XT0NtVG1KT2VhVEpvN18tS0o2VTQxM2l5aTRxZ3V4XzFBbFhHQV9VR1oxeThneGEzYlBsZjk1dnNxVnBTeGFRblRncV9LZC00SzdQLW01SVBTcjNxSU9zS1RPc0hqZz09"
}

Entry Information