Row 7780
Content Data
This page contains data entry 7780 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hey everyone,
I am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study.
Upon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored the accuracy of the validation set and used it for early stopping. But the correct method is to make the validation set from the training set, right? I got confused because I read that in some cases, some people use validation set and test set for the same purpose.
I was surprised that neither my coauthors nor the three reviewers spotted this error during the peer-review process. Does the oversight of neither my coauthors nor the reviewers identifying the error imply that its severity can be downplayed?
Additionally, we conducted measurements on five tests to average our results - but I'm uncertain about whether this helps mitigate the impact of the error. Your input on this matter would be greatly appreciated.
I am ready to mention this error during my presentation for transparency, but I feel a bit dumb.
EDIT: Using that validation set for early stopping or not hasn't changed the capability of our model to be robust on unseen data. So... is the error negligeable?
| Field | Value |
|---|---|
| text | Hey everyone, I am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study. Upon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-17 |
| username_encoded | Z0FBQUFBQm5LakwzR2EyTmFWMUs0dE9NZU4wZXBvRURFNlRnT2NWRGxNQmlURjJKazVfaHNnXzdCeFpqVF9ybVR4aWVNbEowTzBMdWY3R0hTVzc0Mzh5dnZkNE02cm1wVkwtb3NzaUg4ZWlweVE2VUpCUHlCR3M9 |
| url_encoded | Z0FBQUFBQm5Lak9IRDRDelpCR3M2WGJhWk9ZYXJaa1kxamtfeWpkeXRJZVA0UXZ4VHNKdkQyUGNDaDc2Yk1GdDFFNlRoUnQzaEJ6VF95OHU2NV95RV92eTB3YlJTcHI1MnAxbng2S3FtdFh1dTUzVlQ4bDFlWjhUSHRUSDBYcUhqel9XT0NtVG1KT2VhVEpvN18tS0o2VTQxM2l5aTRxZ3V4XzFBbFhHQV9VR1oxeThneGEzYlBsZjk1dnNxVnBTeGFRblRncV9LZC00SzdQLW01SVBTcjNxSU9zS1RPc0hqZz09 |
Raw Record
{
"text": "Hey everyone,\n\nI am sharing a recent experience as a PhD student and seeking advice on interpreting a methodological oversight in our study.\n\nUpon reflection on a recent paper we submitted (I already submitted the camera-ready version), I realized we made a methodological error. Specifically, we used a validation set derived from the test set, potentially introducing bias into our results because we have heterogeneous datasets (training and test datasets are from different origins). We monitored the accuracy of the validation set and used it for early stopping. But the correct method is to make the validation set from the training set, right? I got confused because I read that in some cases, some people use validation set and test set for the same purpose.\n\nI was surprised that neither my coauthors nor the three reviewers spotted this error during the peer-review process. Does the oversight of neither my coauthors nor the reviewers identifying the error imply that its severity can be downplayed?\n\nAdditionally, we conducted measurements on five tests to average our results - but I'm uncertain about whether this helps mitigate the impact of the error. Your input on this matter would be greatly appreciated.\n\nI am ready to mention this error during my presentation for transparency, but I feel a bit dumb.\n\nEDIT: Using that validation set for early stopping or not hasn't changed the capability of our model to be robust on unseen data. So... is the error negligeable?",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-17",
"username_encoded": "Z0FBQUFBQm5LakwzR2EyTmFWMUs0dE9NZU4wZXBvRURFNlRnT2NWRGxNQmlURjJKazVfaHNnXzdCeFpqVF9ybVR4aWVNbEowTzBMdWY3R0hTVzc0Mzh5dnZkNE02cm1wVkwtb3NzaUg4ZWlweVE2VUpCUHlCR3M9",
"url_encoded": "Z0FBQUFBQm5Lak9IRDRDelpCR3M2WGJhWk9ZYXJaa1kxamtfeWpkeXRJZVA0UXZ4VHNKdkQyUGNDaDc2Yk1GdDFFNlRoUnQzaEJ6VF95OHU2NV95RV92eTB3YlJTcHI1MnAxbng2S3FtdFh1dTUzVlQ4bDFlWjhUSHRUSDBYcUhqel9XT0NtVG1KT2VhVEpvN18tS0o2VTQxM2l5aTRxZ3V4XzFBbFhHQV9VR1oxeThneGEzYlBsZjk1dnNxVnBTeGFRblRncV9LZC00SzdQLW01SVBTcjNxSU9zS1RPc0hqZz09"
}
Entry Information
- Entry ID: 7780
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000