Row 94443
Content Data
This page contains data entry 94443 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set.
For the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequences on both accuracy and penalties and error propagation etc etc.
| Field | Value |
|---|---|
| text | there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set. For the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequence… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak10QTgzbjNfMEl4NWhCbF90MThCNlJRTE15aXg5SE1uZlFXdjdPWmQ2Mko2X2xJbFh0cG9BLXk3c0E1U0paMlUwYUY1Mk5NT3RXbTZSTHRNQkJaWU9VNFZsUW5hZEFrRTJqbzUwbWdKRVkyUFk9 |
| url_encoded | Z0FBQUFBQm5Lak9fUTUyUkpxd1hja09zdU1La2YxdDJDTFhzQjUyX1BKVWRRTEFoYXBrM2tqeEswb0p6aVpXQUxjY0FkWFBKRVFjYlA0UFBPTnZ2NE8zXzFTVEg3ZjU0S053TXZBeXdHc3dDZEk5ZW9PbHpsYXpYcWpOekxrQnIzbXlZa2ZSTy16eVBNTW1ZVFVGbndQZGhseDhOenVVeEh2TmxkbnZpS2VuSWlDVEJJSVpDbG01RjdUT0RMMDFCMHhKdVFpcjYwaS1MZ0ROSnF3ZHdnajFEbmRvVm9tbjZsRUd0STBhYVVjazB6ZExTTkVaY0NxYz0= |
Raw Record
{
"text": "there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set. \n\nFor the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequences on both accuracy and penalties and error propagation etc etc.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak10QTgzbjNfMEl4NWhCbF90MThCNlJRTE15aXg5SE1uZlFXdjdPWmQ2Mko2X2xJbFh0cG9BLXk3c0E1U0paMlUwYUY1Mk5NT3RXbTZSTHRNQkJaWU9VNFZsUW5hZEFrRTJqbzUwbWdKRVkyUFk9",
"url_encoded": "Z0FBQUFBQm5Lak9fUTUyUkpxd1hja09zdU1La2YxdDJDTFhzQjUyX1BKVWRRTEFoYXBrM2tqeEswb0p6aVpXQUxjY0FkWFBKRVFjYlA0UFBPTnZ2NE8zXzFTVEg3ZjU0S053TXZBeXdHc3dDZEk5ZW9PbHpsYXpYcWpOekxrQnIzbXlZa2ZSTy16eVBNTW1ZVFVGbndQZGhseDhOenVVeEh2TmxkbnZpS2VuSWlDVEJJSVpDbG01RjdUT0RMMDFCMHhKdVFpcjYwaS1MZ0ROSnF3ZHdnajFEbmRvVm9tbjZsRUd0STBhYVVjazB6ZExTTkVaY0NxYz0="
}
Entry Information
- Entry ID: 94443
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000