Row 94443

Row ID: 94443 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 94443 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set.

For the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequences on both accuracy and penalties and error propagation etc etc.

FieldValue
text there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set. For the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequence…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak10QTgzbjNfMEl4NWhCbF90MThCNlJRTE15aXg5SE1uZlFXdjdPWmQ2Mko2X2xJbFh0cG9BLXk3c0E1U0paMlUwYUY1Mk5NT3RXbTZSTHRNQkJaWU9VNFZsUW5hZEFrRTJqbzUwbWdKRVkyUFk9
url_encoded Z0FBQUFBQm5Lak9fUTUyUkpxd1hja09zdU1La2YxdDJDTFhzQjUyX1BKVWRRTEFoYXBrM2tqeEswb0p6aVpXQUxjY0FkWFBKRVFjYlA0UFBPTnZ2NE8zXzFTVEg3ZjU0S053TXZBeXdHc3dDZEk5ZW9PbHpsYXpYcWpOekxrQnIzbXlZa2ZSTy16eVBNTW1ZVFVGbndQZGhseDhOenVVeEh2TmxkbnZpS2VuSWlDVEJJSVpDbG01RjdUT0RMMDFCMHhKdVFpcjYwaS1MZ0ROSnF3ZHdnajFEbmRvVm9tbjZsRUd0STBhYVVjazB6ZExTTkVaY0NxYz0=

Raw Record

{
  "text": "there is no fit in kNN, you need the query point in order to get the neighbours, and those points are in the test set. \n\nFor the scaling part: that's not really data leakage, is it? You can think about data leakage as a non-empty intersection between any pair of the three sets (Train, Test, Validation). But you're right, scaling is a major issue for metrics where p >=2 due to the fact that bigger features will contribute more and more to the distance (think about infinity norm), with consequences on both accuracy and penalties and error propagation etc etc.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak10QTgzbjNfMEl4NWhCbF90MThCNlJRTE15aXg5SE1uZlFXdjdPWmQ2Mko2X2xJbFh0cG9BLXk3c0E1U0paMlUwYUY1Mk5NT3RXbTZSTHRNQkJaWU9VNFZsUW5hZEFrRTJqbzUwbWdKRVkyUFk9",
  "url_encoded": "Z0FBQUFBQm5Lak9fUTUyUkpxd1hja09zdU1La2YxdDJDTFhzQjUyX1BKVWRRTEFoYXBrM2tqeEswb0p6aVpXQUxjY0FkWFBKRVFjYlA0UFBPTnZ2NE8zXzFTVEg3ZjU0S053TXZBeXdHc3dDZEk5ZW9PbHpsYXpYcWpOekxrQnIzbXlZa2ZSTy16eVBNTW1ZVFVGbndQZGhseDhOenVVeEh2TmxkbnZpS2VuSWlDVEJJSVpDbG01RjdUT0RMMDFCMHhKdVFpcjYwaS1MZ0ROSnF3ZHdnajFEbmRvVm9tbjZsRUd0STBhYVVjazB6ZExTTkVaY0NxYz0="
}

Entry Information