Row 90524

Row ID: 90524 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 90524 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days.

I tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well).

FieldValue
text The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days. I tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well).
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak1yMkhFUEF4RnZTZTJyUnptNjdaa0EtZFpTZUU2d3hzRmthUnlqWFoxa1BXNlRJaE5CcVY2TThta01kUHJZeVlhWnFHckFTZ3NLZDdEMFc3YXZ1Z3lIRXc9PQ==
url_encoded Z0FBQUFBQm5Lak85RU5QTktoQXRES3V0Mk9NakoyakxjR1UtOTVFazFpWThkalBmRkJnZ18xLW9yem5Gc2lNMHdXekNjZnViZUxZQmVWUzhZczZYR195TWZwNV9yemlsTWN3SnVyNGowekZad1BYQ0FmQjdZS3NNNElRV0wzM0Z0NW5ZRGpCNGZ5Yy1oSHR4cGJjSUYtdDF6M2daRDRuY3R1M0xXY19aUng3SkRqd2JpRjFfdm1sLS14WVRLdGFzaVgwdXdNODczMUozOHhOODkwTUlzNzExOVU3SUpXcV9vZz09

Raw Record

{
  "text": "The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days.\n\nI tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well).",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak1yMkhFUEF4RnZTZTJyUnptNjdaa0EtZFpTZUU2d3hzRmthUnlqWFoxa1BXNlRJaE5CcVY2TThta01kUHJZeVlhWnFHckFTZ3NLZDdEMFc3YXZ1Z3lIRXc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak85RU5QTktoQXRES3V0Mk9NakoyakxjR1UtOTVFazFpWThkalBmRkJnZ18xLW9yem5Gc2lNMHdXekNjZnViZUxZQmVWUzhZczZYR195TWZwNV9yemlsTWN3SnVyNGowekZad1BYQ0FmQjdZS3NNNElRV0wzM0Z0NW5ZRGpCNGZ5Yy1oSHR4cGJjSUYtdDF6M2daRDRuY3R1M0xXY19aUng3SkRqd2JpRjFfdm1sLS14WVRLdGFzaVgwdXdNODczMUozOHhOODkwTUlzNzExOVU3SUpXcV9vZz09"
}

Entry Information