Row 90524
Content Data
This page contains data entry 90524 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days.
I tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well).
| Field | Value |
|---|---|
| text | The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days. I tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well). |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1yMkhFUEF4RnZTZTJyUnptNjdaa0EtZFpTZUU2d3hzRmthUnlqWFoxa1BXNlRJaE5CcVY2TThta01kUHJZeVlhWnFHckFTZ3NLZDdEMFc3YXZ1Z3lIRXc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak85RU5QTktoQXRES3V0Mk9NakoyakxjR1UtOTVFazFpWThkalBmRkJnZ18xLW9yem5Gc2lNMHdXekNjZnViZUxZQmVWUzhZczZYR195TWZwNV9yemlsTWN3SnVyNGowekZad1BYQ0FmQjdZS3NNNElRV0wzM0Z0NW5ZRGpCNGZ5Yy1oSHR4cGJjSUYtdDF6M2daRDRuY3R1M0xXY19aUng3SkRqd2JpRjFfdm1sLS14WVRLdGFzaVgwdXdNODczMUozOHhOODkwTUlzNzExOVU3SUpXcV9vZz09 |
Raw Record
{
"text": "The benchmark I'm using is the MIRACL dataset; it seems to be pretty standard for multilingual document retrieval these days.\n\nI tried doing what you proposed, that is, just taking the positive and negative passages for each query and ranking them, then calculating the score. The problem is that, as you rightly pointed out, this leads to a misleading high score considering we're not using the entire corpus (i.e., the entire positive and negative passages of other queries as well).",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1yMkhFUEF4RnZTZTJyUnptNjdaa0EtZFpTZUU2d3hzRmthUnlqWFoxa1BXNlRJaE5CcVY2TThta01kUHJZeVlhWnFHckFTZ3NLZDdEMFc3YXZ1Z3lIRXc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak85RU5QTktoQXRES3V0Mk9NakoyakxjR1UtOTVFazFpWThkalBmRkJnZ18xLW9yem5Gc2lNMHdXekNjZnViZUxZQmVWUzhZczZYR195TWZwNV9yemlsTWN3SnVyNGowekZad1BYQ0FmQjdZS3NNNElRV0wzM0Z0NW5ZRGpCNGZ5Yy1oSHR4cGJjSUYtdDF6M2daRDRuY3R1M0xXY19aUng3SkRqd2JpRjFfdm1sLS14WVRLdGFzaVgwdXdNODczMUozOHhOODkwTUlzNzExOVU3SUpXcV9vZz09"
}
Entry Information
- Entry ID: 90524
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000