Row 90891
Content Data
This page contains data entry 90891 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hm I see your problem. Based off of the miracl paper
> To build MIRACL, we hired native speakers as annotators to provide high-quality queries and relevance judgments. At a high level, our workflow comprised two phases: First, the annotators were asked to generate well-formed queries based on “prompts” (Clark et al., 2020) (details in Section 4.2). Then, they were asked to assess the relevance of the top-k query–passage pairs produced by an ensemble baseline retrieval system.
It looks like they used a model based approach to get their lists. I'm not sure what the common practice is here. I think you should use faiss to retrieve top k results. Then, you should also score any positive documents that are not retrieved in ur top k. Any retrieved document not captured in the original labelled list should default to a label of 0. Then you can compute the ideal ordering and ur ordering and compute ndcg.
The assumption here is that they have a good coverage of the positive documents for each query. If they don't, then ndcg@k can't be reliably calculated.
| Field | Value |
|---|---|
| text | Hm I see your problem. Based off of the miracl paper > To build MIRACL, we hired native speakers as annotators to provide high-quality queries and relevance judgments. At a high level, our workflow comprised two phases: First, the annotators were asked to generate well-formed queries based on “prompts” (Clark et al., 2020) (details in Section 4.2). Then, they were asked to assess the relevance of the top-k query–passage pairs produced by an ensemble baseline retrieval system. It looks like the… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1ybUVyWEFvbDBaQWpjdTFhUUpiTUpBbXFUd2UwQlltbG9zN2FISjJVMm1oZk42bmJfRnloYlJtNi1wUWs1MXJhVF9RWnJweVNDZjBaYVRudDZnRHBMcVE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak85Zmp4Y3I2S2FDeE05dEFhaUI4bmlGQVoyMm5NdXE0SGhfMHF5T3FPalJDYWlROXFsc0xpdHlCbzZMMW1yU0R1SFluWGxFODZ2Z1ZRdnJ6OWl4bFRDODlONk96TmRKYlhEdVMtdTNJQlpRYkZhaVRwNGk1UkZOZWktZklTcDkzbDFwZU5xOU94ZHdYTUpqYzJfVnkzYXVkMUtMTFRqSWlySzJrdG9iZUJrWUY2STh2TzBFNHZVSTZNaVQtNmM4MGRpeXduc2RXQ3NiZXY5Tlo5Snd4Z1gwZz09 |
Raw Record
{
"text": "Hm I see your problem. Based off of the miracl paper\n\n> To build MIRACL, we hired native speakers as annotators to provide high-quality queries and relevance judgments. At a high level, our workflow comprised two phases: First, the annotators were asked to generate well-formed queries based on “prompts” (Clark et al., 2020) (details in Section 4.2). Then, they were asked to assess the relevance of the top-k query–passage pairs produced by an ensemble baseline retrieval system.\n\nIt looks like they used a model based approach to get their lists. I'm not sure what the common practice is here. I think you should use faiss to retrieve top k results. Then, you should also score any positive documents that are not retrieved in ur top k. Any retrieved document not captured in the original labelled list should default to a label of 0. Then you can compute the ideal ordering and ur ordering and compute ndcg.\n\nThe assumption here is that they have a good coverage of the positive documents for each query. If they don't, then ndcg@k can't be reliably calculated.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1ybUVyWEFvbDBaQWpjdTFhUUpiTUpBbXFUd2UwQlltbG9zN2FISjJVMm1oZk42bmJfRnloYlJtNi1wUWs1MXJhVF9RWnJweVNDZjBaYVRudDZnRHBMcVE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak85Zmp4Y3I2S2FDeE05dEFhaUI4bmlGQVoyMm5NdXE0SGhfMHF5T3FPalJDYWlROXFsc0xpdHlCbzZMMW1yU0R1SFluWGxFODZ2Z1ZRdnJ6OWl4bFRDODlONk96TmRKYlhEdVMtdTNJQlpRYkZhaVRwNGk1UkZOZWktZklTcDkzbDFwZU5xOU94ZHdYTUpqYzJfVnkzYXVkMUtMTFRqSWlySzJrdG9iZUJrWUY2STh2TzBFNHZVSTZNaVQtNmM4MGRpeXduc2RXQ3NiZXY5Tlo5Snd4Z1gwZz09"
}
Entry Information
- Entry ID: 90891
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000