Row 7466
Content Data
This page contains data entry 7466 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
**Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack)
**Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/)
**Description**:
>Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo struggles with this benchmark.
>NIAN creates a list of limericks from a large database of limericks and asks a question about a specific limerick that has been placed at a test location. Each test will typically use 5 to 10 test limericks placed at 5 to 10 locations in the prompt. Each test is repeated 2-10 times.
| Field | Value |
|---|---|
| text | **Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack) **Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/) **Description**: >Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo str… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-16 |
| username_encoded | Z0FBQUFBQm5LakwzTktXUXFvZndhV29wTGdzekVOTzk4a2VKc3lZS1FiWFNOWDMtZWpJUVpXZHFtTHNJUkdXa0lsRUViUUNYclJXaHdXLXp5NUpvUldvN1NockdDVHJqcXc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9IY2Z1ZHpHYmRDdVN5QWozdDVkeTROX2dLZ0hBdkxNR05IRnIyTldscnBBNnNHZlZwN2dvUDB1Vm1VV0h2OE45dUtldUJYUFNnU2QtdUl4LThPR0dXTm84d2s5d0pGVzl2SkJNT2FyeWtLYnhzckV6bnIxaEhRdzlQdXFyTHhQS2NaMm15MkZnNWdoamdYMTE2UkxmdjNSbWRvQi1FSndybUkzM25qSWItREZWUjg3U0Y1VVFGdlFnM3pTWC1OMDMx |
Raw Record
{
"text": "**Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack)\n\n**Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/)\n\n**Description**:\n\n>Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo struggles with this benchmark.\n\n>NIAN creates a list of limericks from a large database of limericks and asks a question about a specific limerick that has been placed at a test location. Each test will typically use 5 to 10 test limericks placed at 5 to 10 locations in the prompt. Each test is repeated 2-10 times.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-16",
"username_encoded": "Z0FBQUFBQm5LakwzTktXUXFvZndhV29wTGdzekVOTzk4a2VKc3lZS1FiWFNOWDMtZWpJUVpXZHFtTHNJUkdXa0lsRUViUUNYclJXaHdXLXp5NUpvUldvN1NockdDVHJqcXc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9IY2Z1ZHpHYmRDdVN5QWozdDVkeTROX2dLZ0hBdkxNR05IRnIyTldscnBBNnNHZlZwN2dvUDB1Vm1VV0h2OE45dUtldUJYUFNnU2QtdUl4LThPR0dXTm84d2s5d0pGVzl2SkJNT2FyeWtLYnhzckV6bnIxaEhRdzlQdXFyTHhQS2NaMm15MkZnNWdoamdYMTE2UkxmdjNSbWRvQi1FSndybUkzM25qSWItREZWUjg3U0Y1VVFGdlFnM3pTWC1OMDMx"
}
Entry Information
- Entry ID: 7466
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000