Row 7466

Row ID: 7466 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 7466 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack)

**Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/)

**Description**:

>Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo struggles with this benchmark.

>NIAN creates a list of limericks from a large database of limericks and asks a question about a specific limerick that has been placed at a test location. Each test will typically use 5 to 10 test limericks placed at 5 to 10 locations in the prompt. Each test is repeated 2-10 times.

FieldValue
text **Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack) **Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/) **Description**: >Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo str…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-16
username_encoded Z0FBQUFBQm5LakwzTktXUXFvZndhV29wTGdzekVOTzk4a2VKc3lZS1FiWFNOWDMtZWpJUVpXZHFtTHNJUkdXa0lsRUViUUNYclJXaHdXLXp5NUpvUldvN1NockdDVHJqcXc9PQ==
url_encoded Z0FBQUFBQm5Lak9IY2Z1ZHpHYmRDdVN5QWozdDVkeTROX2dLZ0hBdkxNR05IRnIyTldscnBBNnNHZlZwN2dvUDB1Vm1VV0h2OE45dUtldUJYUFNnU2QtdUl4LThPR0dXTm84d2s5d0pGVzl2SkJNT2FyeWtLYnhzckV6bnIxaEhRdzlQdXFyTHhQS2NaMm15MkZnNWdoamdYMTE2UkxmdjNSbWRvQi1FSndybUkzM25qSWItREZWUjg3U0Y1VVFGdlFnM3pTWC1OMDMx

Raw Record

{
  "text": "**Code**: [https://github.com/llmonpy/needle-in-a-needlestack](https://github.com/llmonpy/needle-in-a-needlestack)\n\n**Website**: [https://nian.llmonpy.ai/](https://nian.llmonpy.ai/)\n\n**Description**:\n\n>Needle in a haystack (NIAH) has been a wildly popular test for evaluating how effectively LLMs can pay attention to the content in their context window. As LLMs have improved NIAH has become too easy. **Needle in a Needlestack** (**NIAN**) is a new, more challenging benchmark. Even GPT-4-turbo struggles with this benchmark.\n\n>NIAN creates a list of limericks from a large database of limericks and asks a question about a specific limerick that has been placed at a test location. Each test will typically use 5 to 10 test limericks placed at 5 to 10 locations in the prompt. Each test is repeated 2-10 times.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-16",
  "username_encoded": "Z0FBQUFBQm5LakwzTktXUXFvZndhV29wTGdzekVOTzk4a2VKc3lZS1FiWFNOWDMtZWpJUVpXZHFtTHNJUkdXa0lsRUViUUNYclJXaHdXLXp5NUpvUldvN1NockdDVHJqcXc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9IY2Z1ZHpHYmRDdVN5QWozdDVkeTROX2dLZ0hBdkxNR05IRnIyTldscnBBNnNHZlZwN2dvUDB1Vm1VV0h2OE45dUtldUJYUFNnU2QtdUl4LThPR0dXTm84d2s5d0pGVzl2SkJNT2FyeWtLYnhzckV6bnIxaEhRdzlQdXFyTHhQS2NaMm15MkZnNWdoamdYMTE2UkxmdjNSbWRvQi1FSndybUkzM25qSWItREZWUjg3U0Y1VVFGdlFnM3pTWC1OMDMx"
}

Entry Information