Row 16324
Content Data
This page contains data entry 16324 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
[https:\/\/leaderboard.lmsys.org\/](https://preview.redd.it/0av7nm84hn1d1.png?width=4000&format=png&auto=webp&s=1b7eaaf5c891bed8dc59be6ee7d1970cf9ae5d19)
Chatbot Arena introduced rankings for the **Hard Prompts category**. This benchmark tests models on highly complex, multi-faceted prompts requiring deep reasoning.
**The prompts are carefully curated** based on seven key criteria: specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy, and real-world application.
**Key Details:**
* The new category ranks models on their ability to handle **difficult prompts** * **Over 20%** of Arena prompts qualify as "hard" * **Llama 3 and Gemini 1.5 Pro** struggle with increased prompt complexity * **GPT-4o shows a slight improvement** in performance while **GPT-4-Turbo sees some decline** * **Claude 3 Opus remains stable** under increased prompt complexity
| Field | Value |
|---|---|
| text | [https:\/\/leaderboard.lmsys.org\/](https://preview.redd.it/0av7nm84hn1d1.png?width=4000&format=png&auto=webp&s=1b7eaaf5c891bed8dc59be6ee7d1970cf9ae5d19) Chatbot Arena introduced rankings for the **Hard Prompts category**. This benchmark tests models on highly complex, multi-faceted prompts requiring deep reasoning. **The prompts are carefully curated** based on seven key criteria: specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy, and real-world applic… |
| label | r/chatgpt |
| dataType | post |
| communityName | r/ChatGPT |
| datetime | 2024-05-20 |
| username_encoded | Z0FBQUFBQm5Lakw5R1dYcnpFRXVYQnMzMVZnX3d4bXh1WFlHR3A5TXhsU1l3WVNNbWFpbnV2YXJ6amxZNmdkZXFFc3U0YXhEWGpSUElLVFR4RzhYOWQxaDk4TG5sdmRNN1Z5SEhpR010UXZIRE1NWWFGVzFnVjQ9 |
| url_encoded | Z0FBQUFBQm5Lak9NZ2YwWjA2MVQySUFSS1JMbTBJa2ZXN0dGek92SVBMSGFReUNVdUUtT1FEdDJabk9VSnJEZkpDTTVPSjRPU3dOTHhsRHpLZHRYZnRzZkM3Rmo3SG5TRXJTc0p6TXJQbkdrRFdXM3NibUE5azdWemVUNTZXdHphVjU1V0ttaTFGX3R2bU82Yjk2NDVIMl9BbXl2RGFmRGlYVWN5ekExeHFEVjdPZlFMQVYyb0R1Z3hIMThBclo3UFZZalZuNXhZVHFiSXZPWTNpQURNbWtLY095aFVxdW94Zz09 |
Raw Record
{
"text": "[https:\\/\\/leaderboard.lmsys.org\\/](https://preview.redd.it/0av7nm84hn1d1.png?width=4000&format=png&auto=webp&s=1b7eaaf5c891bed8dc59be6ee7d1970cf9ae5d19)\n\nChatbot Arena introduced rankings for the **Hard Prompts category**. This benchmark tests models on highly complex, multi-faceted prompts requiring deep reasoning.\n\n**The prompts are carefully curated** based on seven key criteria: specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy, and real-world application.\n\n**Key Details:**\n\n* The new category ranks models on their ability to handle **difficult prompts**\n* **Over 20%** of Arena prompts qualify as \"hard\"\n* **Llama 3 and Gemini 1.5 Pro** struggle with increased prompt complexity\n* **GPT-4o shows a slight improvement** in performance while **GPT-4-Turbo sees some decline**\n* **Claude 3 Opus remains stable** under increased prompt complexity",
"label": "r/chatgpt",
"dataType": "post",
"communityName": "r/ChatGPT",
"datetime": "2024-05-20",
"username_encoded": "Z0FBQUFBQm5Lakw5R1dYcnpFRXVYQnMzMVZnX3d4bXh1WFlHR3A5TXhsU1l3WVNNbWFpbnV2YXJ6amxZNmdkZXFFc3U0YXhEWGpSUElLVFR4RzhYOWQxaDk4TG5sdmRNN1Z5SEhpR010UXZIRE1NWWFGVzFnVjQ9",
"url_encoded": "Z0FBQUFBQm5Lak9NZ2YwWjA2MVQySUFSS1JMbTBJa2ZXN0dGek92SVBMSGFReUNVdUUtT1FEdDJabk9VSnJEZkpDTTVPSjRPU3dOTHhsRHpLZHRYZnRzZkM3Rmo3SG5TRXJTc0p6TXJQbkdrRFdXM3NibUE5azdWemVUNTZXdHphVjU1V0ttaTFGX3R2bU82Yjk2NDVIMl9BbXl2RGFmRGlYVWN5ekExeHFEVjdPZlFMQVYyb0R1Z3hIMThBclo3UFZZalZuNXhZVHFiSXZPWTNpQURNbWtLY095aFVxdW94Zz09"
}
Entry Information
- Entry ID: 16324
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000