Row 34437
Content Data
This page contains data entry 34437 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores.
https://crfm.stanford.edu/2024/05/01/helm-mmlu.html
Here's the ranking of the models based on their HELM scores, from highest to lowest:
1. Claude 3 Opus (84.6 HELM) 2. GPT-4 (0613) (82.4 HELM) 3. Llama 3 (70B) (79.3 HELM) 4. PaLM 2 Unicorn (78.6 HELM) 5. Mixtral (8x22B) (77.8 HELM) 6. Qwen1.5 (72B) (77.4 HELM) 7. Yi (34B) (76.2 HELM) 8. Claude 3 Sonnet (75.9 HELM) 9. Qwen1.5 (32B) (74.4 HELM) 10. Claude 3 Haiku (73.8 HELM) 11. Claude 2.1 (73.5 HELM) 12. Mixtral (8x7B) (71.7 HELM) 13. Gemini 1.0 Pro (70.0 HELM) 14. Llama 2 (70B) (69.5 HELM) 15. Qwen1.5 (14B) (68.6 HELM) 16. Claude Instant (68.8 HELM) 17. Llama 3 (8B) (66.8 HELM) 18. Gemma (7B) (66.1 HELM) 19. Yi (6B) (64.0 HELM) 20. Qwen1.5 (7B) (62.6 HELM) 21. Phi 2 2.7B (58.4 HELM) 22. Mistral v0.1 (7B) (56.6 HELM) 23. Llama 2 (13B) (55.4 HELM) 24. Llama 2 (7B) (45.8 HELM) 25. OLMo (7B) (29.5 HELM)
This ranking is based on the assumption that the HELM scores are more reliable and comparable across models due to the standardized evaluation methodology used by the authors.
https://www.reddit.com/r/OpenAI/s/BO4ArEwPp7
| Field | Value |
|---|---|
| text | The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores. https://crfm.stanford.edu/2024/05/01/helm-mmlu.html Here's the ranking of the models based on their HELM scores, from highest to lowest: 1. Claude 3 Opus (84.6 HELM) 2. GPT-4 (0613) (82.4 HELM) 3. Llama 3 (70B) (79.3 HELM) 4. PaLM 2 Unicorn (78.6 HELM) 5. Mixtra… |
| label | r/openai |
| dataType | post |
| communityName | r/OpenAI |
| datetime | 2024-05-21 |
| username_encoded | Z0FBQUFBQm5Lak1JY1A3ZXozVXI4WlN4M3ZtY0hqYTVUUlR4ZllQRU05cjlzRjNkQ2JnanIxcmpTanF2Y1lTdFBhSkI2ODVsTHVWR0NBbkkxd09aN0ZRYlVfdE94dU9YaHVKS3dTWkpVMjB3M1p6OTdSWVNzQjg9 |
| url_encoded | Z0FBQUFBQm5Lak9YenRTTWNVeHZiT0QycGppWVhEQzh6RTdzcFRfWFh2cnVFU2hxQk55TGNXVlBmNHVWa3I1MHlqeWxud21iZWRaaUVqUFAyX09xdzluclVXQmZaeVdlV09kYlNYNnotVEg5d0NqMnVYSlQydWNJa1RkVHFmWXdrMUdBLWdwOU5yOFAzMnVpWmVVRDkwVW5SU0dwZUZZMGVvdDB0NTAzT1N1TERoT0FaXzJhWWlJV3FNX29oTFI3OWlLeG1VTE40ZnpYczVYWkNWRnJxWnJ0R3F0d2N5SzgtQT09 |
Raw Record
{
"text": "The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores.\n\nhttps://crfm.stanford.edu/2024/05/01/helm-mmlu.html\n\nHere's the ranking of the models based on their HELM scores, from highest to lowest:\n\n1. Claude 3 Opus (84.6 HELM)\n2. GPT-4 (0613) (82.4 HELM)\n3. Llama 3 (70B) (79.3 HELM)\n4. PaLM 2 Unicorn (78.6 HELM)\n5. Mixtral (8x22B) (77.8 HELM)\n6. Qwen1.5 (72B) (77.4 HELM)\n7. Yi (34B) (76.2 HELM)\n8. Claude 3 Sonnet (75.9 HELM)\n9. Qwen1.5 (32B) (74.4 HELM)\n10. Claude 3 Haiku (73.8 HELM)\n11. Claude 2.1 (73.5 HELM)\n12. Mixtral (8x7B) (71.7 HELM)\n13. Gemini 1.0 Pro (70.0 HELM)\n14. Llama 2 (70B) (69.5 HELM)\n15. Qwen1.5 (14B) (68.6 HELM)\n16. Claude Instant (68.8 HELM)\n17. Llama 3 (8B) (66.8 HELM)\n18. Gemma (7B) (66.1 HELM)\n19. Yi (6B) (64.0 HELM)\n20. Qwen1.5 (7B) (62.6 HELM)\n21. Phi 2 2.7B (58.4 HELM)\n22. Mistral v0.1 (7B) (56.6 HELM)\n23. Llama 2 (13B) (55.4 HELM)\n24. Llama 2 (7B) (45.8 HELM)\n25. OLMo (7B) (29.5 HELM)\n\nThis ranking is based on the assumption that the HELM scores are more reliable and comparable across models due to the standardized evaluation methodology used by the authors.\n\nhttps://www.reddit.com/r/OpenAI/s/BO4ArEwPp7",
"label": "r/openai",
"dataType": "post",
"communityName": "r/OpenAI",
"datetime": "2024-05-21",
"username_encoded": "Z0FBQUFBQm5Lak1JY1A3ZXozVXI4WlN4M3ZtY0hqYTVUUlR4ZllQRU05cjlzRjNkQ2JnanIxcmpTanF2Y1lTdFBhSkI2ODVsTHVWR0NBbkkxd09aN0ZRYlVfdE94dU9YaHVKS3dTWkpVMjB3M1p6OTdSWVNzQjg9",
"url_encoded": "Z0FBQUFBQm5Lak9YenRTTWNVeHZiT0QycGppWVhEQzh6RTdzcFRfWFh2cnVFU2hxQk55TGNXVlBmNHVWa3I1MHlqeWxud21iZWRaaUVqUFAyX09xdzluclVXQmZaeVdlV09kYlNYNnotVEg5d0NqMnVYSlQydWNJa1RkVHFmWXdrMUdBLWdwOU5yOFAzMnVpWmVVRDkwVW5SU0dwZUZZMGVvdDB0NTAzT1N1TERoT0FaXzJhWWlJV3FNX29oTFI3OWlLeG1VTE40ZnpYczVYWkNWRnJxWnJ0R3F0d2N5SzgtQT09"
}
Entry Information
- Entry ID: 34437
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000