Row 34437

Row ID: 34437 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 34437 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores.

https://crfm.stanford.edu/2024/05/01/helm-mmlu.html

Here's the ranking of the models based on their HELM scores, from highest to lowest:

1. Claude 3 Opus (84.6 HELM) 2. GPT-4 (0613) (82.4 HELM) 3. Llama 3 (70B) (79.3 HELM) 4. PaLM 2 Unicorn (78.6 HELM) 5. Mixtral (8x22B) (77.8 HELM) 6. Qwen1.5 (72B) (77.4 HELM) 7. Yi (34B) (76.2 HELM) 8. Claude 3 Sonnet (75.9 HELM) 9. Qwen1.5 (32B) (74.4 HELM) 10. Claude 3 Haiku (73.8 HELM) 11. Claude 2.1 (73.5 HELM) 12. Mixtral (8x7B) (71.7 HELM) 13. Gemini 1.0 Pro (70.0 HELM) 14. Llama 2 (70B) (69.5 HELM) 15. Qwen1.5 (14B) (68.6 HELM) 16. Claude Instant (68.8 HELM) 17. Llama 3 (8B) (66.8 HELM) 18. Gemma (7B) (66.1 HELM) 19. Yi (6B) (64.0 HELM) 20. Qwen1.5 (7B) (62.6 HELM) 21. Phi 2 2.7B (58.4 HELM) 22. Mistral v0.1 (7B) (56.6 HELM) 23. Llama 2 (13B) (55.4 HELM) 24. Llama 2 (7B) (45.8 HELM) 25. OLMo (7B) (29.5 HELM)

This ranking is based on the assumption that the HELM scores are more reliable and comparable across models due to the standardized evaluation methodology used by the authors.

https://www.reddit.com/r/OpenAI/s/BO4ArEwPp7

FieldValue
text The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores. https://crfm.stanford.edu/2024/05/01/helm-mmlu.html Here's the ranking of the models based on their HELM scores, from highest to lowest: 1. Claude 3 Opus (84.6 HELM) 2. GPT-4 (0613) (82.4 HELM) 3. Llama 3 (70B) (79.3 HELM) 4. PaLM 2 Unicorn (78.6 HELM) 5. Mixtra…
label r/openai
dataType post
communityName r/OpenAI
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1JY1A3ZXozVXI4WlN4M3ZtY0hqYTVUUlR4ZllQRU05cjlzRjNkQ2JnanIxcmpTanF2Y1lTdFBhSkI2ODVsTHVWR0NBbkkxd09aN0ZRYlVfdE94dU9YaHVKS3dTWkpVMjB3M1p6OTdSWVNzQjg9
url_encoded Z0FBQUFBQm5Lak9YenRTTWNVeHZiT0QycGppWVhEQzh6RTdzcFRfWFh2cnVFU2hxQk55TGNXVlBmNHVWa3I1MHlqeWxud21iZWRaaUVqUFAyX09xdzluclVXQmZaeVdlV09kYlNYNnotVEg5d0NqMnVYSlQydWNJa1RkVHFmWXdrMUdBLWdwOU5yOFAzMnVpWmVVRDkwVW5SU0dwZUZZMGVvdDB0NTAzT1N1TERoT0FaXzJhWWlJV3FNX29oTFI3OWlLeG1VTE40ZnpYczVYWkNWRnJxWnJ0R3F0d2N5SzgtQT09

Raw Record

{
  "text": "The authors mention that they performed a comprehensive MMLU evaluation using the HELM framework with standardized prompts and full transparency, addressing issues with inconsistencies and lack of comparability in the reported scores.\n\nhttps://crfm.stanford.edu/2024/05/01/helm-mmlu.html\n\nHere's the ranking of the models based on their HELM scores, from highest to lowest:\n\n1. Claude 3 Opus (84.6 HELM)\n2. GPT-4 (0613) (82.4 HELM)\n3. Llama 3 (70B) (79.3 HELM)\n4. PaLM 2 Unicorn (78.6 HELM)\n5. Mixtral (8x22B) (77.8 HELM)\n6. Qwen1.5 (72B) (77.4 HELM)\n7. Yi (34B) (76.2 HELM)\n8. Claude 3 Sonnet (75.9 HELM)\n9. Qwen1.5 (32B) (74.4 HELM)\n10. Claude 3 Haiku (73.8 HELM)\n11. Claude 2.1 (73.5 HELM)\n12. Mixtral (8x7B) (71.7 HELM)\n13. Gemini 1.0 Pro (70.0 HELM)\n14. Llama 2 (70B) (69.5 HELM)\n15. Qwen1.5 (14B) (68.6 HELM)\n16. Claude Instant (68.8 HELM)\n17. Llama 3 (8B) (66.8 HELM)\n18. Gemma (7B) (66.1 HELM)\n19. Yi (6B) (64.0 HELM)\n20. Qwen1.5 (7B) (62.6 HELM)\n21. Phi 2 2.7B (58.4 HELM)\n22. Mistral v0.1 (7B) (56.6 HELM)\n23. Llama 2 (13B) (55.4 HELM)\n24. Llama 2 (7B) (45.8 HELM)\n25. OLMo (7B) (29.5 HELM)\n\nThis ranking is based on the assumption that the HELM scores are more reliable and comparable across models due to the standardized evaluation methodology used by the authors.\n\nhttps://www.reddit.com/r/OpenAI/s/BO4ArEwPp7",
  "label": "r/openai",
  "dataType": "post",
  "communityName": "r/OpenAI",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1JY1A3ZXozVXI4WlN4M3ZtY0hqYTVUUlR4ZllQRU05cjlzRjNkQ2JnanIxcmpTanF2Y1lTdFBhSkI2ODVsTHVWR0NBbkkxd09aN0ZRYlVfdE94dU9YaHVKS3dTWkpVMjB3M1p6OTdSWVNzQjg9",
  "url_encoded": "Z0FBQUFBQm5Lak9YenRTTWNVeHZiT0QycGppWVhEQzh6RTdzcFRfWFh2cnVFU2hxQk55TGNXVlBmNHVWa3I1MHlqeWxud21iZWRaaUVqUFAyX09xdzluclVXQmZaeVdlV09kYlNYNnotVEg5d0NqMnVYSlQydWNJa1RkVHFmWXdrMUdBLWdwOU5yOFAzMnVpWmVVRDkwVW5SU0dwZUZZMGVvdDB0NTAzT1N1TERoT0FaXzJhWWlJV3FNX29oTFI3OWlLeG1VTE40ZnpYczVYWkNWRnJxWnJ0R3F0d2N5SzgtQT09"
}

Entry Information