Row 65148
Content Data
This page contains data entry 65148 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
phi models are a prime lesson on how to train on benchmarks to make your model look better than it is
for example they claim phi-3-mini is better than llama3 and mixtral (look at our mmlu scores!!!) yet on lmsys arena leaderboard llama3 is top 20 (mixtral just barely outside at 23) while phi-3 mini is... outside of top 50 in the ranks of mistral 7b
not to mention from several people that actually used or ran their own evals knows phi models are bad. eg https://twitter.com/abacaj/status/1792991309751284123
main lessons to be learned:
1. be very skeptical when someone claims to break the compute/perf frontier (see [here](https://pbs.twimg.com/media/GL0RUOAbEAAnq_n?format=jpg&name=4096x4096)). the best estimate of a model's performance is flops. 2. high quality synthetic data is not going to be the silver bullet that makes your model magically x times more efficient w/ flops. data is not compute agnostic 3. run your own evals and actually use the model, trusting benchmarks blindly is dumb
MSR has really fallen off from the days when they actually released good models like deberta, lets see if hiring most of the inflection team will get them back on track in the llm space
| Field | Value |
|---|---|
| text | phi models are a prime lesson on how to train on benchmarks to make your model look better than it is for example they claim phi-3-mini is better than llama3 and mixtral (look at our mmlu scores!!!) yet on lmsys arena leaderboard llama3 is top 20 (mixtral just barely outside at 23) while phi-3 mini is... outside of top 50 in the ranks of mistral 7b not to mention from several people that actually used or ran their own evals knows phi models are bad. eg https://twitter.com/abacaj/status/1792991… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1iUnBjbWtDaTZaNGoyT0NwbkpzRk14cFhxbzVBYjNuUS1XNE9ZS2NvVXBxYUthMmxLQ0piR2RLQ1RIVGF3TWZVN2Qza0x4cFRwOFVlUlhmZ1pYbE12b3c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9zSGdzRU9xaTJqTEQwY2VjQUhuZDRnNlRVSmRtTk1lZU5MWGpWR01jd3NqVW94cU5nOXluYUdrRjlOYmtqbGtiLUlKRERzRElVWERrcHVwOWEwc1hSYmU4S1Z6VllIMERkR0RRd0hXUzNFc2gtWUpOWWdiOTNqR016N0MtUDh4SkUtSWNESE9SRklfUFFRZkZoVS1wQm1IMjRnTkRZcG1WZG1Fb3c5VXY5V0dfc25hclNXZGYzb2NhZnQ3OUZRQzRhbWd4NzZsV3owUnlHdl84enZGekp6QT09 |
Raw Record
{
"text": "phi models are a prime lesson on how to train on benchmarks to make your model look better than it is\n\nfor example they claim phi-3-mini is better than llama3 and mixtral (look at our mmlu scores!!!) yet on lmsys arena leaderboard llama3 is top 20 (mixtral just barely outside at 23) while phi-3 mini is... outside of top 50 in the ranks of mistral 7b\n\nnot to mention from several people that actually used or ran their own evals knows phi models are bad. eg https://twitter.com/abacaj/status/1792991309751284123\n\nmain lessons to be learned:\n\n1. be very skeptical when someone claims to break the compute/perf frontier (see [here](https://pbs.twimg.com/media/GL0RUOAbEAAnq_n?format=jpg&name=4096x4096)). the best estimate of a model's performance is flops.\n2. high quality synthetic data is not going to be the silver bullet that makes your model magically x times more efficient w/ flops. data is not compute agnostic\n3. run your own evals and actually use the model, trusting benchmarks blindly is dumb\n\nMSR has really fallen off from the days when they actually released good models like deberta, lets see if hiring most of the inflection team will get them back on track in the llm space",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1iUnBjbWtDaTZaNGoyT0NwbkpzRk14cFhxbzVBYjNuUS1XNE9ZS2NvVXBxYUthMmxLQ0piR2RLQ1RIVGF3TWZVN2Qza0x4cFRwOFVlUlhmZ1pYbE12b3c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9zSGdzRU9xaTJqTEQwY2VjQUhuZDRnNlRVSmRtTk1lZU5MWGpWR01jd3NqVW94cU5nOXluYUdrRjlOYmtqbGtiLUlKRERzRElVWERrcHVwOWEwc1hSYmU4S1Z6VllIMERkR0RRd0hXUzNFc2gtWUpOWWdiOTNqR016N0MtUDh4SkUtSWNESE9SRklfUFFRZkZoVS1wQm1IMjRnTkRZcG1WZG1Fb3c5VXY5V0dfc25hclNXZGYzb2NhZnQ3OUZRQzRhbWd4NzZsV3owUnlHdl84enZGekp6QT09"
}
Entry Information
- Entry ID: 65148
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000