Row 28811
Content Data
This page contains data entry 28811 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more.
[Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7)
Utilizing the models, we further demonstrate:
* Inference performance speedups from sparsity alone at 3x for CPUs and 1.7x for GPUs on Neural Magic's platform. * Compounded gains with quantization for up to 8.6X faster inference performance. * Close to theoretical gains for sparse training utilizing Cerebras's CS-3 AI accelerator.
[Prefill and Decode Llama 2 7B Performance at Various Sparsity Levels for FP32 and INT8 on an 8-core CPU.](https://preview.redd.it/jtxsxxio8s1d1.png?width=3636&format=png&auto=webp&s=ce9d228d9ff5891e043b83c27d955fb0645f3ecd)
Paper: [https://arxiv.org/abs/2405.03594](https://arxiv.org/abs/2405.03594) Models: [https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21](https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21)
| Field | Value |
|---|---|
| text | In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more. [Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7) Utilizing the models, we furt… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-21 |
| username_encoded | Z0FBQUFBQm5Lak1FdFF3RXMyOVlFNkNrcGJFTTcwV0tBVlZiRElLOEV2SVd4SVdCTGVsdlZhbTdwWFY2ajh0b2lKOS04SmlMRmJLamNmbUtDLTJWUDhIdkZMWGFQYkQ5VlE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9VcU9CMW5yNERhS1pqR0M4dUR3MjVndUN0eGFSSTN6a3R3TXlsbl93RlZiRmNNZHdmbU5rNkoxcW9nQXJEYlNhSmdzY0lZQy1OMjJGT2pXNXNPY09JaXJ4S1FqZE1Ic25BUnRqZUlpc2hiV2ZGbWVJQVJIbzNVcjZKR2dBb0EtV3RCQU02akxieFZkQ2U0MWpyWng1WlhWcUszX3hybHhtMFkzYnNtX2VVT1VoVDdac3luS0Q5Sm55Mll0V2Z5cVFtanVaUVh4TV9zWmdGaFhDakFHSUxRdz09 |
Raw Record
{
"text": "In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more.\n\n[Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7)\n\nUtilizing the models, we further demonstrate:\n\n* Inference performance speedups from sparsity alone at 3x for CPUs and 1.7x for GPUs on Neural Magic's platform.\n* Compounded gains with quantization for up to 8.6X faster inference performance.\n* Close to theoretical gains for sparse training utilizing Cerebras's CS-3 AI accelerator.\n\n[Prefill and Decode Llama 2 7B Performance at Various Sparsity Levels for FP32 and INT8 on an 8-core CPU.](https://preview.redd.it/jtxsxxio8s1d1.png?width=3636&format=png&auto=webp&s=ce9d228d9ff5891e043b83c27d955fb0645f3ecd)\n\n \nPaper: [https://arxiv.org/abs/2405.03594](https://arxiv.org/abs/2405.03594) \nModels: [https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21](https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21)",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-21",
"username_encoded": "Z0FBQUFBQm5Lak1FdFF3RXMyOVlFNkNrcGJFTTcwV0tBVlZiRElLOEV2SVd4SVdCTGVsdlZhbTdwWFY2ajh0b2lKOS04SmlMRmJLamNmbUtDLTJWUDhIdkZMWGFQYkQ5VlE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9VcU9CMW5yNERhS1pqR0M4dUR3MjVndUN0eGFSSTN6a3R3TXlsbl93RlZiRmNNZHdmbU5rNkoxcW9nQXJEYlNhSmdzY0lZQy1OMjJGT2pXNXNPY09JaXJ4S1FqZE1Ic25BUnRqZUlpc2hiV2ZGbWVJQVJIbzNVcjZKR2dBb0EtV3RCQU02akxieFZkQ2U0MWpyWng1WlhWcUszX3hybHhtMFkzYnNtX2VVT1VoVDdac3luS0Q5Sm55Mll0V2Z5cVFtanVaUVh4TV9zWmdGaFhDakFHSUxRdz09"
}
Entry Information
- Entry ID: 28811
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000