Row 28811

Row ID: 28811 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 28811 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more.

[Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7)

Utilizing the models, we further demonstrate:

* Inference performance speedups from sparsity alone at 3x for CPUs and 1.7x for GPUs on Neural Magic's platform. * Compounded gains with quantization for up to 8.6X faster inference performance. * Close to theoretical gains for sparse training utilizing Cerebras's CS-3 AI accelerator.

[Prefill and Decode Llama 2 7B Performance at Various Sparsity Levels for FP32 and INT8 on an 8-core CPU.](https://preview.redd.it/jtxsxxio8s1d1.png?width=3636&format=png&auto=webp&s=ce9d228d9ff5891e043b83c27d955fb0645f3ecd)

Paper: [https://arxiv.org/abs/2405.03594](https://arxiv.org/abs/2405.03594) Models: [https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21](https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21)

FieldValue
text In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more. [Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7) Utilizing the models, we furt…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1FdFF3RXMyOVlFNkNrcGJFTTcwV0tBVlZiRElLOEV2SVd4SVdCTGVsdlZhbTdwWFY2ajh0b2lKOS04SmlMRmJLamNmbUtDLTJWUDhIdkZMWGFQYkQ5VlE9PQ==
url_encoded Z0FBQUFBQm5Lak9VcU9CMW5yNERhS1pqR0M4dUR3MjVndUN0eGFSSTN6a3R3TXlsbl93RlZiRmNNZHdmbU5rNkoxcW9nQXJEYlNhSmdzY0lZQy1OMjJGT2pXNXNPY09JaXJ4S1FqZE1Ic25BUnRqZUlpc2hiV2ZGbWVJQVJIbzNVcjZKR2dBb0EtV3RCQU02akxieFZkQ2U0MWpyWng1WlhWcUszX3hybHhtMFkzYnNtX2VVT1VoVDdac3luS0Q5Sm55Mll0V2Z5cVFtanVaUVh4TV9zWmdGaFhDakFHSUxRdz09

Raw Record

{
  "text": "In a collaboration across Neural Magic, Cerebras, and IST Austria, we've pushed out, to the best of our knowledge, the first highly sparse, foundational LLMs with full recovery on several fine-tuning tasks, including chat, code generation, summarization, and more.\n\n[Sparsity vs Baseline Accuracy Recovery for Popular Fine-tuning Tasks for Llama 2 7B](https://preview.redd.it/mexhej5i8s1d1.png?width=2936&format=png&auto=webp&s=3a126fe3ca7a36f25cb71cfb02924d2b26aa72f7)\n\nUtilizing the models, we further demonstrate:\n\n* Inference performance speedups from sparsity alone at 3x for CPUs and 1.7x for GPUs on Neural Magic's platform.\n* Compounded gains with quantization for up to 8.6X faster inference performance.\n* Close to theoretical gains for sparse training utilizing Cerebras's CS-3 AI accelerator.\n\n[Prefill and Decode Llama 2 7B Performance at Various Sparsity Levels for FP32 and INT8 on an 8-core CPU.](https://preview.redd.it/jtxsxxio8s1d1.png?width=3636&format=png&auto=webp&s=ce9d228d9ff5891e043b83c27d955fb0645f3ecd)\n\n  \nPaper: [https://arxiv.org/abs/2405.03594](https://arxiv.org/abs/2405.03594)  \nModels: [https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21](https://huggingface.co/collections/neuralmagic/sparse-foundational-llama-2-models-65f48cec6396309f02e74d21)",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1FdFF3RXMyOVlFNkNrcGJFTTcwV0tBVlZiRElLOEV2SVd4SVdCTGVsdlZhbTdwWFY2ajh0b2lKOS04SmlMRmJLamNmbUtDLTJWUDhIdkZMWGFQYkQ5VlE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9VcU9CMW5yNERhS1pqR0M4dUR3MjVndUN0eGFSSTN6a3R3TXlsbl93RlZiRmNNZHdmbU5rNkoxcW9nQXJEYlNhSmdzY0lZQy1OMjJGT2pXNXNPY09JaXJ4S1FqZE1Ic25BUnRqZUlpc2hiV2ZGbWVJQVJIbzNVcjZKR2dBb0EtV3RCQU02akxieFZkQ2U0MWpyWng1WlhWcUszX3hybHhtMFkzYnNtX2VVT1VoVDdac3luS0Q5Sm55Mll0V2Z5cVFtanVaUVh4TV9zWmdGaFhDakFHSUxRdz09"
}

Entry Information