Row 5900
Content Data
This page contains data entry 5900 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
[https://tensordock.com/benchmarks](https://tensordock.com/benchmarks)
Spent the past few hours putting together some data on vLLM (for both Llama 7B and OPT-125M) and Resnet-50 training performance on the TensorDock cloud.
vLLM data is 100% out of the box, with 2048 batch sizes from [this repository](https://github.com/vllm-project/vllm/tree/main/benchmarks).
My learnings:
* H100 and A100 performance is unbeatable, but the price-to-performance of lower-end RTX cards is pretty darn good. Even the L40 and RTX 6000 Ada outperform the A100 at some tasks, as they are 1 generation newer than the A100. **If your application does not need 80GB of VRAM, it probably makes sense to not use an 80GB VRAM card** * Standalone H100 performance isn't as strong as I would have imagined. H100 performance is bottlenecked by memory bandwidth for LLM inference, hence **H100s are only 1.8x faster than A100s for vLLM**. H100s really perform better when interconnected together, but I didn't benchmark that today. * **CPU matters more than I expected**. The OPT-125M vs Llama 7B performance comparison is pretty interesting... somehow all GPUs tend to perform similar on OPT-125M, and I assume that's because relatively more CPU time is used than GPU time, so the GPU performance difference matters less in the grand scheme of things. * The marketplace prices itself pretty well. If cohort GPUs into VRAM amount all GPUs with similar amounts of VRAM share a similar price-to-performance ratio. * Self hosting can save you $$$ if you have sufficient batch sizes. If you built your own inference API, you could serve LLMs utilizing just 50% batches and save money compared to a pay-per-token API (we \[TensorDock\] cost less than $0.07 per million Llama 7 tokens if you use us at 100%)
\--
Let me know which GPU to benchmark next, and I'll add that! Or let me know some other workload to measure, and I'd be happy to add an new section for that too.
P.S. We added some H100s at $1.80/hr for anyone lucky enough to grab them!
| Field | Value |
|---|---|
| text | [https://tensordock.com/benchmarks](https://tensordock.com/benchmarks) Spent the past few hours putting together some data on vLLM (for both Llama 7B and OPT-125M) and Resnet-50 training performance on the TensorDock cloud. vLLM data is 100% out of the box, with 2048 batch sizes from [this repository](https://github.com/vllm-project/vllm/tree/main/benchmarks). My learnings: * H100 and A100 performance is unbeatable, but the price-to-performance of lower-end RTX cards is pretty darn good. E… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-05 |
| username_encoded | Z0FBQUFBQm5LakwySHpiM0NpeFk0aTk3RHlzWUs2NFRaRDFjNWZxWmFzMzljTVR4MGVHb3dnWFprcmZCemlFUjdnTFV4ckNCcUJnTGkzQ1Q4dFN4N096bnNrQmtjN0VMTkE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HbFQ3ODdGUFU1R3NId0RUWldKTGtKa3lHcnBiRE5LRkV3Z1RBbDlFNzJnR0pFNXhmREQ2a0o5OEMyMGc5ZDdJdGJNRWFqc3c3RENOT2tJVF9TRWFoSUxVQkNxMHlGNGZtRDYxclhuRTUxSkhqX1BHSjQzNFZHOHhfVWZ0SXNmNjdVOVFCVW84S2M4dThaa2p4ZjJsa0xtdDhQTk5pc0c4Q2E1QTdmaWxTSXZPZTV0SERZZUxpbDMzQ3ItUFZUMEJu |
Raw Record
{
"text": "[https://tensordock.com/benchmarks](https://tensordock.com/benchmarks)\n\nSpent the past few hours putting together some data on vLLM (for both Llama 7B and OPT-125M) and Resnet-50 training performance on the TensorDock cloud. \n\nvLLM data is 100% out of the box, with 2048 batch sizes from [this repository](https://github.com/vllm-project/vllm/tree/main/benchmarks). \n\nMy learnings:\n\n* H100 and A100 performance is unbeatable, but the price-to-performance of lower-end RTX cards is pretty darn good. Even the L40 and RTX 6000 Ada outperform the A100 at some tasks, as they are 1 generation newer than the A100. **If your application does not need 80GB of VRAM, it probably makes sense to not use an 80GB VRAM card**\n* Standalone H100 performance isn't as strong as I would have imagined. H100 performance is bottlenecked by memory bandwidth for LLM inference, hence **H100s are only 1.8x faster than A100s for vLLM**. H100s really perform better when interconnected together, but I didn't benchmark that today. \n* **CPU matters more than I expected**. The OPT-125M vs Llama 7B performance comparison is pretty interesting... somehow all GPUs tend to perform similar on OPT-125M, and I assume that's because relatively more CPU time is used than GPU time, so the GPU performance difference matters less in the grand scheme of things. \n* The marketplace prices itself pretty well. If cohort GPUs into VRAM amount all GPUs with similar amounts of VRAM share a similar price-to-performance ratio.\n* Self hosting can save you $$$ if you have sufficient batch sizes. If you built your own inference API, you could serve LLMs utilizing just 50% batches and save money compared to a pay-per-token API (we \\[TensorDock\\] cost less than $0.07 per million Llama 7 tokens if you use us at 100%)\n\n\\-- \n\nLet me know which GPU to benchmark next, and I'll add that! Or let me know some other workload to measure, and I'd be happy to add an new section for that too. \n\nP.S. We added some H100s at $1.80/hr for anyone lucky enough to grab them! ",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-05",
"username_encoded": "Z0FBQUFBQm5LakwySHpiM0NpeFk0aTk3RHlzWUs2NFRaRDFjNWZxWmFzMzljTVR4MGVHb3dnWFprcmZCemlFUjdnTFV4ckNCcUJnTGkzQ1Q4dFN4N096bnNrQmtjN0VMTkE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HbFQ3ODdGUFU1R3NId0RUWldKTGtKa3lHcnBiRE5LRkV3Z1RBbDlFNzJnR0pFNXhmREQ2a0o5OEMyMGc5ZDdJdGJNRWFqc3c3RENOT2tJVF9TRWFoSUxVQkNxMHlGNGZtRDYxclhuRTUxSkhqX1BHSjQzNFZHOHhfVWZ0SXNmNjdVOVFCVW84S2M4dThaa2p4ZjJsa0xtdDhQTk5pc0c4Q2E1QTdmaWxTSXZPZTV0SERZZUxpbDMzQ3ItUFZUMEJu"
}
Entry Information
- Entry ID: 5900
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000