Row 5354
Content Data
This page contains data entry 5354 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I tired playing around with open tools to serve multiple models on a Single A100 80GB
I was able to squeeze in 5 language models using VLLM, served via fastchat.
​
Here is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work.
[https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/)
​
Llama 3 BB (8K context) Gemma 8B (8k Context) Phi 3 128K (18K context) DeciLM 7B (8K Context) Stable LM 1.6B (4k Context)
​
Is it stupid? - Maybe, lets discuss down!
Here a writeup on how I did it
[https://supersecurehuman.github.io/Serving-FastChat/](https://supersecurehuman.github.io/Serving-FastChat/)
[https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b](https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b?source=user_profile---------0----------------------------)
| Field | Value |
|---|---|
| text | I tired playing around with open tools to serve multiple models on a Single A100 80GB I was able to squeeze in 5 language models using VLLM, served via fastchat. ​ Here is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work. [https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/) ​ Llama 3 BB (8K context) Gemma 8B (8k Context) Phi 3 128K (18K context) DeciLM 7B (8K Context) Stable LM 1.6B (4k Context) &#x… |
| label | r/deeplearning |
| dataType | post |
| communityName | r/deeplearning |
| datetime | 2024-04-29 |
| username_encoded | Z0FBQUFBQm5LakwyUEZRc1N1XzlqcC1CUTZiZTJsbVA0VERIWGFGMm1JemRYQkhhYnJWcUJENkZqend6VTlNaENNbzRfazFZQXpDajhZM0p4RHpLb2xENnFweGxhRWRvOXhpa3M5WGtkMU1YeFhDUXJaY2g2V2s9 |
| url_encoded | Z0FBQUFBQm5Lak9GOHBnLTJnMURnbS0tLUlkTXBKeGtUSHN3cW5TNnlNemNpMF9DbXc2MmRXQ2dpbGpjbFJQYmtjV0thRzZnV1ZHZDFnaFhpckhUZVloQmNBYnRzX0s0LWhqQTNacjIzekl4Sk5vRG95TG43dWhVX1hEVlB6bm5mdDBZMG5zVDVMOEhQVHZvSTQzQ2ZKQTZGY0h4NXBrNFF5d3o0M1dhNGV1ZVlOWEFCR1ZCV1lSdWVsdnRWNU5xUjFHRkVrdEpjR2VzNFFFMm9HYXdtclNLZ2JJSWVkUUxkQT09 |
Raw Record
{
"text": "I tired playing around with open tools to serve multiple models on a Single A100 80GB\n\nI was able to squeeze in 5 language models using VLLM, served via fastchat.\n\n​\n\nHere is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work.\n\n[https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/)\n\n​\n\nLlama 3 BB (8K context) \nGemma 8B (8k Context) \nPhi 3 128K (18K context) \nDeciLM 7B (8K Context) \nStable LM 1.6B (4k Context)\n\n​\n\nIs it stupid? - Maybe, lets discuss down!\n\nHere a writeup on how I did it\n\n[https://supersecurehuman.github.io/Serving-FastChat/](https://supersecurehuman.github.io/Serving-FastChat/)\n\n[https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b](https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b?source=user_profile---------0----------------------------)",
"label": "r/deeplearning",
"dataType": "post",
"communityName": "r/deeplearning",
"datetime": "2024-04-29",
"username_encoded": "Z0FBQUFBQm5LakwyUEZRc1N1XzlqcC1CUTZiZTJsbVA0VERIWGFGMm1JemRYQkhhYnJWcUJENkZqend6VTlNaENNbzRfazFZQXpDajhZM0p4RHpLb2xENnFweGxhRWRvOXhpa3M5WGtkMU1YeFhDUXJaY2g2V2s9",
"url_encoded": "Z0FBQUFBQm5Lak9GOHBnLTJnMURnbS0tLUlkTXBKeGtUSHN3cW5TNnlNemNpMF9DbXc2MmRXQ2dpbGpjbFJQYmtjV0thRzZnV1ZHZDFnaFhpckhUZVloQmNBYnRzX0s0LWhqQTNacjIzekl4Sk5vRG95TG43dWhVX1hEVlB6bm5mdDBZMG5zVDVMOEhQVHZvSTQzQ2ZKQTZGY0h4NXBrNFF5d3o0M1dhNGV1ZVlOWEFCR1ZCV1lSdWVsdnRWNU5xUjFHRkVrdEpjR2VzNFFFMm9HYXdtclNLZ2JJSWVkUUxkQT09"
}
Entry Information
- Entry ID: 5354
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000