Row 5354

Row ID: 5354 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 5354 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I tired playing around with open tools to serve multiple models on a Single A100 80GB

I was able to squeeze in 5 language models using VLLM, served via fastchat.

​

Here is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work.

[https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/)

​

Llama 3 BB (8K context) Gemma 8B (8k Context) Phi 3 128K (18K context) DeciLM 7B (8K Context) Stable LM 1.6B (4k Context)

​

Is it stupid? - Maybe, lets discuss down!

Here a writeup on how I did it

[https://supersecurehuman.github.io/Serving-FastChat/](https://supersecurehuman.github.io/Serving-FastChat/)

[https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b](https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b?source=user_profile---------0----------------------------)

FieldValue
text I tired playing around with open tools to serve multiple models on a Single A100 80GB I was able to squeeze in 5 language models using VLLM, served via fastchat. ​ Here is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work. [https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/) ​ Llama 3 BB (8K context) Gemma 8B (8k Context) Phi 3 128K (18K context) DeciLM 7B (8K Context) Stable LM 1.6B (4k Context) &#x…
label r/deeplearning
dataType post
communityName r/deeplearning
datetime 2024-04-29
username_encoded Z0FBQUFBQm5LakwyUEZRc1N1XzlqcC1CUTZiZTJsbVA0VERIWGFGMm1JemRYQkhhYnJWcUJENkZqend6VTlNaENNbzRfazFZQXpDajhZM0p4RHpLb2xENnFweGxhRWRvOXhpa3M5WGtkMU1YeFhDUXJaY2g2V2s9
url_encoded Z0FBQUFBQm5Lak9GOHBnLTJnMURnbS0tLUlkTXBKeGtUSHN3cW5TNnlNemNpMF9DbXc2MmRXQ2dpbGpjbFJQYmtjV0thRzZnV1ZHZDFnaFhpckhUZVloQmNBYnRzX0s0LWhqQTNacjIzekl4Sk5vRG95TG43dWhVX1hEVlB6bm5mdDBZMG5zVDVMOEhQVHZvSTQzQ2ZKQTZGY0h4NXBrNFF5d3o0M1dhNGV1ZVlOWEFCR1ZCV1lSdWVsdnRWNU5xUjFHRkVrdEpjR2VzNFFFMm9HYXdtclNLZ2JJSWVkUUxkQT09

Raw Record

{
  "text": "I tired playing around with open tools to serve multiple models on a Single A100 80GB\n\nI was able to squeeze in 5 language models using VLLM, served via fastchat.\n\n​\n\nHere is a public instance for 72 hrs - Pls note that the battle mode is broken, other tabs work.\n\n[https://c8168701070daa5bf3.gradio.live/](https://c8168701070daa5bf3.gradio.live/)\n\n​\n\nLlama 3 BB (8K context)  \nGemma 8B (8k Context)  \nPhi 3 128K (18K context)  \nDeciLM 7B (8K Context)  \nStable LM 1.6B (4k Context)\n\n​\n\nIs it stupid? - Maybe, lets discuss down!\n\nHere a writeup on how I did it\n\n[https://supersecurehuman.github.io/Serving-FastChat/](https://supersecurehuman.github.io/Serving-FastChat/)\n\n[https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b](https://supersecurehuman.medium.com/serving-fastchat-personal-journey-fe84b26ef87b?source=user_profile---------0----------------------------)",
  "label": "r/deeplearning",
  "dataType": "post",
  "communityName": "r/deeplearning",
  "datetime": "2024-04-29",
  "username_encoded": "Z0FBQUFBQm5LakwyUEZRc1N1XzlqcC1CUTZiZTJsbVA0VERIWGFGMm1JemRYQkhhYnJWcUJENkZqend6VTlNaENNbzRfazFZQXpDajhZM0p4RHpLb2xENnFweGxhRWRvOXhpa3M5WGtkMU1YeFhDUXJaY2g2V2s9",
  "url_encoded": "Z0FBQUFBQm5Lak9GOHBnLTJnMURnbS0tLUlkTXBKeGtUSHN3cW5TNnlNemNpMF9DbXc2MmRXQ2dpbGpjbFJQYmtjV0thRzZnV1ZHZDFnaFhpckhUZVloQmNBYnRzX0s0LWhqQTNacjIzekl4Sk5vRG95TG43dWhVX1hEVlB6bm5mdDBZMG5zVDVMOEhQVHZvSTQzQ2ZKQTZGY0h4NXBrNFF5d3o0M1dhNGV1ZVlOWEFCR1ZCV1lSdWVsdnRWNU5xUjFHRkVrdEpjR2VzNFFFMm9HYXdtclNLZ2JJSWVkUUxkQT09"
}

Entry Information