Row 4883

Row ID: 4883 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 4883 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019)

Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.

LLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.

I hope it is useful, and if you have questions please don't hesitate to ask!

Julien

FieldValue
text Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019) Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to prov…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-04-23
username_encoded Z0FBQUFBQm5LakwyUldzYnhfRE5lUmU1bmZLTzQwUUUyNElFSmwyUUp0RWRlQ1pBSm5WTFRXbFJBcjNuN1dDU0Z0RjVtQkRGc04zZnUxaHREU0RwRnZISXBFbEV3S1RtQ0E9PQ==
url_encoded Z0FBQUFBQm5Lak9GYU5wZWxMWC1BdllkejI5djdRZHVzRFNGUF9nOWl6bUR6b0pySFZJWXpDNmtyNWlfTDd0UDhYeXAyZlhHQXdUeWU0NTlkR0tGUHpRLWtYdVRUUUd2dGNEdEhpVUJQT1Nqc2NGZjJDMmM0Tmo1ZS1fYnpPMnFwX0N2cVM2R3dIbWxDSDM1UkVjbkR5Z1pmWGxhdWdGSS1iakFvQjVEaHJPcU1oSEV6RzRnZzBtQUtWVk9CSFdMS1FJVHlGV3I2TTZMYTJISzJDYmJQbGsyaHJrYkFxVHhPZz09

Raw Record

{
  "text": "Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019)\n\nDeploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.\n\nLLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.\n\nI hope it is useful, and if you have questions please don't hesitate to ask!\n\nJulien",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-04-23",
  "username_encoded": "Z0FBQUFBQm5LakwyUldzYnhfRE5lUmU1bmZLTzQwUUUyNElFSmwyUUp0RWRlQ1pBSm5WTFRXbFJBcjNuN1dDU0Z0RjVtQkRGc04zZnUxaHREU0RwRnZISXBFbEV3S1RtQ0E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GYU5wZWxMWC1BdllkejI5djdRZHVzRFNGUF9nOWl6bUR6b0pySFZJWXpDNmtyNWlfTDd0UDhYeXAyZlhHQXdUeWU0NTlkR0tGUHpRLWtYdVRUUUd2dGNEdEhpVUJQT1Nqc2NGZjJDMmM0Tmo1ZS1fYnpPMnFwX0N2cVM2R3dIbWxDSDM1UkVjbkR5Z1pmWGxhdWdGSS1iakFvQjVEaHJPcU1oSEV6RzRnZzBtQUtWVk9CSFdMS1FJVHlGV3I2TTZMYTJISzJDYmJQbGsyaHJrYkFxVHhPZz09"
}

Entry Information