Row 4883
Content Data
This page contains data entry 4883 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019)
Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.
LLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.
I hope it is useful, and if you have questions please don't hesitate to ask!
Julien
| Field | Value |
|---|---|
| text | Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019) Deploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to prov… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-04-23 |
| username_encoded | Z0FBQUFBQm5LakwyUldzYnhfRE5lUmU1bmZLTzQwUUUyNElFSmwyUUp0RWRlQ1pBSm5WTFRXbFJBcjNuN1dDU0Z0RjVtQkRGc04zZnUxaHREU0RwRnZISXBFbEV3S1RtQ0E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9GYU5wZWxMWC1BdllkejI5djdRZHVzRFNGUF9nOWl6bUR6b0pySFZJWXpDNmtyNWlfTDd0UDhYeXAyZlhHQXdUeWU0NTlkR0tGUHpRLWtYdVRUUUd2dGNEdEhpVUJQT1Nqc2NGZjJDMmM0Tmo1ZS1fYnpPMnFwX0N2cVM2R3dIbWxDSDM1UkVjbkR5Z1pmWGxhdWdGSS1iakFvQjVEaHJPcU1oSEV6RzRnZzBtQUtWVk9CSFdMS1FJVHlGV3I2TTZMYTJISzJDYmJQbGsyaHJrYkFxVHhPZz09 |
Raw Record
{
"text": "Many are trying to install and deploy their own LLaMA 3 model, so here is a tutorial I just made showing how to deploy LLaMA 3 on an AWS EC2 instance: [https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html](https://nlpcloud.com/how-to-install-and-deploy-llama-3-into-production.html?utm_source=reddit&utm_campaign=fqwerty13-6816-81ed-a26450242ac140019)\n\nDeploying LLaMA 3 8B is fairly easy but LLaMA 3 70B is another beast. Given the amount of VRAM needed you might want to provision more than one GPU and use a dedicated inference server like vLLM in order to split your model on several GPUs.\n\nLLaMA 3 8B requires around 16GB of disk space and 20GB of VRAM (GPU memory) in FP16. As for LLaMA 3 70B, it requires around 140GB of disk space and 160GB of VRAM in FP16.\n\nI hope it is useful, and if you have questions please don't hesitate to ask!\n\nJulien",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-04-23",
"username_encoded": "Z0FBQUFBQm5LakwyUldzYnhfRE5lUmU1bmZLTzQwUUUyNElFSmwyUUp0RWRlQ1pBSm5WTFRXbFJBcjNuN1dDU0Z0RjVtQkRGc04zZnUxaHREU0RwRnZISXBFbEV3S1RtQ0E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9GYU5wZWxMWC1BdllkejI5djdRZHVzRFNGUF9nOWl6bUR6b0pySFZJWXpDNmtyNWlfTDd0UDhYeXAyZlhHQXdUeWU0NTlkR0tGUHpRLWtYdVRUUUd2dGNEdEhpVUJQT1Nqc2NGZjJDMmM0Tmo1ZS1fYnpPMnFwX0N2cVM2R3dIbWxDSDM1UkVjbkR5Z1pmWGxhdWdGSS1iakFvQjVEaHJPcU1oSEV6RzRnZzBtQUtWVk9CSFdMS1FJVHlGV3I2TTZMYTJISzJDYmJQbGsyaHJrYkFxVHhPZz09"
}
Entry Information
- Entry ID: 4883
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000