Row 78076

Row ID: 78076 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 78076 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

At our organization, the self hosted LLM service deployment was the exact same as any other service, and we had hosted mistral 7b on g5.xlarge instances, so day or parallelism was also not required. We were able to hit around 100 concurrent requests, which was more than enough for our needs, so scaling was also a non issue.

By far the biggest headache was setting up nvidia drivers, getting a base docker image with them installed was easy enough, but it took us a lot of time to get those installed on our node group.

Service startup took around 10 mins as I didn't bundle the model files with the docker image, I had pushed the model files to an S3 bucket and was downloading those during runtime.

FieldValue
text At our organization, the self hosted LLM service deployment was the exact same as any other service, and we had hosted mistral 7b on g5.xlarge instances, so day or parallelism was also not required. We were able to hit around 100 concurrent requests, which was more than enough for our needs, so scaling was also a non issue. By far the biggest headache was setting up nvidia drivers, getting a base docker image with them installed was easy enough, but it took us a lot of time to get those instal…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1qWFR0T1ZPOVBGcUxiekRKaXRpaE9IS0preWZWcXZDeVU5OG5JS3hGaWlIU2p4R2RTWmdMRXpSQ2JtRWJwTWkwa3NuRU5qNjRIei14aGprWVhJazQwWFE9PQ==
url_encoded Z0FBQUFBQm5Lak8wUHlLd0szYkNYckkxbVphV0ZyOFk1QVRfTnU2NjB4ZjZ1WkpkdTZtMzRaZmI3UHRkQVZLczlzam11a1U0ZFdHT01GNnRZcU1uX1VYSkdlcENNUVZ6d0ZiVXcyS0tRWVV3aTNTZ3d2cllwZW03MjRIUl80RzBzMFlERlprQnpnOEc0S2JDRzlxOHFGQ3RWXzBBWWRNT0RrcV9XSENuR051T1JnVzVXWTR1M0o0clcxdWIyMlp5S3hDQzZLXzZxdW1OWktuVFpoLV9pT2dPa3E1VjlTN2tLZz09

Raw Record

{
  "text": "At our organization, the self hosted LLM service deployment was the exact same as any other service, and we had hosted mistral 7b on g5.xlarge instances, so day or parallelism was also not required. We were able to hit around 100 concurrent requests, which was more than enough  for our needs, so scaling was also a non issue.\n\nBy far the biggest headache was setting up nvidia drivers, getting a base docker image with them installed was easy enough, but it took us a lot of time to get those installed on our node group. \n\nService startup took around 10 mins as I didn't bundle the model files with the docker image, I had pushed the model files to an S3 bucket and was downloading those during runtime.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1qWFR0T1ZPOVBGcUxiekRKaXRpaE9IS0preWZWcXZDeVU5OG5JS3hGaWlIU2p4R2RTWmdMRXpSQ2JtRWJwTWkwa3NuRU5qNjRIei14aGprWVhJazQwWFE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak8wUHlLd0szYkNYckkxbVphV0ZyOFk1QVRfTnU2NjB4ZjZ1WkpkdTZtMzRaZmI3UHRkQVZLczlzam11a1U0ZFdHT01GNnRZcU1uX1VYSkdlcENNUVZ6d0ZiVXcyS0tRWVV3aTNTZ3d2cllwZW03MjRIUl80RzBzMFlERlprQnpnOEc0S2JDRzlxOHFGQ3RWXzBBWWRNT0RrcV9XSENuR051T1JnVzVXWTR1M0o0clcxdWIyMlp5S3hDQzZLXzZxdW1OWktuVFpoLV9pT2dPa3E1VjlTN2tLZz09"
}

Entry Information