Row 3980

Row ID: 3980 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 3980 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I have $10,000 to spend on an optimal setup to use large deep learning models and image datasets.

We are currently using two RTX Titan on a linux server but one complete run of my experiments takes around 3-5 days (this is typical for some projects but I am looking for intraday experiment runs). Data size is around 5GB. However, in future projects, data size will increase to around 10 TB. Models used are your typical EfficientNetB1, ResNet50, VGG16, etc. However, I would like to experiment with the larger models as well like EfficientNetB7. Further, the system overheats sometimes.

I understand that first and foremost, optimizing my code should be a priority. Which is better: parallelizing my model or data or both?

As for GPU setup, is it better to buy say 5 RTX 4090 GPUs (have 1 GPU available for other PhD students to use and 4 to run my projects on)? What about TPUs or cloud computing power? Since cloud services pay by the hour, it may not be optimal in the long run as an investment to our group.

Also, I read somewhere that PyTorch has some problems in running models in parallel with RTX 4090. Is that still the case? Would RTX 3090 be better? I understand the VRAM is an issue for large data with this setup, so would A100 or other products be better? As of right now, DataLoader is taking the most time, and I expect that bottleneck to increase with the larger future datasets.

I am extremely new to this so any help would be appreciated.

FieldValue
text I have $10,000 to spend on an optimal setup to use large deep learning models and image datasets. We are currently using two RTX Titan on a linux server but one complete run of my experiments takes around 3-5 days (this is typical for some projects but I am looking for intraday experiment runs). Data size is around 5GB. However, in future projects, data size will increase to around 10 TB. Models used are your typical EfficientNetB1, ResNet50, VGG16, etc. However, I would like to experiment with…
label r/pytorch
dataType post
communityName r/pytorch
datetime 2024-03-31
username_encoded Z0FBQUFBQm5LakwxUWV0Mk9BbUFsUXhTZmZwTWc4ZFktTEdmVV9Uanp2TGxUOHBtMDhrWFp6ckpNM0p6Mm5PMGI0VG5lTTJNb2psZzVWM1lJUHVIbGJ2SWpDVjNHazJzTHc9PQ==
url_encoded Z0FBQUFBQm5Lak9GNTdpU0dEeHJuYTdxU19NNEtpNEdMQi1Cajg5d2NEQUQwNGZDWTJQM2kxS183VGFYb1YxZjMycmM0UFJJNlowamd3d0pZeWFEYjJoRDA4N2ZtOWphOTNhZi1iaVNGQW85TnhuMUR6ODdrcS1GQUZBeHlYV3RqY2xSLWhsbW5JSW42U2o4ZmR5dFdBN3R4QmtCR09lRnc1RUVNUTlrdi1COFJsVC1OcnB3THBySzBRYzJXLThfbXhYNkdsNnJySDdzclF1WkFWODNydDJpSUJKemFWNEZxdz09

Raw Record

{
  "text": "I have $10,000 to spend on an optimal setup to use large deep learning models and image datasets.\n\nWe are currently using two RTX Titan on a linux server but one complete run of my experiments takes around 3-5 days (this is typical for some projects but I am looking for intraday experiment runs). Data size is around 5GB. However, in future projects, data size will increase to around 10 TB. Models used are your typical EfficientNetB1, ResNet50, VGG16, etc. However, I would like to experiment with the larger models as well like EfficientNetB7. Further, the system overheats sometimes.\n\nI understand that first and foremost, optimizing my code should be a priority. Which is better: parallelizing my model or data or both?\n\nAs for GPU setup, is it better to buy say 5 RTX 4090 GPUs (have 1 GPU available for other PhD students to use and 4 to run my projects on)? What about TPUs or cloud computing power? Since cloud services pay by the hour, it may not be optimal in the long run as an investment to our group.\n\nAlso, I read somewhere that PyTorch has some problems in running models in parallel with RTX 4090. Is that still the case? Would RTX 3090 be better? I understand the VRAM is an issue for large data with this setup, so would A100 or other products be better? As of right now, DataLoader is taking the most time, and I expect that bottleneck to increase with the larger future datasets.\n\nI am extremely new to this so any help would be appreciated.",
  "label": "r/pytorch",
  "dataType": "post",
  "communityName": "r/pytorch",
  "datetime": "2024-03-31",
  "username_encoded": "Z0FBQUFBQm5LakwxUWV0Mk9BbUFsUXhTZmZwTWc4ZFktTEdmVV9Uanp2TGxUOHBtMDhrWFp6ckpNM0p6Mm5PMGI0VG5lTTJNb2psZzVWM1lJUHVIbGJ2SWpDVjNHazJzTHc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GNTdpU0dEeHJuYTdxU19NNEtpNEdMQi1Cajg5d2NEQUQwNGZDWTJQM2kxS183VGFYb1YxZjMycmM0UFJJNlowamd3d0pZeWFEYjJoRDA4N2ZtOWphOTNhZi1iaVNGQW85TnhuMUR6ODdrcS1GQUZBeHlYV3RqY2xSLWhsbW5JSW42U2o4ZmR5dFdBN3R4QmtCR09lRnc1RUVNUTlrdi1COFJsVC1OcnB3THBySzBRYzJXLThfbXhYNkdsNnJySDdzclF1WkFWODNydDJpSUJKemFWNEZxdz09"
}

Entry Information