Row 80904

Row ID: 80904 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 80904 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Hi everybody, I am working on a project and I need to train a pretty big model on a Google Colab's 12 GB GPU.

I cannot load the entire model on the GPU at once because it's too big, so I managed to only move the part I need in that moment, in order to save space (this is only a part of my model, my real model is much bigger and uses a lot of vram):

class Analyzer(nn.Module): def __init__(self): super().__init__() self.conv = nn.Sequential( nn.Conv2d(in_channels=1, out_channels=8, kernel_size=4, stride=4), # out -> 8 x 1024 x 256 nn.MaxPool2d(kernel_size=4), # output -> 8 x 256 x 64 ) self.lstm = nn.LSTM(input_size=256 * 64 * 8, hidden_size=1500, num_layers=2) def forward(self, x): device = torch.cuda.current_device() print(f'\nCUDA memory (start): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%') x = x.to('cuda:0') self.conv.to('cuda:0') x = self.conv(x) self.conv.to('cpu') print(f'CUDA memory (after conv): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%') x = x.view(x.size(0), -1) self.lstm.to('cuda:0') x, memory = self.lstm(x) self.lstm.to('cpu') print(f'CUDA memory (after lstm): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%') x = x.view(-1) return x

Actually I am not sure if this method really cleans the gpu vram after each network usage or simply creates a new copy of the network on the cpu. Do you know if this is the right way to do it? Anyway, this seems to work, but when I wanted to compute the backpropagation I didn't really know how to move each network on the gpu to calculate the gradients. I tried this way but it doesn't work:

class Analyzer(nn.Module): # previous part of the model def backpropagation(self, loss): self.conv.to('cuda:0') loss.backward(retain_graph=True) self.conv.to('cpu') self.lstm.to('cuda:0') loss.backward(retain_graph=True) self.lstm.to('cpu') self.head.to('cuda:0') loss.backward() self.head.to('cpu') # training loop for input, label in batch_loader: model.train() optimizer.zero_grad() y_hat = model(input) loss = loss_function(y_hat, label) model.backpropagation(loss) optimizer.step()

Do you have any ideas to make it work or improve its training speed? Thank you, any advice is welcome

FieldValue
text Hi everybody, I am working on a project and I need to train a pretty big model on a Google Colab's 12 GB GPU. I cannot load the entire model on the GPU at once because it's too big, so I managed to only move the part I need in that moment, in order to save space (this is only a part of my model, my real model is much bigger and uses a lot of vram): class Analyzer(nn.Module): def __init__(self): super().__init__() self.conv = nn.Sequential( …
label r/pytorch
dataType post
communityName r/pytorch
datetime 2024-05-24
username_encoded Z0FBQUFBQm5Lak1sd29KQzBNOGtDbEhvVUFKUkZMaUl3M01tSGU3QXFiX1E5NmduWXZlVkllcDV6S004bVVCMUJGUUtIdzQ2QmJXbGtZcGQwUFk4UEp4bHJEOXlTdHBiSkRFbGFRQTdVNmxUSVFMOU95NjEwSUU9
url_encoded Z0FBQUFBQm5Lak8yNXZKQ1hjcGw3cXVOeDFCZnZNOXIyemhNVjBwejJmQmE3MWhyY21NamNobUVBSFlOeWJtWGZqMVdZaENvZ2pIMTRxQm1JR0ZseXBfU29wQ0tkWXlob1VienE2ZlBteWlCQWJuazVpVGpBSjNjQUh4Xzhra20teHpueUZGSnZJR01Wd3Bvd2hBaVJCTlFZRXRrNGlQUzlxYVVOLTVSWHhKWEctcUNKQTFBaHBMUUVya2tmWkxDMTBkZmw3SXJJSjhuUkI2ZDFHeDI3bFBrd2xFVlpsQy0xQT09

Raw Record

{
  "text": "Hi everybody, I am working on a project and I need to train a pretty big model on a Google Colab's 12 GB GPU.\n\nI cannot load the entire model on the GPU at once because it's too big, so I managed to only move the part I need in that moment, in order to save space (this is only a part of my model, my real model is much bigger and uses a lot of vram):\n\n    class Analyzer(nn.Module):\n        def __init__(self):\n            super().__init__()\n    \n            self.conv = nn.Sequential(\n                nn.Conv2d(in_channels=1, out_channels=8, kernel_size=4, stride=4),  # out -> 8 x 1024 x 256\n                nn.MaxPool2d(kernel_size=4),  # output -> 8 x 256 x 64\n            )\n    \n            self.lstm = nn.LSTM(input_size=256 * 64 * 8, hidden_size=1500, num_layers=2)\n    \n        def forward(self, x):\n            device = torch.cuda.current_device()\n            print(f'\\nCUDA memory (start): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%')\n    \n            x = x.to('cuda:0')\n            self.conv.to('cuda:0')\n            x = self.conv(x)\n            self.conv.to('cpu')\n            print(f'CUDA memory (after conv): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%')\n    \n            x = x.view(x.size(0), -1)\n    \n            self.lstm.to('cuda:0')\n            x, memory = self.lstm(x)\n            self.lstm.to('cpu')\n            print(f'CUDA memory (after lstm): {torch.cuda.memory_allocated(device) / torch.cuda.get_device_properties(device).total_memory * 100:0.3f}%')\n    \n            x = x.view(-1)\n    \n            return x\n\n  \nActually I am not sure if this method really cleans the gpu vram after each network usage or simply creates a new copy of the network on the cpu. Do you know if this is the right way to do it?  \n  \nAnyway, this seems to work, but when I wanted to compute the backpropagation I didn't really know how to move each network on the gpu to calculate the gradients. I tried this way but it doesn't work:\n\n    class Analyzer(nn.Module):\n        # previous part of the model\n        def backpropagation(self, loss):\n            self.conv.to('cuda:0')\n            loss.backward(retain_graph=True)\n            self.conv.to('cpu')\n    \n            self.lstm.to('cuda:0')\n            loss.backward(retain_graph=True)\n            self.lstm.to('cpu')\n    \n            self.head.to('cuda:0')\n            loss.backward()\n            self.head.to('cpu')\n    \n    # training loop\n    for input, label in batch_loader:\n        model.train()\n    \n        optimizer.zero_grad()\n    \n        y_hat = model(input)\n        loss = loss_function(y_hat, label)\n        \n        model.backpropagation(loss)\n        optimizer.step()\n\n  \nDo you have any ideas to make it work or improve its training speed?  \nThank you, any advice is welcome",
  "label": "r/pytorch",
  "dataType": "post",
  "communityName": "r/pytorch",
  "datetime": "2024-05-24",
  "username_encoded": "Z0FBQUFBQm5Lak1sd29KQzBNOGtDbEhvVUFKUkZMaUl3M01tSGU3QXFiX1E5NmduWXZlVkllcDV6S004bVVCMUJGUUtIdzQ2QmJXbGtZcGQwUFk4UEp4bHJEOXlTdHBiSkRFbGFRQTdVNmxUSVFMOU95NjEwSUU9",
  "url_encoded": "Z0FBQUFBQm5Lak8yNXZKQ1hjcGw3cXVOeDFCZnZNOXIyemhNVjBwejJmQmE3MWhyY21NamNobUVBSFlOeWJtWGZqMVdZaENvZ2pIMTRxQm1JR0ZseXBfU29wQ0tkWXlob1VienE2ZlBteWlCQWJuazVpVGpBSjNjQUh4Xzhra20teHpueUZGSnZJR01Wd3Bvd2hBaVJCTlFZRXRrNGlQUzlxYVVOLTVSWHhKWEctcUNKQTFBaHBMUUVya2tmWkxDMTBkZmw3SXJJSjhuUkI2ZDFHeDI3bFBrd2xFVlpsQy0xQT09"
}

Entry Information