Row 53645

Row ID: 53645 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 53645 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

Hi Everyone,

for learning purposes I've been replicating the methods from Kaiming He's 2015 ['Deep Residual Learning for Image Recognition'](https://arxiv.org/pdf/1512.03385). I've built the VGG inspired plain-CNN as well as the ResNet architectures (standard & bottleneck).

However, I have been unable to replicate the degradation (saturation of accuracy) problem highlighted in the publication. The %-error figures in the publication show clear drops in %-error as training progresses followed by stagnation.

My figures appear to stagnate, but its clear the model is generalizing horribly to the validation data. I've included one of their figures as reference. Any recommendations to better replicate the error rate saturation from this paper? Note: For Kaiming He figure, bold lines are testing error & dashed are training.

Parameters:

* 162 Epochs w/ batch size of 128 for 64k iterations. * Lr: 0.1 * Momentum: 0.9 * Weight decay: 0.0001 * Multi-step scheduler dividing the lr by 10 at 32k and 48k iterations

[My %-error](https://preview.redd.it/du9j19ay322d1.png?width=729&format=png&auto=webp&s=15d461e662e5ca823bef852b73626c5cb3f043c8)

[From Paper](https://preview.redd.it/zsvtfxb0422d1.png?width=701&format=png&auto=webp&s=019edcbf04a5eff3ffb0902fbb3057fd501e9dd4)

**Edit:** Adjustment to training transformation by implementing normalization prior to random cropping & re-ran all models. First image is change to plain 18 error rate curve, second is all error rates for tested architectures.

https://preview.redd.it/g7gpcuql482d1.png?width=1000&format=png&auto=webp&s=0db10a04a16585a0401c56a7f4c2ae25f9f8249b

https://preview.redd.it/7d35ztql482d1.png?width=1000&format=png&auto=webp&s=a83604401d1971bc7472d0b08231985a222ba8f4

FieldValue
text Hi Everyone, for learning purposes I've been replicating the methods from Kaiming He's 2015 ['Deep Residual Learning for Image Recognition'](https://arxiv.org/pdf/1512.03385). I've built the VGG inspired plain-CNN as well as the ResNet architectures (standard & bottleneck). However, I have been unable to replicate the degradation (saturation of accuracy) problem highlighted in the publication. The %-error figures in the publication show clear drops in %-error as training progresses followed b…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-22
username_encoded Z0FBQUFBQm5Lak1VdzVORW9hNjZMLV85QjlRWWZRNlQxa292bFQyU0xfdzVrdks2SER5Vk9YMVdGR3JHUnAyQzJZRWlPRjdDd3JOck9XQklOM05yRXJQY0hZYzVfRC1mNVU3OVlSYWdWWnB4VDBBcTZzWEwwMDA9
url_encoded Z0FBQUFBQm5Lak9rcHNzaVNPQVJjWmVGaU5DQ040OU43LU1ZVTkwSFBnclM2RlhlcmZSNlk3blc5NUR2cTUwbjlBRkU4eFE4M3NyRDJhYkRPYUtBVFdtTzBERXBEMnFNQi1fd3lZbTZsUjh5dzJ0bHVUUTctSDFsX2xuRWFqdFNjNFg5RFlCblg4TUhxT0tCbThyYzlOblVCVkFrNzA0NU5CRnV5Ym5QZklpNFN2bktYMDJVa043SHpuaGZ0YlF2MTRBaEhsazVmajM1VUxrUUF0T0tILUdrMzVuVmIySFRxZz09

Raw Record

{
  "text": "Hi Everyone,\n\nfor learning purposes I've been replicating the methods from Kaiming He's 2015 ['Deep Residual Learning for Image Recognition'](https://arxiv.org/pdf/1512.03385). I've built the VGG inspired plain-CNN as well as the ResNet architectures (standard & bottleneck).\n\nHowever, I have been unable to replicate the degradation (saturation of accuracy) problem highlighted in the publication.  The %-error figures in the publication show clear drops in %-error as training progresses followed by stagnation.\n\nMy figures appear to stagnate, but its clear the model is generalizing horribly to the validation data. I've included one of their figures as reference. Any recommendations to better replicate the error rate saturation from this paper? Note: For Kaiming He figure, bold lines are testing error & dashed are training.\n\nParameters:\n\n* 162 Epochs w/ batch size of 128 for 64k iterations.\n* Lr: 0.1\n* Momentum: 0.9\n* Weight decay: 0.0001\n* Multi-step scheduler dividing the lr by 10 at 32k and 48k iterations\n\n[My %-error](https://preview.redd.it/du9j19ay322d1.png?width=729&format=png&auto=webp&s=15d461e662e5ca823bef852b73626c5cb3f043c8)\n\n[From Paper](https://preview.redd.it/zsvtfxb0422d1.png?width=701&format=png&auto=webp&s=019edcbf04a5eff3ffb0902fbb3057fd501e9dd4)\n\n**Edit:** Adjustment to training transformation by implementing normalization prior to random cropping & re-ran all models. First image is change to plain 18 error rate curve, second is all error rates for tested architectures. \n\nhttps://preview.redd.it/g7gpcuql482d1.png?width=1000&format=png&auto=webp&s=0db10a04a16585a0401c56a7f4c2ae25f9f8249b\n\nhttps://preview.redd.it/7d35ztql482d1.png?width=1000&format=png&auto=webp&s=a83604401d1971bc7472d0b08231985a222ba8f4\n\n",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-22",
  "username_encoded": "Z0FBQUFBQm5Lak1VdzVORW9hNjZMLV85QjlRWWZRNlQxa292bFQyU0xfdzVrdks2SER5Vk9YMVdGR3JHUnAyQzJZRWlPRjdDd3JOck9XQklOM05yRXJQY0hZYzVfRC1mNVU3OVlSYWdWWnB4VDBBcTZzWEwwMDA9",
  "url_encoded": "Z0FBQUFBQm5Lak9rcHNzaVNPQVJjWmVGaU5DQ040OU43LU1ZVTkwSFBnclM2RlhlcmZSNlk3blc5NUR2cTUwbjlBRkU4eFE4M3NyRDJhYkRPYUtBVFdtTzBERXBEMnFNQi1fd3lZbTZsUjh5dzJ0bHVUUTctSDFsX2xuRWFqdFNjNFg5RFlCblg4TUhxT0tCbThyYzlOblVCVkFrNzA0NU5CRnV5Ym5QZklpNFN2bktYMDJVa043SHpuaGZ0YlF2MTRBaEhsazVmajM1VUxrUUF0T0tILUdrMzVuVmIySFRxZz09"
}

Entry Information