Row 27739

Row ID: 27739 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 27739 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

https://preview.redd.it/jg317tkdur1d1.png?width=1827&format=png&auto=webp&s=de12fa29edd5acbea05545404ab00cf2984ca185

Left: Dino v2 embeddings. (384 dimensions) Right: PyTorch ResNet18 pretrained without last linear layer. (512 dimensions)

Input: [https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion](https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion)

I didn't do any finetuning, just took the dataset and feed into Dino and ResNet, then I used t-SNE to reduce the dimension of the embeddings to 2.

Why pretrained ResNet seems to do a better job clustering pictures by color compared to Dino ? They have been trained using totally different approach, i know. But none of them have been trained for color distinction. Just I would like to start a discussion with you. Thanks!

FieldValue
text https://preview.redd.it/jg317tkdur1d1.png?width=1827&format=png&auto=webp&s=de12fa29edd5acbea05545404ab00cf2984ca185 Left: Dino v2 embeddings. (384 dimensions) Right: PyTorch ResNet18 pretrained without last linear layer. (512 dimensions) Input: [https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion](https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion) I didn't do any finetuning, just took the dataset a…
label r/deeplearning
dataType post
communityName r/deeplearning
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1FS0VqWGtfMzItZnlKdUFYVEFpek9BUVNfYlpZMWRzcElLaHdBTkJFQlJxTmpjQzduQ3lxSFlEMkh3MUxJWGxPVGhUMTlyeVNfSmVheE5WT0hYNUU2cEE9PQ==
url_encoded Z0FBQUFBQm5Lak9URlR3SWc3VDdfNFFpVnBGYU9BSnYzbG1sUjVPOWFDZF9saG5mSUFpdGtnd1VVaDlrNGxwUjVNZkxWWDJBNGlNVDhwUHBJTWdtTVJzeGFaSTBTRjhoRFBNVTJmTmdrSzFSYW9sdzg1SDl1T2FQZmFjb2FlYmpOWjlYS3U2eW5wYWprR1l3MjIxckQ5OTRiN2ZGMjJCWnZwTVd1eC02R080QW45Uld2eEZNZl9MSUZ0U2ZCbklYYXU4eDJ2MElIcUFW

Raw Record

{
  "text": "https://preview.redd.it/jg317tkdur1d1.png?width=1827&format=png&auto=webp&s=de12fa29edd5acbea05545404ab00cf2984ca185\n\nLeft: Dino v2 embeddings. (384 dimensions)  \nRight: PyTorch ResNet18 pretrained without last linear layer. (512 dimensions)  \n\n\nInput: [https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion](https://www.kaggle.com/datasets/landrykezebou/vcor-vehicle-color-recognition-dataset/discussion)\n\nI didn't do any finetuning, just took the dataset and feed into Dino and ResNet, then I used t-SNE to reduce the dimension of the embeddings to 2.\n\nWhy pretrained ResNet seems to do a better job clustering pictures by color compared to Dino ? They have been trained using totally different approach, i know. But none of them have been trained for color distinction. Just I would like to start a discussion with you. Thanks!",
  "label": "r/deeplearning",
  "dataType": "post",
  "communityName": "r/deeplearning",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1FS0VqWGtfMzItZnlKdUFYVEFpek9BUVNfYlpZMWRzcElLaHdBTkJFQlJxTmpjQzduQ3lxSFlEMkh3MUxJWGxPVGhUMTlyeVNfSmVheE5WT0hYNUU2cEE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9URlR3SWc3VDdfNFFpVnBGYU9BSnYzbG1sUjVPOWFDZF9saG5mSUFpdGtnd1VVaDlrNGxwUjVNZkxWWDJBNGlNVDhwUHBJTWdtTVJzeGFaSTBTRjhoRFBNVTJmTmdrSzFSYW9sdzg1SDl1T2FQZmFjb2FlYmpOWjlYS3U2eW5wYWprR1l3MjIxckQ5OTRiN2ZGMjJCWnZwTVd1eC02R080QW45Uld2eEZNZl9MSUZ0U2ZCbklYYXU4eDJ2MElIcUFW"
}

Entry Information