Row 7079
Content Data
This page contains data entry 7079 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
MRL \[1\] for CLIP allows smaller dimension embeddings to be used without loss in fidelity. Training is modified to optimize for truncated embeddings (multiple target dimensions at once) across both vision and text encoders.
Key findings:
* Reducing embeddings size by 4x retains \~95 performance * Projection layers for sub-embeddings did not help performance * Works in and out (zero-shot) of domain on multi-modal retrieval * Using too many sub-embeddings degrades performance (i.e. {512, 256, 128} vs {512, 256, 128, 64, 32, 16, 8} * The number of sub-embeddings impacts convergence (same as above) * Works with rank-tuning methods like GCL * Relative importance (weighting, wi) of sub-dimensions matters (e.g. w1\*L\_512 + w2\*L\_256 + w3\*L\_128) * MRL trained models can improve if the original sized embedding is used. i.e. performance is improved even if the smaller embeddings are not used.
Article: [https://www.marqo.ai/blog/matryoshka-representation-learning-with-clip-for-multimodal-retrieval-and-ranking](https://www.marqo.ai/blog/matryoshka-representation-learning-with-clip-for-multimodal-retrieval-and-ranking)
\[1\] MRL [https://arxiv.org/abs/2205.13147](https://arxiv.org/abs/2205.13147)
| Field | Value |
|---|---|
| text | MRL \[1\] for CLIP allows smaller dimension embeddings to be used without loss in fidelity. Training is modified to optimize for truncated embeddings (multiple target dimensions at once) across both vision and text encoders. Key findings: * Reducing embeddings size by 4x retains \~95 performance * Projection layers for sub-embeddings did not help performance * Works in and out (zero-shot) of domain on multi-modal retrieval * Using too many sub-embeddings degrades performance (i.e. {512, 2… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-15 |
| username_encoded | Z0FBQUFBQm5LakwzT2JYMC10SEZyU2o3a2lEQWRTTU00bmJjcHg5R2ZUYmRXTF9jVDgxTU9Dblptclg1eE9PcllWaDAxMERJdk00bEtFaU9oVUFXSTRMUDlSOUJyWmVTenc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HLXMwNkN3LWJiM2tPeU8yTDlsOFJaUXhJNkt5bWFraENHbFc1ekRISEFPSG4ybkxRbkJic3V4M05yMlZpS3UtZEcyeXF5TjdCTzhfVVg0ejQ3N3Q4Q0dCZUViZjJWbHloVzg4allBS2VFdW8tYk5pYm9oUXF3MXZfLW9nd2daUC1nSVQ2cExsSWRZcWNkOGNBMk42ek1EbXcyOEl4WFoyZ3E4aUhYZGNLbnVBWmk1ZHdzRlR4U1l0WGE4bzY0NzE3NnpGejlveFZTYm9NVUlITnN4dFI1dz09 |
Raw Record
{
"text": "MRL \\[1\\] for CLIP allows smaller dimension embeddings to be used without loss in fidelity. Training is modified to optimize for truncated embeddings (multiple target dimensions at once) across both vision and text encoders. \n\n \nKey findings:\n\n* Reducing embeddings size by 4x retains \\~95 performance\n* Projection layers for sub-embeddings did not help performance\n* Works in and out (zero-shot) of domain on multi-modal retrieval\n* Using too many sub-embeddings degrades performance (i.e. {512, 256, 128} vs {512, 256, 128, 64, 32, 16, 8}\n* The number of sub-embeddings impacts convergence (same as above)\n* Works with rank-tuning methods like GCL\n* Relative importance (weighting, wi) of sub-dimensions matters (e.g. w1\\*L\\_512 + w2\\*L\\_256 + w3\\*L\\_128)\n* MRL trained models can improve if the original sized embedding is used. i.e. performance is improved even if the smaller embeddings are not used.\n\nArticle: \n[https://www.marqo.ai/blog/matryoshka-representation-learning-with-clip-for-multimodal-retrieval-and-ranking](https://www.marqo.ai/blog/matryoshka-representation-learning-with-clip-for-multimodal-retrieval-and-ranking)\n\n\\[1\\] MRL [https://arxiv.org/abs/2205.13147](https://arxiv.org/abs/2205.13147)",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-15",
"username_encoded": "Z0FBQUFBQm5LakwzT2JYMC10SEZyU2o3a2lEQWRTTU00bmJjcHg5R2ZUYmRXTF9jVDgxTU9Dblptclg1eE9PcllWaDAxMERJdk00bEtFaU9oVUFXSTRMUDlSOUJyWmVTenc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HLXMwNkN3LWJiM2tPeU8yTDlsOFJaUXhJNkt5bWFraENHbFc1ekRISEFPSG4ybkxRbkJic3V4M05yMlZpS3UtZEcyeXF5TjdCTzhfVVg0ejQ3N3Q4Q0dCZUViZjJWbHloVzg4allBS2VFdW8tYk5pYm9oUXF3MXZfLW9nd2daUC1nSVQ2cExsSWRZcWNkOGNBMk42ek1EbXcyOEl4WFoyZ3E4aUhYZGNLbnVBWmk1ZHdzRlR4U1l0WGE4bzY0NzE3NnpGejlveFZTYm9NVUlITnN4dFI1dz09"
}
Entry Information
- Entry ID: 7079
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000