Row 6220
Content Data
This page contains data entry 6220 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using "query: " for both input texts for linear probing classification and symmetric tasks such as semantic similarity.
Currently my vector database stores text documents embedded with this embedding model and prefixed with "passage: " because I also read that documents should be embedded with prefix "passage: ". I want to avoid storing another vector database with the only difference being each text embedding is prefixed with "query: ".
Wondering if there's any implication on using input texts both prefixed with "passage: " and used for symmetric tasks?
Any advice or guidance is greatly appreciated! Thanks :)
| Field | Value |
|---|---|
| text | I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using "query: " for both input texts for linear probing classification and symmetric tasks such as semantic similarity. Currently my vector database stores text documents embedded with this embedding model and prefixed with "passage: " because I also read that documents should be embedded with prefix "passage: ". I want to avoid storing another vector database with the onl… |
| label | r/datascience |
| dataType | post |
| communityName | r/datascience |
| datetime | 2024-05-08 |
| username_encoded | Z0FBQUFBQm5LakwycGhhbFdDRzN2YmJvRGgxUUJSR3VaTGxSN2JITzhtaTMxZUpJN1QxUWQ1a1E0MnVtRktNMFdnMTlZc3hsY3Yzb25LdU5fVTViR2t3aFVmMk43dHBnQ1E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9HUWtDMTQxTGtmZUhjdm5vcDk1dkdaeWdjZnp4WHJJRENwUnJHRmpENk12OFhZMVFHNkxrcEJKNm42UFkyT212ODZOXzAtRklDcVBUWHh3V2xXZVlUdEZXYVhFcWpoQjRJam9zdEhhbkMtOU1NSDVGV2dWNi1KaFA2cF9lc2JKa2ppaFVWRXk0OTVWLXJMbkhJWGgtLVc0SHhDcm9ZV3o3a3phVmFTV05GbTBDdGdtSFA5UmUyR1RZOTJGVEJ0ZWVabk5udnVQNG1JY3NEdHF0TGhKbnl1QT09 |
Raw Record
{
"text": "I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using \"query: \" for both input texts for linear probing classification and symmetric tasks such as semantic similarity.\n\nCurrently my vector database stores text documents embedded with this embedding model and prefixed with \"passage: \" because I also read that documents should be embedded with prefix \"passage: \". I want to avoid storing another vector database with the only difference being each text embedding is prefixed with \"query: \".\n\nWondering if there's any implication on using input texts both prefixed with \"passage: \" and used for symmetric tasks?\n\nAny advice or guidance is greatly appreciated! Thanks :)\n",
"label": "r/datascience",
"dataType": "post",
"communityName": "r/datascience",
"datetime": "2024-05-08",
"username_encoded": "Z0FBQUFBQm5LakwycGhhbFdDRzN2YmJvRGgxUUJSR3VaTGxSN2JITzhtaTMxZUpJN1QxUWQ1a1E0MnVtRktNMFdnMTlZc3hsY3Yzb25LdU5fVTViR2t3aFVmMk43dHBnQ1E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9HUWtDMTQxTGtmZUhjdm5vcDk1dkdaeWdjZnp4WHJJRENwUnJHRmpENk12OFhZMVFHNkxrcEJKNm42UFkyT212ODZOXzAtRklDcVBUWHh3V2xXZVlUdEZXYVhFcWpoQjRJam9zdEhhbkMtOU1NSDVGV2dWNi1KaFA2cF9lc2JKa2ppaFVWRXk0OTVWLXJMbkhJWGgtLVc0SHhDcm9ZV3o3a3phVmFTV05GbTBDdGdtSFA5UmUyR1RZOTJGVEJ0ZWVabk5udnVQNG1JY3NEdHF0TGhKbnl1QT09"
}
Entry Information
- Entry ID: 6220
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000