Row 6220

Row ID: 6220 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 6220 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using "query: " for both input texts for linear probing classification and symmetric tasks such as semantic similarity.

Currently my vector database stores text documents embedded with this embedding model and prefixed with "passage: " because I also read that documents should be embedded with prefix "passage: ". I want to avoid storing another vector database with the only difference being each text embedding is prefixed with "query: ".

Wondering if there's any implication on using input texts both prefixed with "passage: " and used for symmetric tasks?

Any advice or guidance is greatly appreciated! Thanks :)

FieldValue
text I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using "query: " for both input texts for linear probing classification and symmetric tasks such as semantic similarity. Currently my vector database stores text documents embedded with this embedding model and prefixed with "passage: " because I also read that documents should be embedded with prefix "passage: ". I want to avoid storing another vector database with the onl…
label r/datascience
dataType post
communityName r/datascience
datetime 2024-05-08
username_encoded Z0FBQUFBQm5LakwycGhhbFdDRzN2YmJvRGgxUUJSR3VaTGxSN2JITzhtaTMxZUpJN1QxUWQ1a1E0MnVtRktNMFdnMTlZc3hsY3Yzb25LdU5fVTViR2t3aFVmMk43dHBnQ1E9PQ==
url_encoded Z0FBQUFBQm5Lak9HUWtDMTQxTGtmZUhjdm5vcDk1dkdaeWdjZnp4WHJJRENwUnJHRmpENk12OFhZMVFHNkxrcEJKNm42UFkyT212ODZOXzAtRklDcVBUWHh3V2xXZVlUdEZXYVhFcWpoQjRJam9zdEhhbkMtOU1NSDVGV2dWNi1KaFA2cF9lc2JKa2ppaFVWRXk0OTVWLXJMbkhJWGgtLVc0SHhDcm9ZV3o3a3phVmFTV05GbTBDdGdtSFA5UmUyR1RZOTJGVEJ0ZWVabk5udnVQNG1JY3NEdHF0TGhKbnl1QT09

Raw Record

{
  "text": "I was reading multilingual-e5-large [documentation](https://huggingface.co/intfloat/multilingual-e5-large) and it suggested using \"query: \" for both input texts for linear probing classification and symmetric tasks such as semantic similarity.\n\nCurrently my vector database stores text documents embedded with this embedding model and prefixed with \"passage: \" because I also read that documents should be embedded with prefix \"passage: \". I want to avoid storing another vector database with the only difference being each text embedding is prefixed with \"query: \".\n\nWondering if there's any implication on using input texts both prefixed with \"passage: \" and used for symmetric tasks?\n\nAny advice or guidance is greatly appreciated! Thanks :)\n",
  "label": "r/datascience",
  "dataType": "post",
  "communityName": "r/datascience",
  "datetime": "2024-05-08",
  "username_encoded": "Z0FBQUFBQm5LakwycGhhbFdDRzN2YmJvRGgxUUJSR3VaTGxSN2JITzhtaTMxZUpJN1QxUWQ1a1E0MnVtRktNMFdnMTlZc3hsY3Yzb25LdU5fVTViR2t3aFVmMk43dHBnQ1E9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9HUWtDMTQxTGtmZUhjdm5vcDk1dkdaeWdjZnp4WHJJRENwUnJHRmpENk12OFhZMVFHNkxrcEJKNm42UFkyT212ODZOXzAtRklDcVBUWHh3V2xXZVlUdEZXYVhFcWpoQjRJam9zdEhhbkMtOU1NSDVGV2dWNi1KaFA2cF9lc2JKa2ppaFVWRXk0OTVWLXJMbkhJWGgtLVc0SHhDcm9ZV3o3a3phVmFTV05GbTBDdGdtSFA5UmUyR1RZOTJGVEJ0ZWVabk5udnVQNG1JY3NEdHF0TGhKbnl1QT09"
}

Entry Information