Row 62638
Content Data
This page contains data entry 62638 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Okay so I have been exploring the world of RAGs and I am thinking of a concept called index of indices.
I am working on a large database of a company \[Consider a Fortune 2500\]. There are multiple functions (100+) across the company which do not really overlap too much - marketing & manufacturing for instance.
Now for various teams within the company I want to give each department a RAG enabled chat system to play with. Someone from marketing department does some search on a topic like - "What was CAC for Product A in the year 2023?". Now it goes and looks for that information across vectors in an index, but in parallel it also looks for the information in 99 other indices (considering 100 departments = 100 indices). If top\_k=10, then likely, it's going to come up with 1000 results, which would be passed as context to LLM which may become very large resulting in increased latency and pricing.
I am thinking of creating a master index, that has top\_k = 10 (example) which then selects 10 most relevant indices based on what a user is looking for and then searches within those 10 indices, giving 100 results instead of 1000.
I know these numbers would need to be optimized a lot, but has anyone heard about creating a vector database (on chroma or pinecone) where there's a master index and sub indices. Is it recommended for a RAG model? Is there any reading material on it?
Thanks!
| Field | Value |
|---|---|
| text | Okay so I have been exploring the world of RAGs and I am thinking of a concept called index of indices. I am working on a large database of a company \[Consider a Fortune 2500\]. There are multiple functions (100+) across the company which do not really overlap too much - marketing & manufacturing for instance. Now for various teams within the company I want to give each department a RAG enabled chat system to play with. Someone from marketing department does some search on a topic like - "W… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1heXo2OENCbHQ0cG5zRktfdEdLbmlFWTBOSFQyY3Z1S0NNOVZpZ2NFcFdOUlpDQzc4Y3o1UlRpVXZJa2J4dmRXUWZSNkFIUGt5bC1DTE9sbmdpbjZocjZ3dXRvck9QSmxDbU53NzRXZGZTRGs9 |
| url_encoded | Z0FBQUFBQm5Lak9xZHZQQzdPSlc3bE84cFIwdVVhU21QQTM3YUZuSVNRRTRTVDVfbk5YNXNabEwwTU5lTEliMjQtYWtqUDBIVThUdHgwLUhuMWVGR0gxUVBpWEp1Q0xmVmtmbVpST2JaOUNtcTZlWkU2Rmp3YUp5MFdoRHZONkg4ZXN3Q0VfblVjd3dKcjZ1Uk1GSkFHc2dvVXVudm1aRXFsZXZBN0Y5MXpXVF81dnBwUERsb3hrPQ== |
Raw Record
{
"text": "Okay so I have been exploring the world of RAGs and I am thinking of a concept called index of indices. \n\nI am working on a large database of a company \\[Consider a Fortune 2500\\]. There are multiple functions (100+) across the company which do not really overlap too much - marketing & manufacturing for instance. \n\nNow for various teams within the company I want to give each department a RAG enabled chat system to play with. Someone from marketing department does some search on a topic like - \"What was CAC for Product A in the year 2023?\". Now it goes and looks for that information across vectors in an index, but in parallel it also looks for the information in 99 other indices (considering 100 departments = 100 indices). If top\\_k=10, then likely, it's going to come up with 1000 results, which would be passed as context to LLM which may become very large resulting in increased latency and pricing. \n\nI am thinking of creating a master index, that has top\\_k = 10 (example) which then selects 10 most relevant indices based on what a user is looking for and then searches within those 10 indices, giving 100 results instead of 1000. \n\nI know these numbers would need to be optimized a lot, but has anyone heard about creating a vector database (on chroma or pinecone) where there's a master index and sub indices. Is it recommended for a RAG model? Is there any reading material on it? \n\nThanks! ",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1heXo2OENCbHQ0cG5zRktfdEdLbmlFWTBOSFQyY3Z1S0NNOVZpZ2NFcFdOUlpDQzc4Y3o1UlRpVXZJa2J4dmRXUWZSNkFIUGt5bC1DTE9sbmdpbjZocjZ3dXRvck9QSmxDbU53NzRXZGZTRGs9",
"url_encoded": "Z0FBQUFBQm5Lak9xZHZQQzdPSlc3bE84cFIwdVVhU21QQTM3YUZuSVNRRTRTVDVfbk5YNXNabEwwTU5lTEliMjQtYWtqUDBIVThUdHgwLUhuMWVGR0gxUVBpWEp1Q0xmVmtmbVpST2JaOUNtcTZlWkU2Rmp3YUp5MFdoRHZONkg4ZXN3Q0VfblVjd3dKcjZ1Uk1GSkFHc2dvVXVudm1aRXFsZXZBN0Y5MXpXVF81dnBwUERsb3hrPQ=="
}
Entry Information
- Entry ID: 62638
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000