Row 24881

Row ID: 24881 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 24881 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

When dealing with RAG, or information retrieval in general, extraction and chunking along with indexing are the most relevant sliders to fine tune the process and therefore the retrieval quality.

Are there tools available to experiment with different extraction and chunking methods? I know there's like 1000 No-Code UIs to create a Chat-Bot, but the RAG part is mostly just a black box that says "drop your PDF here".

I'm thinking about features like

* Clean the content before processing (HTML to Markdown) * Work with Summaries vs Full Text * Extract Facts & Questions * Extract Short Snippets vs Paragraphs * Extract Relations and Graph Information * Sentence vs Token Chunking * Vector Index vs Full Text Search

Basically everything that happens before passing the context to the LLM. Doesn't have to be super fancy, but is there anything better than just creating a bunch of Jupyter Notebooks and running benchmarks?

FieldValue
text When dealing with RAG, or information retrieval in general, extraction and chunking along with indexing are the most relevant sliders to fine tune the process and therefore the retrieval quality. Are there tools available to experiment with different extraction and chunking methods? I know there's like 1000 No-Code UIs to create a Chat-Bot, but the RAG part is mostly just a black box that says "drop your PDF here". I'm thinking about features like * Clean the content before processing (HTML t…
label r/datascience
dataType post
communityName r/datascience
datetime 2024-05-21
username_encoded Z0FBQUFBQm5Lak1DZmVsLTFDX0hiMWZILWgwM1lVVm1hNm8zNWYtN2ZvWUZLMjF2QXVCOEYyQmhGZV9vSnZhOXFLdWJsaWg1MHZlbkJDMHBIdkM3cGtMOUlyZVJXSUJvbnc9PQ==
url_encoded Z0FBQUFBQm5Lak9SdktiNENrcnpiaU9tY2tnR05rMVJBOHR5YVpxN3FmTkxHTEVEZHRUZEt3eUk5VUZFNy1mZDVrQ2VxbkF5ZXFKQ3RPa1VIOUd4OHp4M2l5YUV5RElVSUFSY0dWZ1JpMkZScVRiQUJTZHNCMTk4TnZyWm9pUF9wLVJnM3A2TWJmR0VGV3NMa041cmdkQTdxSG00OEcxS0Q5cFFoSlBVVlV2OW1YMHBieHg1VThtcHpBdzFJZ2ZCcW00Ny1KNU44aFh6

Raw Record

{
  "text": "When dealing with RAG, or information retrieval in general, extraction and chunking along with indexing are the most relevant sliders to fine tune the process and therefore the retrieval quality.\n\nAre there tools available to experiment with different extraction and chunking methods? I know there's like 1000 No-Code UIs to create a Chat-Bot, but the RAG part is mostly just a black box that says \"drop your PDF here\".\n\nI'm thinking about features like\n\n* Clean the content before processing (HTML to Markdown)\n* Work with Summaries vs Full Text\n* Extract Facts & Questions\n* Extract Short Snippets vs Paragraphs\n* Extract Relations and Graph Information\n* Sentence vs Token Chunking\n* Vector Index vs Full Text Search\n\nBasically everything that happens before passing the context to the LLM. Doesn't have to be super fancy, but is there anything better than just creating a bunch of Jupyter Notebooks and running benchmarks?",
  "label": "r/datascience",
  "dataType": "post",
  "communityName": "r/datascience",
  "datetime": "2024-05-21",
  "username_encoded": "Z0FBQUFBQm5Lak1DZmVsLTFDX0hiMWZILWgwM1lVVm1hNm8zNWYtN2ZvWUZLMjF2QXVCOEYyQmhGZV9vSnZhOXFLdWJsaWg1MHZlbkJDMHBIdkM3cGtMOUlyZVJXSUJvbnc9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9SdktiNENrcnpiaU9tY2tnR05rMVJBOHR5YVpxN3FmTkxHTEVEZHRUZEt3eUk5VUZFNy1mZDVrQ2VxbkF5ZXFKQ3RPa1VIOUd4OHp4M2l5YUV5RElVSUFSY0dWZ1JpMkZScVRiQUJTZHNCMTk4TnZyWm9pUF9wLVJnM3A2TWJmR0VGV3NMa041cmdkQTdxSG00OEcxS0Q5cFFoSlBVVlV2OW1YMHBieHg1VThtcHpBdzFJZ2ZCcW00Ny1KNU44aFh6"
}

Entry Information