Row 26066
Content Data
This page contains data entry 26066 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
>Thanks, Unstructured and Docugami look interesting, but they both do blackbox magic on top of documents.
This is true for Docugami, that is why I was not sure whether to included it in my comment. But Unstructured provides the full suite opensource as well with docs - besides the paid API and the platform. [https://docs.unstructured.io/open-source/introduction/overview](https://docs.unstructured.io/open-source/introduction/overview) [https://github.com/Unstructured-IO](https://github.com/Unstructured-IO)
A few days ago, I had to quickly get into prepossessing pdf-s for LLM-s (the task itself is unfortunately a hornets' nest, its not an easy thing to solve), that's how I have found Unstructured and a bunch of other stuff. Other materials - which I have found, and might be interest to you: [https://medium.com/@jerryjliu98/how-unstructured-and-llamaindex-can-help-bring-the-power-of-llms-to-your-own-data-3657d063e30d](https://medium.com/@jerryjliu98/how-unstructured-and-llamaindex-can-help-bring-the-power-of-llms-to-your-own-data-3657d063e30d) [https://medium.com/@anuragmishra\_27746/five-levels-of-chunking-strategies-in-rag-notes-from-gregs-video-7b735895694d#b123](https://medium.com/@anuragmishra_27746/five-levels-of-chunking-strategies-in-rag-notes-from-gregs-video-7b735895694d#b123) [https://github.com/tstanislawek/awesome-document-understanding](https://github.com/tstanislawek/awesome-document-understanding)
| Field | Value |
|---|---|
| text | >Thanks, Unstructured and Docugami look interesting, but they both do blackbox magic on top of documents. This is true for Docugami, that is why I was not sure whether to included it in my comment. But Unstructured provides the full suite opensource as well with docs - besides the paid API and the platform. [https://docs.unstructured.io/open-source/introduction/overview](https://docs.unstructured.io/open-source/introduction/overview) [https://github.com/Unstructured-IO](https://github.com/U… |
| label | r/datascience |
| dataType | comment |
| communityName | r/datascience |
| datetime | 2024-05-21 |
| username_encoded | Z0FBQUFBQm5Lak1EVE5tcFZVZ1ptUUVTNHplN05jTHR0X1VvVjNQNWpwR1VyandZOGR5TjRWUUUyVG10RGdjV3VEWFNkRHV1WXJyYktrQU9vYmg1NmFKN2FTRUdFUnBEWEE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9TVFRCQmw1REJqUjUySkt2T1pLRzE5NjhidFhfMWdqSHZHNVVyU3hkRHh6VnVHR185ckZZQkFJSlp2cmV0MkhjeEE3Tkw1TEdzNkpNMlRrOHA0dllBQ2h4dXBZanBKbWNKZmdCYjROVV91Z1BLeEtzZXFRZEpPemI2U1N3c003LVh0RDBBOW5YbE9zelUzZE43YV9WUkV5N3NiZXdKSlhGQURJeFZTaEU1UXpDVGw2Um5ad2tZa0JEWjVOU0x4bU5KeU5CMi1JLURlUTlrNmg0Z0p6a3dKQT09 |
Raw Record
{
"text": ">Thanks, Unstructured and Docugami look interesting, but they both do blackbox magic on top of documents.\n\nThis is true for Docugami, that is why I was not sure whether to included it in my comment. But Unstructured provides the full suite opensource as well with docs - besides the paid API and the platform. \n[https://docs.unstructured.io/open-source/introduction/overview](https://docs.unstructured.io/open-source/introduction/overview) \n[https://github.com/Unstructured-IO](https://github.com/Unstructured-IO)\n\nA few days ago, I had to quickly get into prepossessing pdf-s for LLM-s (the task itself is unfortunately a hornets' nest, its not an easy thing to solve), that's how I have found Unstructured and a bunch of other stuff. Other materials - which I have found, and might be interest to you: \n[https://medium.com/@jerryjliu98/how-unstructured-and-llamaindex-can-help-bring-the-power-of-llms-to-your-own-data-3657d063e30d](https://medium.com/@jerryjliu98/how-unstructured-and-llamaindex-can-help-bring-the-power-of-llms-to-your-own-data-3657d063e30d) \n[https://medium.com/@anuragmishra\\_27746/five-levels-of-chunking-strategies-in-rag-notes-from-gregs-video-7b735895694d#b123](https://medium.com/@anuragmishra_27746/five-levels-of-chunking-strategies-in-rag-notes-from-gregs-video-7b735895694d#b123) \n[https://github.com/tstanislawek/awesome-document-understanding](https://github.com/tstanislawek/awesome-document-understanding)",
"label": "r/datascience",
"dataType": "comment",
"communityName": "r/datascience",
"datetime": "2024-05-21",
"username_encoded": "Z0FBQUFBQm5Lak1EVE5tcFZVZ1ptUUVTNHplN05jTHR0X1VvVjNQNWpwR1VyandZOGR5TjRWUUUyVG10RGdjV3VEWFNkRHV1WXJyYktrQU9vYmg1NmFKN2FTRUdFUnBEWEE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9TVFRCQmw1REJqUjUySkt2T1pLRzE5NjhidFhfMWdqSHZHNVVyU3hkRHh6VnVHR185ckZZQkFJSlp2cmV0MkhjeEE3Tkw1TEdzNkpNMlRrOHA0dllBQ2h4dXBZanBKbWNKZmdCYjROVV91Z1BLeEtzZXFRZEpPemI2U1N3c003LVh0RDBBOW5YbE9zelUzZE43YV9WUkV5N3NiZXdKSlhGQURJeFZTaEU1UXpDVGw2Um5ad2tZa0JEWjVOU0x4bU5KeU5CMi1JLURlUTlrNmg0Z0p6a3dKQT09"
}
Entry Information
- Entry ID: 26066
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000