Row 99109
Content Data
This page contains data entry 99109 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I have three and posted but not really getting answers. Hope you can help, I am pretty new to this.
This is around LLMs.
**First question**
I think I have the concept around LLMs, I have been looking at tensorflow and keras and llama2. I know this gets into the detail but I like to roll my own stuff for learning for better or worse. There is a model reader in tensor flow to read llama2 binary files. I still can't get a binary format for it. What is it? Pickle based? I even asked chatgpt and it says there is no format. How can you not have a standard format. What is there if I were to byte by yte look at one. What is an example one from hugging face. Can i visualize a small one?
**Second Question**
Same lines. I am still not clear how people build the llama2 binaries. I need to read more and watch videos. I know there is a binary, they will see wizard of oz and then hey, here is a chat. Hold on, what are all the steps? What are the weights? How are they built? Can I tweak them? Can I pre-train and how?
**Third Question**
With that said, I have a blog, crappy one but I figure I can build MY own llm against that, also tweaked with public book data. What are steps to do that, step by step for dumb newbies. I see steps from wizard oz then cuda, pytorth. I dont know, if it is a simple demo, I wouldn't gpu accel in it.
I also want to build a language, llm around povray ray tracing see here. This is mix of programming and docs. How to do that too? How do they build llms around programming?
[https://www.povray.org/](https://www.povray.org/) Possbly one for libgdx [https://libgdx.com/](https://libgdx.com/)
**OK Fourth Question - Legal**
I am surprised the legal question doesnt come up. I guess it doesn't matter. For example, I see the spaces in hugging face and think, this can't be legal. Some of it. Meaning, taking CNN data and putting it through a LLM. Also, I ask because I want to run my blogthrough a llm and then repost things. But it is my data, it is public to me. But what about reposting llm data from say llama2. What license would allow that?
| Field | Value |
|---|---|
| text | I have three and posted but not really getting answers. Hope you can help, I am pretty new to this. This is around LLMs. **First question** I think I have the concept around LLMs, I have been looking at tensorflow and keras and llama2. I know this gets into the detail but I like to roll my own stuff for learning for better or worse. There is a model reader in tensor flow to read llama2 binary files. I still can't get a binary format for it. What is it? Pickle based? I even asked chatgpt a… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak13Wk5YU1BhbWh5amFWQjlpLU11VFFhb3lNYWlQTmpnRkRTQnhlMnBmSmpWYkVidDFuWTluYlNqN0JLMnlsbWg2cnozTEg4WVV5QXQ0SWVNSGVsX0V2dHc9PQ== |
| url_encoded | Z0FBQUFBQm5LalBDNXp2bEJtZ3pjRDBaNU8xV3V2UUoyZkZCbW1KclVSOTRTNXhtMVBQRnpmN0dOV01aRUZ1azl2bGhrdEtzODNDd0x6Tkk1SFIwLVhyZFhmUHZkdS1aLWFKengyT2pfYlZjZUU0MG54a2FZQndBS1YyS0RmTHlZOXE2SVFRVktpekkta1VTN2pPMGdHYlJEUGpZWHBIdFh2N0c1WmFQVXhhaURqVkJZVG15T1U4TW8xZmVxQkhna3VFem5PVVF5SVRu |
Raw Record
{
"text": "I have three and posted but not really getting answers. Hope you can help, I am pretty new to this.\n\nThis is around LLMs.\n\n**First question**\n\nI think I have the concept around LLMs, I have been looking at tensorflow and keras and llama2. I know this gets into the detail but I like to roll my own stuff for learning for better or worse. There is a model reader in tensor flow to read llama2 binary files. I still can't get a binary format for it. What is it? Pickle based? I even asked chatgpt and it says there is no format. How can you not have a standard format. What is there if I were to byte by yte look at one. What is an example one from hugging face. Can i visualize a small one?\n\n**Second Question**\n\nSame lines. I am still not clear how people build the llama2 binaries. I need to read more and watch videos. I know there is a binary, they will see wizard of oz and then hey, here is a chat. Hold on, what are all the steps? What are the weights? How are they built? Can I tweak them? Can I pre-train and how?\n\n**Third Question**\n\nWith that said, I have a blog, crappy one but I figure I can build MY own llm against that, also tweaked with public book data. What are steps to do that, step by step for dumb newbies. I see steps from wizard oz then cuda, pytorth. I dont know, if it is a simple demo, I wouldn't gpu accel in it.\n\nI also want to build a language, llm around povray ray tracing see here. This is mix of programming and docs. How to do that too? How do they build llms around programming? \n\n[https://www.povray.org/](https://www.povray.org/) \nPossbly one for libgdx \n[https://libgdx.com/](https://libgdx.com/)\n\n**OK Fourth Question - Legal**\n\nI am surprised the legal question doesnt come up. I guess it doesn't matter. For example, I see the spaces in hugging face and think, this can't be legal. Some of it. Meaning, taking CNN data and putting it through a LLM. Also, I ask because I want to run my blogthrough a llm and then repost things. But it is my data, it is public to me. But what about reposting llm data from say llama2. What license would allow that?",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak13Wk5YU1BhbWh5amFWQjlpLU11VFFhb3lNYWlQTmpnRkRTQnhlMnBmSmpWYkVidDFuWTluYlNqN0JLMnlsbWg2cnozTEg4WVV5QXQ0SWVNSGVsX0V2dHc9PQ==",
"url_encoded": "Z0FBQUFBQm5LalBDNXp2bEJtZ3pjRDBaNU8xV3V2UUoyZkZCbW1KclVSOTRTNXhtMVBQRnpmN0dOV01aRUZ1azl2bGhrdEtzODNDd0x6Tkk1SFIwLVhyZFhmUHZkdS1aLWFKengyT2pfYlZjZUU0MG54a2FZQndBS1YyS0RmTHlZOXE2SVFRVktpekkta1VTN2pPMGdHYlJEUGpZWHBIdFh2N0c1WmFQVXhhaURqVkJZVG15T1U4TW8xZmVxQkhna3VFem5PVVF5SVRu"
}
Entry Information
- Entry ID: 99109
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000