Row 4711
Content Data
This page contains data entry 4711 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text.
Assuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems?
Secondly, what are some ways to do such training, where the model learn on lots of rejects and lots of accepts for a single prompt.
Question Inspiration:
[https://arxiv.org/pdf/2305.20050.pdf](https://arxiv.org/pdf/2305.20050.pdf)
[https://www.youtube.com/watch?v=Zc03IYnnuIA](https://www.youtube.com/watch?v=Zc03IYnnuIA) \- The part where Altman talks about reliability of LLM outputs.
| Field | Value |
|---|---|
| text | So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text. Assuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems? Secondly, what are some ways to do such training, where the model learn on lots of… |
| label | r/deeplearning |
| dataType | post |
| communityName | r/deeplearning |
| datetime | 2024-04-20 |
| username_encoded | Z0FBQUFBQm5Lakwxc1lPQVlPbURSUzR1alY3ci0xTy15R0lZR3ZOajJHQWh0SHltdE9FZ2lIbVF1MmNjRUhWTk4xMXpkSDZxTGdIaTU3RThNZ2h6bl9vTDBuWm53Y3hZTkM0UWl1YW1mQnRQTEFnTXhyWm1RNE09 |
| url_encoded | Z0FBQUFBQm5Lak9GeVFjNXd6YnBKcUsxLTdSMGhCcC12cTE3SjRFTnUyTC1hbzhyVXZNRFBQRE1FbU9maHI1M01jaFhjNjVmaHJGREpBeTVRVUMxNDRFYWlMLUNJa1NpS3cwZHY5ZjdVUWlLUW8tamhocmp4WDBSaDZlVURRTGp3a204RVlHTkdYbUdlVlpzMzZ6VTVvNEV4a2ZoMXZoc2htNGhKajE1R2g5d3FxMS14aks4dmJaeTNzalcyVDBVcU9qSGR4bFd4TkFjdXFjVTlSRlV3bDlqY0dWX1FRRVR6dz09 |
Raw Record
{
"text": "So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text.\n\nAssuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems?\n\nSecondly, what are some ways to do such training, where the model learn on lots of rejects and lots of accepts for a single prompt.\n\nQuestion Inspiration:\n\n[https://arxiv.org/pdf/2305.20050.pdf](https://arxiv.org/pdf/2305.20050.pdf)\n\n[https://www.youtube.com/watch?v=Zc03IYnnuIA](https://www.youtube.com/watch?v=Zc03IYnnuIA) \\- The part where Altman talks about reliability of LLM outputs.",
"label": "r/deeplearning",
"dataType": "post",
"communityName": "r/deeplearning",
"datetime": "2024-04-20",
"username_encoded": "Z0FBQUFBQm5Lakwxc1lPQVlPbURSUzR1alY3ci0xTy15R0lZR3ZOajJHQWh0SHltdE9FZ2lIbVF1MmNjRUhWTk4xMXpkSDZxTGdIaTU3RThNZ2h6bl9vTDBuWm53Y3hZTkM0UWl1YW1mQnRQTEFnTXhyWm1RNE09",
"url_encoded": "Z0FBQUFBQm5Lak9GeVFjNXd6YnBKcUsxLTdSMGhCcC12cTE3SjRFTnUyTC1hbzhyVXZNRFBQRE1FbU9maHI1M01jaFhjNjVmaHJGREpBeTVRVUMxNDRFYWlMLUNJa1NpS3cwZHY5ZjdVUWlLUW8tamhocmp4WDBSaDZlVURRTGp3a204RVlHTkdYbUdlVlpzMzZ6VTVvNEV4a2ZoMXZoc2htNGhKajE1R2g5d3FxMS14aks4dmJaeTNzalcyVDBVcU9qSGR4bFd4TkFjdXFjVTlSRlV3bDlqY0dWX1FRRVR6dz09"
}
Entry Information
- Entry ID: 4711
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000