Row 4711

Row ID: 4711 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 4711 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text.

Assuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems?

Secondly, what are some ways to do such training, where the model learn on lots of rejects and lots of accepts for a single prompt.

Question Inspiration:

[https://arxiv.org/pdf/2305.20050.pdf](https://arxiv.org/pdf/2305.20050.pdf)

[https://www.youtube.com/watch?v=Zc03IYnnuIA](https://www.youtube.com/watch?v=Zc03IYnnuIA) \- The part where Altman talks about reliability of LLM outputs.

FieldValue
text So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text. Assuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems? Secondly, what are some ways to do such training, where the model learn on lots of…
label r/deeplearning
dataType post
communityName r/deeplearning
datetime 2024-04-20
username_encoded Z0FBQUFBQm5Lakwxc1lPQVlPbURSUzR1alY3ci0xTy15R0lZR3ZOajJHQWh0SHltdE9FZ2lIbVF1MmNjRUhWTk4xMXpkSDZxTGdIaTU3RThNZ2h6bl9vTDBuWm53Y3hZTkM0UWl1YW1mQnRQTEFnTXhyWm1RNE09
url_encoded Z0FBQUFBQm5Lak9GeVFjNXd6YnBKcUsxLTdSMGhCcC12cTE3SjRFTnUyTC1hbzhyVXZNRFBQRE1FbU9maHI1M01jaFhjNjVmaHJGREpBeTVRVUMxNDRFYWlMLUNJa1NpS3cwZHY5ZjdVUWlLUW8tamhocmp4WDBSaDZlVURRTGp3a204RVlHTkdYbUdlVlpzMzZ6VTVvNEV4a2ZoMXZoc2htNGhKajE1R2g5d3FxMS14aks4dmJaeTNzalcyVDBVcU9qSGR4bFd4TkFjdXFjVTlSRlV3bDlqY0dWX1FRRVR6dz09

Raw Record

{
  "text": "So, when we ask an LLM a query lots of time, there might 1 among all of that, which is correct, and highly preferred. Current DPO method, can learn between one chosen text, and one rejected text.\n\nAssuming there comes a method to show model n-number of rejected samples and m-number of accepted samples, what are your thoughts that, this type of training, will dramatically increase the reliability of these systems?\n\nSecondly, what are some ways to do such training, where the model learn on lots of rejects and lots of accepts for a single prompt.\n\nQuestion Inspiration:\n\n[https://arxiv.org/pdf/2305.20050.pdf](https://arxiv.org/pdf/2305.20050.pdf)\n\n[https://www.youtube.com/watch?v=Zc03IYnnuIA](https://www.youtube.com/watch?v=Zc03IYnnuIA) \\- The part where Altman talks about reliability of LLM outputs.",
  "label": "r/deeplearning",
  "dataType": "post",
  "communityName": "r/deeplearning",
  "datetime": "2024-04-20",
  "username_encoded": "Z0FBQUFBQm5Lakwxc1lPQVlPbURSUzR1alY3ci0xTy15R0lZR3ZOajJHQWh0SHltdE9FZ2lIbVF1MmNjRUhWTk4xMXpkSDZxTGdIaTU3RThNZ2h6bl9vTDBuWm53Y3hZTkM0UWl1YW1mQnRQTEFnTXhyWm1RNE09",
  "url_encoded": "Z0FBQUFBQm5Lak9GeVFjNXd6YnBKcUsxLTdSMGhCcC12cTE3SjRFTnUyTC1hbzhyVXZNRFBQRE1FbU9maHI1M01jaFhjNjVmaHJGREpBeTVRVUMxNDRFYWlMLUNJa1NpS3cwZHY5ZjdVUWlLUW8tamhocmp4WDBSaDZlVURRTGp3a204RVlHTkdYbUdlVlpzMzZ6VTVvNEV4a2ZoMXZoc2htNGhKajE1R2g5d3FxMS14aks4dmJaeTNzalcyVDBVcU9qSGR4bFd4TkFjdXFjVTlSRlV3bDlqY0dWX1FRRVR6dz09"
}

Entry Information