Row 92921
Content Data
This page contains data entry 92921 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
In PPO, we clip the per-token advantage weight which is normally
P_policy(action|state) / P_previous(action|state)
to (1-eps, 1+eps) to prevent too-destructive updates of the policy.
My question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting point, wouldn't we actually want to clip the probability ratio at (0.5, 2)? Which would be (1 / eps, 1 * eps) instead of (1 - eps, 1 + eps)
| Field | Value |
|---|---|
| text | In PPO, we clip the per-token advantage weight which is normally P_policy(action|state) / P_previous(action|state) to (1-eps, 1+eps) to prevent too-destructive updates of the policy. My question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting … |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-25 |
| username_encoded | Z0FBQUFBQm5Lak1zYy0xQjZNUkpGd05nT1BDYlBlT0l3dEYwR19fbjEyZnNlMWxGbUNlVlNGU2ZBVGRjU01iRkdpcHVvR0JGajJjTE1ZZVMwNFQ3MzRHcWpSVmdHMWVwQ3c9PQ== |
| url_encoded | Z0FBQUFBQm5Lak8tRWZxLWpMdjJLY2paS2JYVm8zNWo0dlNucmZKbHNzamFSdnJxaTRhLWhlWE9LWUZ0aFNPTm9LTjFXS3JfZ0cxbDBQU1h6MUI4SG1rYzlYeXRRaHBmOVBscC1iS0hNQ0FKSnFUY0ZVdWQwTHVJb1FUREU1UHQtYjNLQUttY2RzM0hXWExIbW8xLXVlRWFteDJvVlVnOU45aWVqaDhxQldFNFBpdXo5MWV3SXotTnpVT2FtYktXd0p6NXRXUHFkbVpqaVB1T0FOS2tEWWM0MUhZbjlOdHYwZz09 |
Raw Record
{
"text": "In PPO, we clip the per-token advantage weight which is normally\n\nP_policy(action|state) / P_previous(action|state) \n\nto (1-eps, 1+eps) to prevent too-destructive updates of the policy.\n\nMy question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting point, wouldn't we actually want to clip the probability ratio at (0.5, 2)? Which would be (1 / eps, 1 * eps) instead of (1 - eps, 1 + eps)",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-25",
"username_encoded": "Z0FBQUFBQm5Lak1zYy0xQjZNUkpGd05nT1BDYlBlT0l3dEYwR19fbjEyZnNlMWxGbUNlVlNGU2ZBVGRjU01iRkdpcHVvR0JGajJjTE1ZZVMwNFQ3MzRHcWpSVmdHMWVwQ3c9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak8tRWZxLWpMdjJLY2paS2JYVm8zNWo0dlNucmZKbHNzamFSdnJxaTRhLWhlWE9LWUZ0aFNPTm9LTjFXS3JfZ0cxbDBQU1h6MUI4SG1rYzlYeXRRaHBmOVBscC1iS0hNQ0FKSnFUY0ZVdWQwTHVJb1FUREU1UHQtYjNLQUttY2RzM0hXWExIbW8xLXVlRWFteDJvVlVnOU45aWVqaDhxQldFNFBpdXo5MWV3SXotTnpVT2FtYktXd0p6NXRXUHFkbVpqaVB1T0FOS2tEWWM0MUhZbjlOdHYwZz09"
}
Entry Information
- Entry ID: 92921
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000