Row 92921

Row ID: 92921 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 92921 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

In PPO, we clip the per-token advantage weight which is normally

P_policy(action|state) / P_previous(action|state)

to (1-eps, 1+eps) to prevent too-destructive updates of the policy.

My question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting point, wouldn't we actually want to clip the probability ratio at (0.5, 2)? Which would be (1 / eps, 1 * eps) instead of (1 - eps, 1 + eps)

FieldValue
text In PPO, we clip the per-token advantage weight which is normally P_policy(action|state) / P_previous(action|state) to (1-eps, 1+eps) to prevent too-destructive updates of the policy. My question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting …
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-25
username_encoded Z0FBQUFBQm5Lak1zYy0xQjZNUkpGd05nT1BDYlBlT0l3dEYwR19fbjEyZnNlMWxGbUNlVlNGU2ZBVGRjU01iRkdpcHVvR0JGajJjTE1ZZVMwNFQ3MzRHcWpSVmdHMWVwQ3c9PQ==
url_encoded Z0FBQUFBQm5Lak8tRWZxLWpMdjJLY2paS2JYVm8zNWo0dlNucmZKbHNzamFSdnJxaTRhLWhlWE9LWUZ0aFNPTm9LTjFXS3JfZ0cxbDBQU1h6MUI4SG1rYzlYeXRRaHBmOVBscC1iS0hNQ0FKSnFUY0ZVdWQwTHVJb1FUREU1UHQtYjNLQUttY2RzM0hXWExIbW8xLXVlRWFteDJvVlVnOU45aWVqaDhxQldFNFBpdXo5MWV3SXotTnpVT2FtYktXd0p6NXRXUHFkbVpqaVB1T0FOS2tEWWM0MUhZbjlOdHYwZz09

Raw Record

{
  "text": "In PPO, we clip the per-token advantage weight which is normally\n\nP_policy(action|state) / P_previous(action|state) \n\nto (1-eps, 1+eps) to prevent too-destructive updates of the policy.\n\nMy question is, since we are clipping the probability ratio of the current policy vs the previous policy, why is the clipping not symmetric? Ie, shouldn't eps be divided / multiplied, instead of subtracted / added? For example, if we wanted to prevent the current policy from deviating within 0.5 of its starting point, wouldn't we actually want to clip the probability ratio at (0.5, 2)? Which would be (1 / eps, 1 * eps) instead of (1 - eps, 1 + eps)",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-25",
  "username_encoded": "Z0FBQUFBQm5Lak1zYy0xQjZNUkpGd05nT1BDYlBlT0l3dEYwR19fbjEyZnNlMWxGbUNlVlNGU2ZBVGRjU01iRkdpcHVvR0JGajJjTE1ZZVMwNFQ3MzRHcWpSVmdHMWVwQ3c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak8tRWZxLWpMdjJLY2paS2JYVm8zNWo0dlNucmZKbHNzamFSdnJxaTRhLWhlWE9LWUZ0aFNPTm9LTjFXS3JfZ0cxbDBQU1h6MUI4SG1rYzlYeXRRaHBmOVBscC1iS0hNQ0FKSnFUY0ZVdWQwTHVJb1FUREU1UHQtYjNLQUttY2RzM0hXWExIbW8xLXVlRWFteDJvVlVnOU45aWVqaDhxQldFNFBpdXo5MWV3SXotTnpVT2FtYktXd0p6NXRXUHFkbVpqaVB1T0FOS2tEWWM0MUhZbjlOdHYwZz09"
}

Entry Information