Row 66038
Content Data
This page contains data entry 66038 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
I understand the intuition about the I-projection and M-projection approximating the posterior over latents differently and how that would be desirable/non-desirable for different models. My question is about the consequences of the different objectives on the predictive distribution of the data.
Ok, so being more specific. Model is given as p(x,z) = p(z)p(x|z) where x is 'data' and z is 'latent' variable. Given x, I approximate the intractable posterior p(z|x) with q(z) by either minimizing KL(q(z)||p(z|x)) or KL(p(z|x)||q(z)) wrt q. My predictive distribution over x is then given as p(x) = ∫q(z)p(x|z)dz.
Question: what is the difference in p(x) when approximating q(z) w/ either rKL or fKL? Anything we can say about the variance of p(x) that generally holds for certain models? Any works that study this question? For example for sparse GP models, where z is a latent function, \[4\] finds that the rKL overestimates the variance of p(x).
| Field | Value |
|---|---|
| text | I understand the intuition about the I-projection and M-projection approximating the posterior over latents differently and how that would be desirable/non-desirable for different models. My question is about the consequences of the different objectives on the predictive distribution of the data. Ok, so being more specific. Model is given as p(x,z) = p(z)p(x|z) where x is 'data' and z is 'latent' variable. Given x, I approximate the intractable posterior p(z|x) with q(z) by either minimizing KL… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1jUXNiVG8zdXJiSUc2R09VSHgxN2NZR1N3T1FYSmo2Qzg5amExQVpFYjZqNlZrUzQwdkUySm1ZRXc0RmZVaVdXMlBHdE90QU84bm9ablVjNGlWYU9LRFE9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9zNXRacURsaU9ZOUZnaUpBWDA3LU1lUDE2MjVlYUxFTzlzZ2hFNEF3dXJubFM1UUdMSHQyNTFKN1hBT2gyZmlYUjdsbGp3MDFoQUZER050WDlxX2xQemFzM0IzaW9QbjN6MU5tZ0p0SkJlSElQaGxPdmVDcjhfNi0zN0lFRkJYTUtMVW9GOGNFdTl2ZUVFR2pFN05HV2NyaENSZk9aX3JXbHNlaURmWnFFSVl6YjB1U3M2dDdQTk1TUXJvRzNqQnRadHF0VE13UkNFc0p4eFlWY0NhWG9PLS10WDdjOXY0Q3lGeXZVTGEzNUNiZz0= |
Raw Record
{
"text": "I understand the intuition about the I-projection and M-projection approximating the posterior over latents differently and how that would be desirable/non-desirable for different models. My question is about the consequences of the different objectives on the predictive distribution of the data.\n\nOk, so being more specific. Model is given as p(x,z) = p(z)p(x|z) where x is 'data' and z is 'latent' variable. Given x, I approximate the intractable posterior p(z|x) with q(z) by either minimizing KL(q(z)||p(z|x)) or KL(p(z|x)||q(z)) wrt q. My predictive distribution over x is then given as p(x) = ∫q(z)p(x|z)dz. \n\nQuestion: what is the difference in p(x) when approximating q(z) w/ either rKL or fKL? Anything we can say about the variance of p(x) that generally holds for certain models? Any works that study this question? For example for sparse GP models, where z is a latent function, \\[4\\] finds that the rKL overestimates the variance of p(x).",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1jUXNiVG8zdXJiSUc2R09VSHgxN2NZR1N3T1FYSmo2Qzg5amExQVpFYjZqNlZrUzQwdkUySm1ZRXc0RmZVaVdXMlBHdE90QU84bm9ablVjNGlWYU9LRFE9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9zNXRacURsaU9ZOUZnaUpBWDA3LU1lUDE2MjVlYUxFTzlzZ2hFNEF3dXJubFM1UUdMSHQyNTFKN1hBT2gyZmlYUjdsbGp3MDFoQUZER050WDlxX2xQemFzM0IzaW9QbjN6MU5tZ0p0SkJlSElQaGxPdmVDcjhfNi0zN0lFRkJYTUtMVW9GOGNFdTl2ZUVFR2pFN05HV2NyaENSZk9aX3JXbHNlaURmWnFFSVl6YjB1U3M2dDdQTk1TUXJvRzNqQnRadHF0VE13UkNFc0p4eFlWY0NhWG9PLS10WDdjOXY0Q3lGeXZVTGEzNUNiZz0="
}
Entry Information
- Entry ID: 66038
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000