Row 65770
Content Data
This page contains data entry 65770 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
> I'm interested in understanding these downstream differences in the context of VI, but haven't found any works that explain these claims theoretically instead of empirically.
Researchers like to ask for "theory", but if you're not specific, it's really hard to know what you want or if it's even possible.
Imagine a complex posterior that's approximated by a Gaussian. The M-projection (forward KL) will match the mean and variance. The I-projection (reverse KL) will fit a mode of the posterior. If your true posterior is also Gaussian (i.e., in sparse GPs), the KL you use is much-of-a-muchness. However, if your true posterior looks like a set of delta functions that are sufficiently "far apart," the M-projection would give you a pretty bad fit that assigns probability mass where there shouldn't be anything. Conversely, fitting a mode is the best thing you can do with a Gaussian.
| Field | Value |
|---|---|
| text | > I'm interested in understanding these downstream differences in the context of VI, but haven't found any works that explain these claims theoretically instead of empirically. Researchers like to ask for "theory", but if you're not specific, it's really hard to know what you want or if it's even possible. Imagine a complex posterior that's approximated by a Gaussian. The M-projection (forward KL) will match the mean and variance. The I-projection (reverse KL) will fit a mode of the posterior… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1jZXhaY1hBbzFFb1J2dEI1c2VOaFpYTkFVa0I2ei1WcW9ZNUR1MGJmNGJ0b1pmdFpvRnBfR0Q5anJPd0twVzNvdHJjc3hQLWZtblctenhNeHR2NUtSX1E9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9zSVk0S0hTSkk3cUxfTVp0SnpORFRtWWR3X1F0c3VsTDNPS0FlTTNDT04zUkhlWEFQMk5Mb2J0VzZlVmJKMFVxQ012NEVDLW1kRl9ESkkzME5wbzNkelB3Q3lxN2JiTl9uVm1USnlCcXl1bXVtVXExWU9nZ1E1ZnQ5VmQ1Q0MyLXJjck5BWUpmc2daVEs4VlFaVGNFdHgwYTR6YUxaSDR0Qkd6a3FObnZKODk0VEt4Z2tHMXZYNTlWajNONk1udE80b1paTFFNSXJ6ZnNvUXFXRFR3WjRJRWtPX2Y0UjVJRDNSWHpub09GS1NqWT0= |
Raw Record
{
"text": "> I'm interested in understanding these downstream differences in the context of VI, but haven't found any works that explain these claims theoretically instead of empirically.\n\nResearchers like to ask for \"theory\", but if you're not specific, it's really hard to know what you want or if it's even possible.\n\nImagine a complex posterior that's approximated by a Gaussian. \nThe M-projection (forward KL) will match the mean and variance. The I-projection (reverse KL) will fit a mode of the posterior. If your true posterior is also Gaussian (i.e., in sparse GPs), the KL you use is much-of-a-muchness. However, if your true posterior looks like a set of delta functions that are sufficiently \"far apart,\" the M-projection would give you a pretty bad fit that assigns probability mass where there shouldn't be anything. Conversely, fitting a mode is the best thing you can do with a Gaussian.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1jZXhaY1hBbzFFb1J2dEI1c2VOaFpYTkFVa0I2ei1WcW9ZNUR1MGJmNGJ0b1pmdFpvRnBfR0Q5anJPd0twVzNvdHJjc3hQLWZtblctenhNeHR2NUtSX1E9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9zSVk0S0hTSkk3cUxfTVp0SnpORFRtWWR3X1F0c3VsTDNPS0FlTTNDT04zUkhlWEFQMk5Mb2J0VzZlVmJKMFVxQ012NEVDLW1kRl9ESkkzME5wbzNkelB3Q3lxN2JiTl9uVm1USnlCcXl1bXVtVXExWU9nZ1E1ZnQ5VmQ1Q0MyLXJjck5BWUpmc2daVEs4VlFaVGNFdHgwYTR6YUxaSDR0Qkd6a3FObnZKODk0VEt4Z2tHMXZYNTlWajNONk1udE80b1paTFFNSXJ6ZnNvUXFXRFR3WjRJRWtPX2Y0UjVJRDNSWHpub09GS1NqWT0="
}
Entry Information
- Entry ID: 65770
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000