Row 64711
Content Data
This page contains data entry 64711 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
Hi all,
I'm working on variational inference methods, mainly in the context of BNNs. Using the reverse (exclusive) KL as the variational objective is the common approach, though lately I stumbled upon some interesting works that use the forward (inclusive) KL as an objective instead, e.g \[1\]\[2\]\[3\]. Also in the context of VI for GPs both divergence measures have been used, see e.g \[4\].
While I'm familiar with the well-known difference between the objectives that the reverse KL is 'mode-seeking' and the forward KL is 'mode covering', I see some of these works making claims about downstream differences of these VI objectives such as (paraphrasing here) *"the reverse KL underestimates predictive variance" \[4\]* and *"the forward KL is useful for applications benefiting from conservative uncertainty quantification" \[3\]*.
I'm interested in understanding these downstream differences in the context of VI, but haven't found any works that explain these claims theoretically instead of empirically. Anyone who can point me in the right direction or have a go at explaining this?
Cheers
\[1\] Naesseth, Christian, Fredrik Lindsten, and David Blei. "Markovian score climbing: Variational inference with KL (p|| q)." *Advances in Neural Information Processing Systems* 33 (2020): 15499-15510.
\[2\] Zhang, L., Blei, D. M., & Naesseth, C. A. (2022). Transport score climbing: Variational inference using forward KL and adaptive neural transport. *arXiv preprint arXiv:2202.01841*.
\[3\] McNamara, D., Loper, J., & Regier, J. (2024, April). Sequential Monte Carlo for Inclusive KL Minimization in Amortized Variational Inference. In *International Conference on Artificial Intelligence and Statistics* (pp. 4312-4320). PMLR.
\[4\] Bauer, M., Van der Wilk, M., & Rasmussen, C. E. (2016). Understanding probabilistic sparse Gaussian process approximations. *Advances in neural information processing systems*, *29*.
| Field | Value |
|---|---|
| text | Hi all, I'm working on variational inference methods, mainly in the context of BNNs. Using the reverse (exclusive) KL as the variational objective is the common approach, though lately I stumbled upon some interesting works that use the forward (inclusive) KL as an objective instead, e.g \[1\]\[2\]\[3\]. Also in the context of VI for GPs both divergence measures have been used, see e.g \[4\]. While I'm familiar with the well-known difference between the objectives that the reverse KL is 'mode-… |
| label | r/machinelearning |
| dataType | post |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1iclRKSmpDaEV1TEZzSUhQa0lSR0FsQ0FoRHJEZ0d3alFnektQOTNhSXNYZWtYR2IwMUEzRmg2bGJydzdSU2ZpX2tqeF82UXR5elBmejcxTUJxNFJEanc9PQ== |
| url_encoded | Z0FBQUFBQm5Lak9yVnRHYWhNc2JQTmlkT2d2SVEtcFRJakVOSGlRLWVJcXRLTTQtSllkdTBZLVhjX1Q1MDFzblMxLWE4TTRUc3JfSEFxQXJHRTlWQW9kLVVHSVJEY1docTFNMklOa3hoNVY2U0VjSHQxUkdTbzJlc2RLV21IQVhHQ0ZRQjlCTkFLaGItNVljZS1hOEFJYmtNRUFmSlFxZjRrai03WnNHdHBHXzE0TmtURzFqUXBUa0pGZ0U3STZITF9meEgwdHdfWGVDVUFVc19BbTdUejRrTklMSGEzV2tjQT09 |
Raw Record
{
"text": "Hi all,\n\nI'm working on variational inference methods, mainly in the context of BNNs. Using the reverse (exclusive) KL as the variational objective is the common approach, though lately I stumbled upon some interesting works that use the forward (inclusive) KL as an objective instead, e.g \\[1\\]\\[2\\]\\[3\\]. Also in the context of VI for GPs both divergence measures have been used, see e.g \\[4\\].\n\nWhile I'm familiar with the well-known difference between the objectives that the reverse KL is 'mode-seeking' and the forward KL is 'mode covering', I see some of these works making claims about downstream differences of these VI objectives such as (paraphrasing here) *\"the reverse KL underestimates predictive variance\" \\[4\\]* and *\"the forward KL is useful for applications benefiting from conservative uncertainty quantification\" \\[3\\]*.\n\nI'm interested in understanding these downstream differences in the context of VI, but haven't found any works that explain these claims theoretically instead of empirically. Anyone who can point me in the right direction or have a go at explaining this?\n\nCheers\n\n\\[1\\] Naesseth, Christian, Fredrik Lindsten, and David Blei. \"Markovian score climbing: Variational inference with KL (p|| q).\" *Advances in Neural Information Processing Systems* 33 (2020): 15499-15510.\n\n\\[2\\] Zhang, L., Blei, D. M., & Naesseth, C. A. (2022). Transport score climbing: Variational inference using forward KL and adaptive neural transport. *arXiv preprint arXiv:2202.01841*.\n\n\\[3\\] McNamara, D., Loper, J., & Regier, J. (2024, April). Sequential Monte Carlo for Inclusive KL Minimization in Amortized Variational Inference. In *International Conference on Artificial Intelligence and Statistics* (pp. 4312-4320). PMLR.\n\n\\[4\\] Bauer, M., Van der Wilk, M., & Rasmussen, C. E. (2016). Understanding probabilistic sparse Gaussian process approximations. *Advances in neural information processing systems*, *29*.",
"label": "r/machinelearning",
"dataType": "post",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1iclRKSmpDaEV1TEZzSUhQa0lSR0FsQ0FoRHJEZ0d3alFnektQOTNhSXNYZWtYR2IwMUEzRmg2bGJydzdSU2ZpX2tqeF82UXR5elBmejcxTUJxNFJEanc9PQ==",
"url_encoded": "Z0FBQUFBQm5Lak9yVnRHYWhNc2JQTmlkT2d2SVEtcFRJakVOSGlRLWVJcXRLTTQtSllkdTBZLVhjX1Q1MDFzblMxLWE4TTRUc3JfSEFxQXJHRTlWQW9kLVVHSVJEY1docTFNMklOa3hoNVY2U0VjSHQxUkdTbzJlc2RLV21IQVhHQ0ZRQjlCTkFLaGItNVljZS1hOEFJYmtNRUFmSlFxZjRrai03WnNHdHBHXzE0TmtURzFqUXBUa0pGZ0U3STZITF9meEgwdHdfWGVDVUFVc19BbTdUejRrTklMSGEzV2tjQT09"
}
Entry Information
- Entry ID: 64711
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000