Row 65994
Content Data
This page contains data entry 65994 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.
My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve this task, returning a regret bounded decision given the distributional shift, its policy is functionally equivalent to a causal model of the environment.
This is a bit idealised, but plausibly you would want an agent that knows how to solve a task given any change to its environment. They have this example where an ideal doctor would know how to diagnose someone given their symptoms, and if they knew some intervention had happened to the patient like they were forced to take a drug, they would also know how to diagnose them.
They also argue that adapting given you know the domain shift is strictly easier than adapting when you don't (and have to figure it out from your inputs). So if the agent has to learn a causal model for this strictly easier adaptation task, it would also have to learn it when you don't know the domain shift.
| Field | Value |
|---|---|
| text | My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve t… |
| label | r/machinelearning |
| dataType | comment |
| communityName | r/MachineLearning |
| datetime | 2024-05-23 |
| username_encoded | Z0FBQUFBQm5Lak1jTjNLM0xMcFIzdG5BT2g1aURxZ0xNdklpVnNyU1NUaGdRZGRaUmlWb0RBd1pqS3Vrazl0dzhqMWJPdTJmOGtqRnJYRXdJMm5oT01OYWRqeko0Sk1pUzNtV2pma2FSRS02ZkZjQUFjYUpBVEU9 |
| url_encoded | Z0FBQUFBQm5Lak9zazlRdnQwZDREbnh6UDI2TU1vM0lkakMtbzBvUFc0WDRMM0p6d1BYb21DQzFhU3Jnc3UzTThRU05QWk9vb09TTWRneFNDM1BtUWEzeVM2X2RjWVR0NjVLX2NpR2tJcm56WlFhVTI4LVRKb1RHbEdVV1plemp5UlBHc1gwMVpESEpGSUFlRDhWdU1ZaUF0ZzItN0hCUkltaW1FSEQ0Rm45M2Y3Tk1CNHg5RHc1ejZuR1doenhxTzB6VUFLOERMNEZuNVRQNm1UNTZySG5Cai11WEVTUW5udz09 |
Raw Record
{
"text": "My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve this task, returning a regret bounded decision given the distributional shift, its policy is functionally equivalent to a causal model of the environment. \n\nThis is a bit idealised, but plausibly you would want an agent that knows how to solve a task given any change to its environment. They have this example where an ideal doctor would know how to diagnose someone given their symptoms, and if they knew some intervention had happened to the patient like they were forced to take a drug, they would also know how to diagnose them. \n\nThey also argue that adapting given you know the domain shift is strictly easier than adapting when you don't (and have to figure it out from your inputs). So if the agent has to learn a causal model for this strictly easier adaptation task, it would also have to learn it when you don't know the domain shift.",
"label": "r/machinelearning",
"dataType": "comment",
"communityName": "r/MachineLearning",
"datetime": "2024-05-23",
"username_encoded": "Z0FBQUFBQm5Lak1jTjNLM0xMcFIzdG5BT2g1aURxZ0xNdklpVnNyU1NUaGdRZGRaUmlWb0RBd1pqS3Vrazl0dzhqMWJPdTJmOGtqRnJYRXdJMm5oT01OYWRqeko0Sk1pUzNtV2pma2FSRS02ZkZjQUFjYUpBVEU9",
"url_encoded": "Z0FBQUFBQm5Lak9zazlRdnQwZDREbnh6UDI2TU1vM0lkakMtbzBvUFc0WDRMM0p6d1BYb21DQzFhU3Jnc3UzTThRU05QWk9vb09TTWRneFNDM1BtUWEzeVM2X2RjWVR0NjVLX2NpR2tJcm56WlFhVTI4LVRKb1RHbEdVV1plemp5UlBHc1gwMVpESEpGSUFlRDhWdU1ZaUF0ZzItN0hCUkltaW1FSEQ0Rm45M2Y3Tk1CNHg5RHc1ejZuR1doenhxTzB6VUFLOERMNEZuNVRQNm1UNTZySG5Cai11WEVTUW5udz09"
}
Entry Information
- Entry ID: 65994
- Repository: Axioma AXP
- Dataset: arrmlet/reddit_dataset_36
- Total Entries: 100,000