Row 65994

Row ID: 65994 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 65994 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve this task, returning a regret bounded decision given the distributional shift, its policy is functionally equivalent to a causal model of the environment.

This is a bit idealised, but plausibly you would want an agent that knows how to solve a task given any change to its environment. They have this example where an ideal doctor would know how to diagnose someone given their symptoms, and if they knew some intervention had happened to the patient like they were forced to take a drug, they would also know how to diagnose them.

They also argue that adapting given you know the domain shift is strictly easier than adapting when you don't (and have to figure it out from your inputs). So if the agent has to learn a causal model for this strictly easier adaptation task, it would also have to learn it when you don't know the domain shift.

FieldValue
text My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve t…
label r/machinelearning
dataType comment
communityName r/MachineLearning
datetime 2024-05-23
username_encoded Z0FBQUFBQm5Lak1jTjNLM0xMcFIzdG5BT2g1aURxZ0xNdklpVnNyU1NUaGdRZGRaUmlWb0RBd1pqS3Vrazl0dzhqMWJPdTJmOGtqRnJYRXdJMm5oT01OYWRqeko0Sk1pUzNtV2pma2FSRS02ZkZjQUFjYUpBVEU9
url_encoded Z0FBQUFBQm5Lak9zazlRdnQwZDREbnh6UDI2TU1vM0lkakMtbzBvUFc0WDRMM0p6d1BYb21DQzFhU3Jnc3UzTThRU05QWk9vb09TTWRneFNDM1BtUWEzeVM2X2RjWVR0NjVLX2NpR2tJcm56WlFhVTI4LVRKb1RHbEdVV1plemp5UlBHc1gwMVpESEpGSUFlRDhWdU1ZaUF0ZzItN0hCUkltaW1FSEQ0Rm45M2Y3Tk1CNHg5RHc1ejZuR1doenhxTzB6VUFLOERMNEZuNVRQNm1UNTZySG5Cai11WEVTUW5udz09

Raw Record

{
  "text": "My understanding is they then have something like a thought experiment ('interpretation' section) where you have an agent tasked with adapting to know domain shifts, ie the agents policy conditions on the domain shift. I imagine this as something like an LLM that you query with 'given this observation and this shift on the environment, like I force X=x, what is the optional policy?'. And it doesn't have to be optimal, just satisfy a regret bound. Anyway, they show that any agent that can solve this task, returning a regret bounded decision given the distributional shift, its policy is functionally equivalent to a causal model of the environment. \n\nThis is a bit idealised, but plausibly you would want an agent that knows how to solve a task given any change to its environment. They have this example where an ideal doctor would know how to diagnose someone given their symptoms, and if they knew some intervention had happened to the patient like they were forced to take a drug, they would also know how to diagnose them. \n\nThey also argue that adapting given you know the domain shift is strictly easier than adapting when you don't (and have to figure it out from your inputs). So if the agent has to learn a causal model for this strictly easier adaptation task, it would also have to learn it when you don't know the domain shift.",
  "label": "r/machinelearning",
  "dataType": "comment",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-23",
  "username_encoded": "Z0FBQUFBQm5Lak1jTjNLM0xMcFIzdG5BT2g1aURxZ0xNdklpVnNyU1NUaGdRZGRaUmlWb0RBd1pqS3Vrazl0dzhqMWJPdTJmOGtqRnJYRXdJMm5oT01OYWRqeko0Sk1pUzNtV2pma2FSRS02ZkZjQUFjYUpBVEU9",
  "url_encoded": "Z0FBQUFBQm5Lak9zazlRdnQwZDREbnh6UDI2TU1vM0lkakMtbzBvUFc0WDRMM0p6d1BYb21DQzFhU3Jnc3UzTThRU05QWk9vb09TTWRneFNDM1BtUWEzeVM2X2RjWVR0NjVLX2NpR2tJcm56WlFhVTI4LVRKb1RHbEdVV1plemp5UlBHc1gwMVpESEpGSUFlRDhWdU1ZaUF0ZzItN0hCUkltaW1FSEQ0Rm45M2Y3Tk1CNHg5RHc1ejZuR1doenhxTzB6VUFLOERMNEZuNVRQNm1UNTZySG5Cai11WEVTUW5udz09"
}

Entry Information