Row 9003

Row ID: 9003 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 9003 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**Website**: [https://agentclinic.github.io/](https://agentclinic.github.io/) **Arxiv**: [https://arxiv.org/pdf/2405.07960](https://arxiv.org/pdf/2405.07960)

**TLDR:** AgentClinic turns static medical QA problems into agents in a clinical environment (doctor, patient, medical devices) in order to present a more clinically relevant challenge for medical language models.

**Abstract:** Diagnosing and managing a patient is a complex, sequential decision making process that requires physicians to obtain information---such as which tests to perform---and to act upon it. Recent advances in artificial intelligence (AI) and large language models (LLMs) promise to profoundly impact clinical care. However, current evaluation schemes overrely on static medical question-answering benchmarks, falling short on interactive decision-making that is required in real-life clinical work. Here, we present AgentClinic: a multimodal benchmark to evaluate LLMs in their ability to operate as agents in simulated clinical environments. In our benchmark, the doctor agent must uncover the patient's diagnosis through dialogue and active data collection. We present two open benchmarks: a multimodal image and dialogue environment, AgentClinic-NEJM, and a dialogue-only environment, AgentClinic-MedQA. Agents in AgentClinic-MedQA are grounded in cases from the US Medical Licensing Exam\~(USMLE) and AgentClinic-NEJM are grounded in multimodal New England Journal of Medicine (NEJM) case challenges. We embed cognitive and implicit biases both in patient and doctor agents to emulate realistic interactions between biased agents. We find that introducing bias leads to large reductions in diagnostic accuracy of the doctor agents, as well as reduced compliance, confidence, and follow-up consultation willingness in patient agents. Evaluating a suite of state-of-the-art LLMs, we find that several models that excel in benchmarks like MedQA are performing poorly in AgentClinic-MedQA. We find that the LLM used in the patient agent is an important factor for performance in the AgentClinic benchmark. We show that both having limited interactions as well as too many interaction reduces diagnostic accuracy in doctor agents.

FieldValue
text **Website**: [https://agentclinic.github.io/](https://agentclinic.github.io/) **Arxiv**: [https://arxiv.org/pdf/2405.07960](https://arxiv.org/pdf/2405.07960) **TLDR:** AgentClinic turns static medical QA problems into agents in a clinical environment (doctor, patient, medical devices) in order to present a more clinically relevant challenge for medical language models. **Abstract:** Diagnosing and managing a patient is a complex, sequential decision making process that requires physicians to…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-05-20
username_encoded Z0FBQUFBQm5Lakw0bHg3TVFqd1drRThodDFDZnBaSnZPNng3Q1B6MVpaVjVKcEw2Z3B5SUlIVXhJWUZRY3dxSHZCNHNQRzFNbTBTc29HNi1jUlVGeGM2SG5VQUVKRjJUM2c9PQ==
url_encoded Z0FBQUFBQm5Lak9JV0o1M1FWQkMxMWRaNG14SG45S044SHpTVVNqbnl2TFVDajd3c29YaUY1Tk91dm01ZHVjeUZWb3YyU3dUaER4RllEZHcwYVZob2FtY0ZFN0pOMUY1elBhMWN4X3l3RkxZUWw0RndIWlZLVUI4c0Qyeng1UkxPVXJES0tLQVRQTTI4QWMyMzhWTm9IdGtObUpHNUZJS2xIMlZkNTEzQnNnZ1NjV19HamtnamxMb1JOUm9TYldVRWZJcnhsdGlqYlU3MzU2M3dISWNmUUFiLTQyLXkyZVdOdz09

Raw Record

{
  "text": "**Website**: [https://agentclinic.github.io/](https://agentclinic.github.io/)  \n**Arxiv**: [https://arxiv.org/pdf/2405.07960](https://arxiv.org/pdf/2405.07960)\n\n**TLDR:** AgentClinic turns static medical QA problems into agents in a clinical environment (doctor, patient, medical devices) in order to present a more clinically relevant challenge for medical language models.\n\n**Abstract:** Diagnosing and managing a patient is a complex, sequential decision making process that requires physicians to obtain information---such as which tests to perform---and to act upon it. Recent advances in artificial intelligence (AI) and large language models (LLMs) promise to profoundly impact clinical care. However, current evaluation schemes overrely on static medical question-answering benchmarks, falling short on interactive decision-making that is required in real-life clinical work. Here, we present AgentClinic: a multimodal benchmark to evaluate LLMs in their ability to operate as agents in simulated clinical environments. In our benchmark, the doctor agent must uncover the patient's diagnosis through dialogue and active data collection. We present two open benchmarks: a multimodal image and dialogue environment, AgentClinic-NEJM, and a dialogue-only environment, AgentClinic-MedQA. Agents in AgentClinic-MedQA are grounded in cases from the US Medical Licensing Exam\\~(USMLE) and AgentClinic-NEJM are grounded in multimodal New England Journal of Medicine (NEJM) case challenges. We embed cognitive and implicit biases both in patient and doctor agents to emulate realistic interactions between biased agents. We find that introducing bias leads to large reductions in diagnostic accuracy of the doctor agents, as well as reduced compliance, confidence, and follow-up consultation willingness in patient agents. Evaluating a suite of state-of-the-art LLMs, we find that several models that excel in benchmarks like MedQA are performing poorly in AgentClinic-MedQA. We find that the LLM used in the patient agent is an important factor for performance in the AgentClinic benchmark. We show that both having limited interactions as well as too many interaction reduces diagnostic accuracy in doctor agents.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-05-20",
  "username_encoded": "Z0FBQUFBQm5Lakw0bHg3TVFqd1drRThodDFDZnBaSnZPNng3Q1B6MVpaVjVKcEw2Z3B5SUlIVXhJWUZRY3dxSHZCNHNQRzFNbTBTc29HNi1jUlVGeGM2SG5VQUVKRjJUM2c9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9JV0o1M1FWQkMxMWRaNG14SG45S044SHpTVVNqbnl2TFVDajd3c29YaUY1Tk91dm01ZHVjeUZWb3YyU3dUaER4RllEZHcwYVZob2FtY0ZFN0pOMUY1elBhMWN4X3l3RkxZUWw0RndIWlZLVUI4c0Qyeng1UkxPVXJES0tLQVRQTTI4QWMyMzhWTm9IdGtObUpHNUZJS2xIMlZkNTEzQnNnZ1NjV19HamtnamxMb1JOUm9TYldVRWZJcnhsdGlqYlU3MzU2M3dISWNmUUFiLTQyLXkyZVdOdz09"
}

Entry Information