Row 5448

Row ID: 5448 | Dataset Entry | Axioma AXP Content Repository

Content Data

This page contains data entry 5448 from the Axioma AXP content repository. The structured data below represents the complete record for this entry.

**Paper**: [https://arxiv.org/abs/2404.14662](https://arxiv.org/abs/2404.14662)

**Abstract**:

>A fundamental skill among human developers is the ability to understand and reason about program execution. As an example, a programmer can mentally simulate code execution in natural language to debug and repair code (aka. rubber duck debugging). However, large language models (LLMs) of code are typically trained on the surface textual form of programs, thus may lack a semantic understanding of how programs execute at run-time. To address this issue, we propose **NExT**, a method to teach LLMs to inspect the execution traces of programs (variable states of executed lines) and reason about their run-time behavior through chain-of-thought (CoT) rationales. Specifically, NExT uses self-training to bootstrap a synthetic training set of execution-aware rationales that lead to correct task solutions (e.g., fixed programs) without laborious manual annotation. Experiments on program repair tasks based **on MBPP and HumanEval demonstrate that NExT improves the fix rate of a PaLM 2 model, by 26.1% and 14.3% absolute, respectively**, with significantly improved rationale quality as verified by automated metrics and human raters. Our model can also generalize to scenarios where program traces are absent at test-time.

FieldValue
text **Paper**: [https://arxiv.org/abs/2404.14662](https://arxiv.org/abs/2404.14662) **Abstract**: >A fundamental skill among human developers is the ability to understand and reason about program execution. As an example, a programmer can mentally simulate code execution in natural language to debug and repair code (aka. rubber duck debugging). However, large language models (LLMs) of code are typically trained on the surface textual form of programs, thus may lack a semantic understanding of…
label r/machinelearning
dataType post
communityName r/MachineLearning
datetime 2024-04-30
username_encoded Z0FBQUFBQm5LakwyMkJ6bmtoWkpPVUpjTGFBcUVuRi1nZVUwZTRNY3k1LVdSYm1qek5JMzY3MVc5VnU5NjNTVjAtSG9BNHZYb3ZLVl84MkhBSTlIWkNsV2xJd3M2ZWNMZUE9PQ==
url_encoded Z0FBQUFBQm5Lak9GLV9pcXNVLVV4NkdWN1dwbW1CRGRsOWZkbXlSajhYZmhWWnVQenVFWW5FaHVlWjQ1cUpnRGRWa0R3aU1CN2VnSzd3ampGNVRNLUlqVndGNkJKMl93am50ZGlEYjFrSXNJMF9FazNPb2QxSDFTS0hEa1VEQVpTMTktX3pHeUdQVHhKU2huWUtxd3ZILXJ0VUYwYlQ5TE44dXRVSE1aYjcxSE9aaWZvSU94aDNLcC1ic2ZKREVfSlNEWTNXXzkwR00tdGQ0UXJjcGJETVo4N2pSMFloUkZzZz09

Raw Record

{
  "text": "**Paper**: [https://arxiv.org/abs/2404.14662](https://arxiv.org/abs/2404.14662)\n\n**Abstract**:\n\n>A fundamental skill among human developers is the ability to understand  and reason about program execution. As an example, a programmer can  mentally simulate code execution in natural language to debug and repair  code (aka. rubber duck debugging). However, large language models  (LLMs) of code are typically trained on the surface textual form of  programs, thus may lack a semantic understanding of how programs execute  at run-time. To address this issue, we propose **NExT**, a method to teach  LLMs to inspect the execution traces of programs (variable states of  executed lines) and reason about their run-time behavior through  chain-of-thought (CoT) rationales. Specifically, NExT uses self-training  to bootstrap a synthetic training set of execution-aware rationales  that lead to correct task solutions (e.g., fixed programs) without  laborious manual annotation. Experiments on program repair tasks based  **on MBPP and HumanEval demonstrate that NExT improves the fix rate of a  PaLM 2 model, by 26.1% and 14.3% absolute, respectively**, with  significantly improved rationale quality as verified by automated  metrics and human raters. Our model can also generalize to scenarios  where program traces are absent at test-time.",
  "label": "r/machinelearning",
  "dataType": "post",
  "communityName": "r/MachineLearning",
  "datetime": "2024-04-30",
  "username_encoded": "Z0FBQUFBQm5LakwyMkJ6bmtoWkpPVUpjTGFBcUVuRi1nZVUwZTRNY3k1LVdSYm1qek5JMzY3MVc5VnU5NjNTVjAtSG9BNHZYb3ZLVl84MkhBSTlIWkNsV2xJd3M2ZWNMZUE9PQ==",
  "url_encoded": "Z0FBQUFBQm5Lak9GLV9pcXNVLVV4NkdWN1dwbW1CRGRsOWZkbXlSajhYZmhWWnVQenVFWW5FaHVlWjQ1cUpnRGRWa0R3aU1CN2VnSzd3ampGNVRNLUlqVndGNkJKMl93am50ZGlEYjFrSXNJMF9FazNPb2QxSDFTS0hEa1VEQVpTMTktX3pHeUdQVHhKU2huWUtxd3ZILXJ0VUYwYlQ5TE44dXRVSE1aYjcxSE9aaWZvSU94aDNLcC1ic2ZKREVfSlNEWTNXXzkwR00tdGQ0UXJjcGJETVo4N2pSMFloUkZzZz09"
}

Entry Information