DOI: 10.3390/ai7100392 ISSN: 2673-2688

Inducing Code Change in LLM Self-Debugging: Effects on Program Repair

Jindae Kim

Execution-guided self-debugging can return unchanged code when revision is requested, and even a changed program may still fail to repair the observed failure. A final pass/fail result alone does not distinguish these different outcomes. We study them using 9790 Python programs generated by six open-weight models on three benchmarks. The models generally classified programs as PASS or FAIL from benchmark evaluation records and selected PRESERVE or REVISE in separate judgment tasks, but separate revision requests changed only 57.39% of 4168 initially failing programs and repaired 16.34% under the same benchmark test suite. Because many revision requests returned unchanged code, we tested a minimal intervention that simply added an instruction stating that the returned code must change. This instruction increased the rate of structural code changes, measured as differences between the labeled ASTs of the initial and revised programs, by 7.92–31.86 percentage points and produced modest repair gains of 0.82–1.87 points across the tested failure-information settings. Most code changes nevertheless remained unsuccessful, and concrete execution feedback provided a substantial additional benefit: providing the first reported failure instead of only a failure status increased repair by 12.93–13.43 points across four revision prompts, while providing all reported failures yielded a further benefit on BigCodeBench-Instruct when multiple failures were reported. Repeating revision after reevaluation recovered additional repairs, but the incremental gain decreased sharply over successive turns. Prompts that explicitly required a code change generally retained a small repair advantage across repeated turns even as the benefit of each additional turn diminished. Overall, the results identify complementary ways to improve LLM self-debugging: a simple code-change instruction can reduce unchanged revisions and provide modest repair gains, while concrete execution feedback helps direct changed code toward repair, and repeated reevaluation provides further but diminishing opportunities for repair.