DOI: 10.3390/a19090805 ISSN: 1999-4893

EvoSort: An Audit Protocol for LLM-Driven Program Search—Correctness, Ablation, and the Cost of Too Few Seeds

Karrar Maher Khudhair, Bareq Maher Khudhair

EvoSort evolves compiled C++ sorting routines. A language model proposes mutations of a baseline sorter, a per-context upper confidence bound (UCB1) bandit picks which operator to try, and a four-gate harness decides whether a candidate may be timed at all. This paper reports what happened when that system was audited at ten seeds per condition. Two conclusions drawn by an earlier version of this work, one from three seeds and one from a single seed, failed in opposite directions. Conditioning operator selection on the input distribution reaches 2.43× over std::sort (standard deviation, SD 1.18) against 2.98× (SD 1.00) for one global UCB1 table shared by every island, losing seven of ten paired seeds, so the conditioning buys nothing we can measure. Generalisation, previously reported as a collapse on the evidence of one champion, holds up across ten: they average 3.51× out of sample, spanning 1.50× to 8.39×, none is slower than std::sort at ten million keys, and six of the ten beat a plain least-significant-digit (LSD) radix sort. Whether radix beats a champion on one of thirty cells or on twenty-four is settled by the seed alone. Correctness held throughout, with zero failures in 420 measured cells covering eleven evolved programs at up to 100 times the training size. Hand-written baselines lose ground at much the same rate across a tenfold increase in input size, so that decline is not evidence of overfitting.