LLM-Assisted Repair Gating for Improving Hidden-Test Success in Automated Code-Validation Pipelines
DOI:
https://doi.org/10.31861/sisiot2026.1.01005Keywords:
large language models, automated program repair, hidden tests, code-validation pipeline, AI-assisted debuggingAbstract
Large language models (LLMs) are increasingly explored as tools for code generation, debugging, and automated program repair. However, their reliable use in software-validation pipelines remains constrained by the risk of accepting plausible but incorrect patches, especially in environments where final correctness is determined by hidden evaluation tests rather than visible checks alone. This paper investigates an LLM-assisted repair framework integrated into a gated code-validation pipeline. Instead of treating the LLM as an unconstrained code generator, the proposed approach introduces a repair stage into a controlled workflow with sequential validation checks and final hidden-test assessment. The framework was evaluated on a benchmark of fifty programming tasks under two paired conditions: a baseline execution pipeline without automated repair and an LLM+gating configuration in which a candidate repair was applied before validation. The results show a substantial improvement in final correctness: in the baseline condition, only 5 of 50 tasks passed hidden tests, whereas the LLM+gating configuration achieved hidden-test success on all 50 tasks. The observed difference was highly significant under paired statistical analysis, while runtime measurements indicated no meaningful computational overhead compared with the baseline workflow. These findings show that LLM-assisted repair can become an effective component of automated validation pipelines when deployed within an explicit repair-and-gating workflow. The study also provides a reproducible experimental protocol and supplementary research package for independent verification of the reported results.
Downloads
References
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A Survey on Large Language Models for Code Generation,” ACM Transactions on Software Engineering and Methodology, 2024, doi: 10.1145/3747588.
N. Huynh and B. Lin, “Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications,” arXiv, 2025, doi: 10.48550/arXiv.2503.01245.
Z. Rasheed, M. Waseem, K.-K. Kemell, A. Ahmad, M. A. Sami, J. Rasku, K. Systä, and P. Abrahamsson, “Large Language Models for Code Generation: The Practitioners Perspective,” arXiv, 2025, doi: 10.48550/arXiv.2501.16998.
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica, “LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code,” arXiv, 2024, doi: 10.48550/arXiv.2403.07974.
L. Yang, R. Jin, L. Shi, J. Peng, Y. Chen, and D. Xiong, “ProBench: Benchmarking Large Language Models in Competitive Programming,” arXiv, 2025, doi: 10.48550/arXiv.2502.20868.
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation,” IEEE Transactions on Software Engineering, vol. 49, pp. 3675–3691, 2023, doi: 10.1109/TSE.2023.3267446.
D. G. Paul, H. Zhu, and I. Bayley, “Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review,” in Proc. 2024 IEEE Int. Conf. on Artificial Intelligence Testing (AITest), 2024, pp. 87–94, doi: 10.1109/AITest62860.2024.00019.
J. Liu, S. Xie, J. Wang, Y. Wei, Y. Ding, and L. Zhang, “Evaluating Language Models for Efficient Code Generation,” arXiv, 2024, doi: 10.48550/arXiv.2408.06450.
J. Zheng, B. Cao, Z. Ma, R. Pan, H. Lin, Y. Lu, X. Han, and L. Sun, “Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models,” arXiv, 2024, doi: 10.48550/arXiv.2407.11470.
S. B. Hossain, N. Jiang, Q. Zhou, X. Li, W.-H. Chiang, Y. Lyu, H. Nguyen, and O. Tripp, “A Deep Dive into Large Language Models for Automated Bug Localization and Repair,” Proceedings of the ACM on Software Engineering, vol. 1, pp. 1471–1493, 2024, doi: 10.1145/3660773.
N. Jiang, K. Liu, T. Lutellier, and L. Tan, “Impact of Code Language Models on Automated Program Repair,” in Proc. 2023 IEEE/ACM 45th Int. Conf. on Software Engineering (ICSE), 2023, pp. 1430–1442, doi: 10.1109/ICSE48619.2023.00125.
C. Xia and L. Zhang, “Automated Program Repair in the Era of Large Pre-trained Language Models,” in Proc. 2023 IEEE/ACM 45th Int. Conf. on Software Engineering (ICSE), 2023, pp. 1482–1494, doi: 10.1109/ICSE48619.2023.00129.
K. Huang, Z. Xu, S. Yang, H. Sun, X. Li, Z. Yan, and Y. Zhang, “A Survey on Automated Program Repair Techniques,” arXiv, 2023, doi: 10.48550/arXiv.2303.18184.
J. Renzullo, P. Reiter, W. Weimer, and S. Forrest, “Automated Program Repair: Emerging Trends Pose and Expose Problems for Benchmarks,” ACM Computing Surveys, vol. 57, 2024, doi: 10.1145/3704997.
M. Monperrus, “Automatic Software Repair,” ACM Computing Surveys, vol. 51, 2018, doi: 10.1145/3105906.
Q. Zhang, C. Fang, Y. Xie, Y. Ma, W. Sun, Y. Yang, and Z. Chen, “A Systematic Literature Review on Large Language Models for Automated Program Repair,” arXiv, 2024, doi: 10.48550/arXiv.2405.01466.
K. H. Levin, N. van Kempen, E. D. Berger, and S. N. Freund, “ChatDBG: Augmenting Debugging with Large Language Models,” Proceedings of the ACM on Software Engineering, 2025, doi: 10.1145/3729355.
Z. Englhardt, R. Li, D. Nissanka, Z. Zhang, G. Narayanswamy, J. Breda, X. Liu, S. N. Patel, and V. Iyer, “Exploring and Characterizing Large Language Models for Embedded System Development and Debugging,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2023, doi: 10.1145/3613905.3650764.
Published
Issue
Section
License
Copyright (c) 2026 Security of Infocommunication Systems and Internet of Things

This work is licensed under a Creative Commons Attribution 4.0 International License.









