LLM-Assisted Repair Gating for Improving Hidden-Test Success in Automated Code-Validation Pipelines

Authors

DOI:

https://doi.org/10.31861/sisiot2026.1.01005

Keywords:

large language models, automated program repair, hidden tests, code-validation pipeline, AI-assisted debugging

Abstract

Large language models (LLMs) are increasingly explored as tools for code generation, debugging, and automated program repair. However, their reliable use in software-validation pipelines remains constrained by the risk of accepting plausible but incorrect patches, especially in environments where final correctness is determined by hidden evaluation tests rather than visible checks alone. This paper investigates an LLM-assisted repair framework integrated into a gated code-validation pipeline. Instead of treating the LLM as an unconstrained code generator, the proposed approach introduces a repair stage into a controlled workflow with sequential validation checks and final hidden-test assessment. The framework was evaluated on a benchmark of fifty programming tasks under two paired conditions: a baseline execution pipeline without automated repair and an LLM+gating configuration in which a candidate repair was applied before validation. The results show a substantial improvement in final correctness: in the baseline condition, only 5 of 50 tasks passed hidden tests, whereas the LLM+gating configuration achieved hidden-test success on all 50 tasks. The observed difference was highly significant under paired statistical analysis, while runtime measurements indicated no meaningful computational overhead compared with the baseline workflow. These findings show that LLM-assisted repair can become an effective component of automated validation pipelines when deployed within an explicit repair-and-gating workflow. The study also provides a reproducible experimental protocol and supplementary research package for independent verification of the reported results.

Downloads

Download data is not yet available.

Author Biographies

  • Roman Zaiats, Yuriy Fedkovych Chernivtsi National University

    Roman Zaiats is a Ph.D. candidate in his final year at the Department of Radio Engineering and Information Security, Chernivtsi National University, Ukraine. His research interests focus on the application of artificial intelligence in data processing.

  • Myroslav Strynadko, Yuriy Fedkovych Chernivtsi National University

    Myroslav Strynadko is a Senior Research Fellow and Associate Professor at the Department of Correlation Optics, Yuriy Fedkovych Chernivtsi National University, Chernivtsi, Ukraine. His research interests include probabilistic computing systems, quantum interferometry, photonic circuits, fiber-optic systems, signal encoding methods, and applied studies in industrial automation, photonics and robotics, data encoding systems, and microprocessor technologies.

  • Halyna Lastivka, Yuriy Fedkovych Chernivtsi National University

    Received BS and MS degrees in Radio Engineering from Yuriy Fedkovych Chernivtsi National University, Ukraine. She received a Ph.D. in solid state electronics from Yuriy Fedkovych Chernivtsi National University. She is currently an associate professor of the Radio Engineering Department of Yuriy Fedkovych Chernivtsi National  University. Her research interests encompass the fields of cybersecurity, radiospectroscopy, and materials science. The primary research activity is focused on the design and development of comprehensive and technical data protection systems, the development of a portable digital multi-pulse nuclear quadrupole resonance spectrometer for analyzing the structure and sensing properties of semiconductors.

References

J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A Survey on Large Language Models for Code Generation,” ACM Transactions on Software Engineering and Methodology, 2024, doi: 10.1145/3747588.

N. Huynh and B. Lin, “Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications,” arXiv, 2025, doi: 10.48550/arXiv.2503.01245.

Z. Rasheed, M. Waseem, K.-K. Kemell, A. Ahmad, M. A. Sami, J. Rasku, K. Systä, and P. Abrahamsson, “Large Language Models for Code Generation: The Practitioners Perspective,” arXiv, 2025, doi: 10.48550/arXiv.2501.16998.

N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica, “LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code,” arXiv, 2024, doi: 10.48550/arXiv.2403.07974.

L. Yang, R. Jin, L. Shi, J. Peng, Y. Chen, and D. Xiong, “ProBench: Benchmarking Large Language Models in Competitive Programming,” arXiv, 2025, doi: 10.48550/arXiv.2502.20868.

F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation,” IEEE Transactions on Software Engineering, vol. 49, pp. 3675–3691, 2023, doi: 10.1109/TSE.2023.3267446.

D. G. Paul, H. Zhu, and I. Bayley, “Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review,” in Proc. 2024 IEEE Int. Conf. on Artificial Intelligence Testing (AITest), 2024, pp. 87–94, doi: 10.1109/AITest62860.2024.00019.

J. Liu, S. Xie, J. Wang, Y. Wei, Y. Ding, and L. Zhang, “Evaluating Language Models for Efficient Code Generation,” arXiv, 2024, doi: 10.48550/arXiv.2408.06450.

J. Zheng, B. Cao, Z. Ma, R. Pan, H. Lin, Y. Lu, X. Han, and L. Sun, “Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models,” arXiv, 2024, doi: 10.48550/arXiv.2407.11470.

S. B. Hossain, N. Jiang, Q. Zhou, X. Li, W.-H. Chiang, Y. Lyu, H. Nguyen, and O. Tripp, “A Deep Dive into Large Language Models for Automated Bug Localization and Repair,” Proceedings of the ACM on Software Engineering, vol. 1, pp. 1471–1493, 2024, doi: 10.1145/3660773.

N. Jiang, K. Liu, T. Lutellier, and L. Tan, “Impact of Code Language Models on Automated Program Repair,” in Proc. 2023 IEEE/ACM 45th Int. Conf. on Software Engineering (ICSE), 2023, pp. 1430–1442, doi: 10.1109/ICSE48619.2023.00125.

C. Xia and L. Zhang, “Automated Program Repair in the Era of Large Pre-trained Language Models,” in Proc. 2023 IEEE/ACM 45th Int. Conf. on Software Engineering (ICSE), 2023, pp. 1482–1494, doi: 10.1109/ICSE48619.2023.00129.

K. Huang, Z. Xu, S. Yang, H. Sun, X. Li, Z. Yan, and Y. Zhang, “A Survey on Automated Program Repair Techniques,” arXiv, 2023, doi: 10.48550/arXiv.2303.18184.

J. Renzullo, P. Reiter, W. Weimer, and S. Forrest, “Automated Program Repair: Emerging Trends Pose and Expose Problems for Benchmarks,” ACM Computing Surveys, vol. 57, 2024, doi: 10.1145/3704997.

M. Monperrus, “Automatic Software Repair,” ACM Computing Surveys, vol. 51, 2018, doi: 10.1145/3105906.

Q. Zhang, C. Fang, Y. Xie, Y. Ma, W. Sun, Y. Yang, and Z. Chen, “A Systematic Literature Review on Large Language Models for Automated Program Repair,” arXiv, 2024, doi: 10.48550/arXiv.2405.01466.

K. H. Levin, N. van Kempen, E. D. Berger, and S. N. Freund, “ChatDBG: Augmenting Debugging with Large Language Models,” Proceedings of the ACM on Software Engineering, 2025, doi: 10.1145/3729355.

Z. Englhardt, R. Li, D. Nissanka, Z. Zhang, G. Narayanswamy, J. Breda, X. Liu, S. N. Patel, and V. Iyer, “Exploring and Characterizing Large Language Models for Embedded System Development and Debugging,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2023, doi: 10.1145/3613905.3650764.

Downloads


Abstract views: 0

Published

2026-06-30

Issue

Section

Articles

How to Cite

[1]
R. Zaiats, M. Strynadko, and H. Lastivka, “LLM-Assisted Repair Gating for Improving Hidden-Test Success in Automated Code-Validation Pipelines”, SISIOT, vol. 4, no. 1, p. 01005, Jun. 2026, doi: 10.31861/sisiot2026.1.01005.

Most read articles by the same author(s)