Automating software testing using artificial intelligence and machine learning
DOI:
https://doi.org/10.63688/cognitivatech202653Keywords:
artificial intelligence, machine learning, software testing, test automation, software engineering.Abstract
Introduction: Software test automation is an essential strategy to reduce times, improve coverage and detect defects during development. The incorporation of artificial intelligence (AI) and machine learning (ML) opens up new possibilities for generating, prioritizing, and optimizing test cases. Objective: To evaluate the effectiveness of AI- and ML-assisted software test automation versus a conventional approach in Software Technology Engineering students at the Universidad Autónoma de Nuevo León, Mexico. Methodology: A quantitative, applied, experimental, prospective and comparative study was carried out with 128 students, randomly distributed into a control group (n=64) and an experimental group (n=64). Execution time, compilable and executable tests, coverage of lines and branches, mutation score and detected defects were evaluated. Results: The experimental group reduced the mean resolution time from 44.8 to 36.2 minutes and obtained higher values of line coverage (82.6 % vs. 74.8 %), branch coverage (75.4 % vs. 66.1 %) and mutation score (70.8 % vs. 59.7 %). It also detected a greater number of defects, with statistically significant differences (p<0.001). Conclusions: The incorporation of AI and ML showed advantages in efficiency and effectiveness of automated tests, although their application requires technical supervision and human validation.References
Martins G, Tenório N, Bernardino J. AI-driven software testing: a review. Big Data Cogn Comput. 2026;10(7):233. https://doi.org/10.3390/bdcc10070233 DOI: https://doi.org/10.3390/bdcc10070233
Ajorloo S, Jamarani A, Kashfi M, Haghi Kashani M, Najafizadeh A. A systematic review of machine learning methods in software testing. Appl Soft Comput. 2024;162:111805. https://doi.org/10.1016/j.asoc.2024.111805 DOI: https://doi.org/10.1016/j.asoc.2024.111805
Fontes A, Gay G. The integration of machine learning into automated test generation: a systematic mapping study. Softw Test Verif Reliab. 2023;33(4). https://doi.org/10.1002/stvr.1845 DOI: https://doi.org/10.1002/stvr.1845
Rafi DM, Moses KRK, Petersen K, Mäntylä MV. Benefits and limitations of automated software testing: systematic literature review and practitioner survey. In: 2012 7th International Workshop on Automation of Software Test. 2012. p. 36-42. https://doi.org/10.1109/IWAST.2012.6228988 DOI: https://doi.org/10.1109/IWAST.2012.6228988
Pan R, Bagherzadeh M, Ghaleb TA, Briand L. Test case selection and prioritization using machine learning: a systematic literature review. Empir Softw Eng. 2022;27(2):29. https://doi.org/10.1007/s10664-021-10066-6 DOI: https://doi.org/10.1007/s10664-021-10066-6
Spieker H, Gotlieb A, Marijan D, Mossige M. Reinforcement learning for automatic test case prioritization and selection in continuous integration. In: Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis. 2017. p. 12-22. https://doi.org/10.1145/3092703.3092709 DOI: https://doi.org/10.1145/3092703.3092709
Bertolino A, Guerriero A, Miranda B, Pietrantuono R, Russo S. Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integration. In: Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 2020. p. 1261-1272. https://doi.org/10.1145/3377811.3380369 DOI: https://doi.org/10.1145/3377811.3380369
Fraser G, Arcuri A. EvoSuite: automatic test suite generation for object-oriented software. In: Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering. 2011. p. 416-419. https://doi.org/10.1145/2025113.2025179 DOI: https://doi.org/10.1145/2025113.2025179
Lukasczyk S, Fraser G. Pynguin: automated unit test generation for Python. In: 44th International Conference on Software Engineering Companion. 2022. p. 168-172. https://doi.org/10.1145/3510454.3516829 DOI: https://doi.org/10.1145/3510454.3516829
Zhang YW, Jin Z, Wang ZJ, Xing Y, Li G. SAGA: summarization-guided assert statement generation. J Comput Sci Technol. 2025;40(1):138-157. https://doi.org/10.1007/s11390-023-2878-6 DOI: https://doi.org/10.1007/s11390-023-2878-6
Dinella E, Ryan G, Mytkowicz T, Lahiri SK. TOGA: a neural method for test oracle generation. In: Proceedings of the 44th International Conference on Software Engineering. 2022. p. 2130-2141. https://doi.org/10.1145/3510003.3510141 DOI: https://doi.org/10.1145/3510003.3510141
Alagarsamy S, Tantithamthavorn C, Aleti A. A3Test: assertion-augmented automated test case generation. Inf Softw Technol. 2024;176:107565. https://doi.org/10.1016/j.infsof.2024.107565 DOI: https://doi.org/10.1016/j.infsof.2024.107565
Lemieux C, Inala JP, Lahiri SK, Sen S. CodaMOSA: escaping coverage plateaus in test generation with pre-trained large language models. In: 2023 IEEE/ACM 45th International Conference on Software Engineering. 2023. p. 919-931. https://doi.org/10.1109/ICSE48619.2023.00085 DOI: https://doi.org/10.1109/ICSE48619.2023.00085
Wang J, Huang Y, Chen C, Liu Z, Wang S, Wang Q. Software testing with large language models: survey, landscape, and vision. IEEE Trans Softw Eng. 2024;50(4):911-936. https://doi.org/10.1109/TSE.2024.3368208 DOI: https://doi.org/10.1109/TSE.2024.3368208
Li Y, Liu P, Wang H, Chu J, Wong WE. Evaluating large language models for software testing. Comput Stand Interfaces. 2025;93:103942. https://doi.org/10.1016/j.csi.2024.103942 DOI: https://doi.org/10.1016/j.csi.2024.103942
Tasarsu M, Tokmak AV, Catal C. Test case generation using large language models: a systematic literature review. Cluster Comput. 2026;29:227. https://doi.org/10.1007/s10586-026-06021-z DOI: https://doi.org/10.1007/s10586-026-06021-z
Moradi Dakhel A, Nikanjam A, Majdinasab V, Khomh F, Desmarais MC. Effective test generation using pre-trained large language models and mutation testing. Inf Softw Technol. 2024;171:107468. https://doi.org/10.1016/j.infsof.2024.107468 DOI: https://doi.org/10.1016/j.infsof.2024.107468
Rehan S, Al-Bander B, Al-Said Ahmad A. Harnessing large language models for automated software testing: a leap towards scalable test case generation. Electronics. 2025;14(7):1463. https://doi.org/10.3390/electronics14071463 DOI: https://doi.org/10.3390/electronics14071463
Yang L, Yang C, Gao S, Wang W, Wang B, Zhu Q, et al. On the evaluation of large language models in unit test generation. In: 2024 39th IEEE/ACM International Conference on Automated Software Engineering. 2024. p. 1607-1619. https://doi.org/10.1145/3691620.3695529 DOI: https://doi.org/10.1145/3691620.3695529
Santos R, Santos I, Magalhães CVC, Santos RS. Are we testing or being tested? Exploring the practical applications of large language models in software testing. In: 2024 IEEE Conference on Software Testing, Verification and Validation. 2024. p. 353-360. https://doi.org/10.1109/ICST60714.2024.00039 DOI: https://doi.org/10.1109/ICST60714.2024.00039
Lukasczyk S, Kroiß F, Fraser G. An empirical study of automated unit test generation for Python. Empir Softw Eng. 2023;28:36. https://doi.org/10.1007/s10664-022-10248-w DOI: https://doi.org/10.1007/s10664-022-10248-w
Lukasczyk S, Kroiß F, Fraser G. Automated unit test generation for Python. In: Search-Based Software Engineering. Lecture Notes in Computer Science. 2020;12420:9-24. https://doi.org/10.1007/978-3-030-59762-7_2 DOI: https://doi.org/10.1007/978-3-030-59762-7_2
Tufano M, Drain D, Svyatkovskiy A, Sundaresan N. Generating accurate assert statements for unit test cases using pretrained transformers. In: IEEE/ACM International Conference on Automation of Software Test. 2022. p. 54-64. https://doi.org/10.1145/3524481.3527220 DOI: https://doi.org/10.1145/3524481.3527220
Liu XJ, Yu P, Ma XX. An empirical study on automated test generation tools for Java: effectiveness and challenges. J Comput Sci Technol. 2024;39(3):715-736. https://doi.org/10.1007/s11390-023-1935-5 DOI: https://doi.org/10.1007/s11390-023-1935-5
Chen YT, Gopinath R, Tadakamalla A, Ernst MD, Holmes R, Fraser G, et al. Revisiting the relationship between fault detection, test adequacy criteria, and test set size. In: Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering. 2020. p. 237-249. https://doi.org/10.1145/3324884.3416667 DOI: https://doi.org/10.1145/3324884.3416667
Garousi V, Mäntylä MV. When and what to automate in software testing? A multi-vocal literature review. Inf Softw Technol. 2016;76:92-117. https://doi.org/10.1016/j.infsof.2016.04.015 DOI: https://doi.org/10.1016/j.infsof.2016.04.015
Published
Issue
Section
License
Copyright (c) 2026 Carlos Hernández Salas, Carmen Carolina Ortega Hernández (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
The article is distributed under the Creative Commons Attribution 4.0 License. Unless otherwise stated, associated published material is distributed under the same licence.
