Metamorphic testing and AI agents in ensuring the reliability of corporate software

Main Article Content

Mokhammed Ali Patvari

Abstract

The article is devoted to studying the importance of AI agents in ensuring the reliability of corporate software. The purpose of the article is to substantiate the role of metamorphic testing and AI agents in ensuring the reliability, semantic consistency and security of corporate software based on artificial intelligence. The research process used methods of analysis and synthesis to generalize scientific approaches to testing AI systems; the comparative method was used to compare traditional QA, Metamorphic Testing and the agent approach. The results of the study showed that the reliability of corporate software based on AI cannot be ensured only by traditional QA approaches, since LLM, RAG systems, chat bots, recommendation services and AI agents operate in a non-deterministic mode and can change the answer depending on the formulation of the query. It has been established that Metamorphic Testing is an appropriate method for testing such systems, since it allows us to evaluate not a single result, but the constancy of the relationship between responses after a controlled transformation of the input data: paraphrasing, changing the format, rating scale, word order, or adding a small amount of noise. It has been proven that if the content of the query does not change, then the main fact, conclusion, class, recommendation, or logic of the response should remain stable, and their significant change indicates semantic inconsistency, the risk of hallucination, or weak robustness of the model. On this basis, the author's concept of AI-Agentic Testing is proposed, in which the AI-agent automates the generation of follow-up inputs, checking metamorphic relations, detecting unstable responses, hallucinations, logical contradictions, and potential security breaches in corporate AI systems. This approach moves QA from static test scenarios to dynamic control of semantic consistency, reliability, and security of AI solutions in a corporate environment. Its practical value lies in the fact that it can be used as a basis for testing LLM, RAG systems, AI agents and enterprise AI workflows before their integration into critical business processes.

Downloads

Download data is not yet available.

| Abstract views: 10 | PDF Downloads: 3 |

Article Details

How to Cite
Patvari, M. A. (2026). Metamorphic testing and AI agents in ensuring the reliability of corporate software. Global Prosperity, 6(3). https://doi.org/10.66556/2787-9364.3-6.patvari-m
Section
Articles

References

Asaftei, G.M., Roberts, R., Sticha, A. (2026, March 25). State of AI trust in 2026: Shifting to the agentic era. McKinsey & Company. URL: https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-forward/state-of-ai-trust-in-2026-shifting-to-the-agentic-era

Autio, C., Barrett, M., Newman, J., Nunn, J., Oprea, A., Pinelis, L., Rice, A., Tabassi, E., & Vassilev, A. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

Bengio Y. International AI safety report 2026. International AI Safety Report. (2026). URL: https://internationalaisafetyreport.org/sites/default/files/2026-02/international-ai-safety-report-2026_1.pdf

Cho, S., Terragni, V., Kang, S., Ahmed, T., & Yoo, S. (2025). Metamorphic testing of large language models for natural language processing. arXiv. https://arxiv.org/pdf/2511.02108

Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 150–158. https://aclanthology.org/2024.eacl-demo.16.pdf

Gil, Y., Perrault R. (2026). The 2026 AI Index report. Stanford Institute for Human-Centered Artificial Intelligence. URL: https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf

Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv. https://arxiv.org/pdf/2509.04664

Kanstren T. (2020, June 23). Metamorphic testing of machine-learning based systems. Medium. https://medium.com/data-science/metamorphic-testing-of-machine-learning-based-systems-e1fe13baf048

Khirbat M., Ren Y., Castells P., and Sanderson M. (2024). Metamorphic evaluation of ChatGPT as a recommender system. Conference acronym ’XX, June 03–05, 2024. URL: https://arxiv.org/pdf/2411.12121

Manino, E., Rozanova, J., & Morante, R. (2022). Systematicity, Compositionality and Transitivity of Deep NLP Models: a Metamorphic Testing Perspective. Findings of the Association for Computational Linguistics: ACL 2022, pages 2355 – 2366. URL: https://aclanthology.org/2022.findings-acl.185.pdf

OWASP Foundation. (2025). OWASP AI Testing Guide. https://owasp.org/www-project-ai-testing-guide/

Pesaranghader, A., & Li, E. (2026). Hallucination detection and mitigation in large language models. arXiv. https://arxiv.org/pdf/2601.09929

Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). A survey on metamorphic testing. IEEE Transactions on Software Engineering, 42(9), 805–824. https://doi.org/10.1109/TSE.2016.2532875

Singh R. (2026, April 28). Citigroup lifts AI market view to over $4 trillion on enterprise adoption. Reuters. URL: https://www.reuters.com/business/finance/citigroup-lifts-ai-market-view-over-4-trillion-enterprise-adoption-2026-04-28/

Sok, C., Luz, D., & Haddam, Y. (2025). MetaRAG: Metamorphic testing for hallucination detection in RAG systems. CEUR Workshop Proceedings. https://ceur-ws.org/Vol-4136/iaai6.pdf

Xie, X., Ho, J. W. K., Murphy, C., Kaiser, G., Xu, B., & Chen, T. Y. (2011). Testing and validating machine learning classifiers by metamorphic testing. Journal of Systems and Software, 84(4), 544–558. https://doi.org/10.1016/j.jss.2010.11.920

Yang, J., Chen, D., Sun, Y., Li, R., Feng, Z., & Peng, W. (2024). Enhancing semantic consistency of large language models through model editing: An interpretability-oriented approach. Findings of the Association for Computational Linguistics: ACL 2024, 3343–3353. https://aclanthology.org/2024.findings-acl.199.pdf

Zhang, W., & Zhang, J. (2025). Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics, 13(5), 856. https://doi.org/10.3390/math13050856

Lavrik, N. (2026). Transparent financial architecture as a tool for feducing graud risks. Global Prosperity, 6(1). https://gprosperity.org/index.php/journal/article/view/256

Lavrik, N. (2026). Transformation of accounting into a strategic asset based on the practice of clear finance architecture. SWorldJournal, 4(36-04), 16–30. https://doi.org/10.30888/2663-5712.2026-36-04-013