Adaptive Verification and Confidence-Aware Routing for Reliable Multi-Agent Large Language Model Reasoning

Authors

  • Andrew Hollis epartment of Computer Science, University of Texas at Dallas, USA Author
  • Uday Naidu Department of Computer Science, George Mason University, USA Author
  • Darren Lowe Department of Computer Science, University of Central Florida, USA Author

Keywords:

large language models; multi-agent systems; confidence calibration; uncertainty estimation; adaptive verification; dynamic routing; reliable artificial intelligence

Abstract

Large-language-model (LLM) multi-agent systems can improve difficult reasoning by decomposing a problem, assigning specialized roles, and checking intermediate outputs.  Most existing systems, however, invoke verification according to a fixed workflow and route subtasks using static role descriptions or uncalibrated self-reports.  They consequently  spend  computation on  low-risk  states while allowing  confidently wrong,  high-impact  states to propagate.  We propose CAVR-MAS (Confidence-Aware Verification and Routing for Multi-Agent Systems), a control framework that couples feature-augmented confidence calibration, domain-conditioned reliability memory, uncertainty-aware routing, and multi-level verification.  For every node in a subtask dependency graph, CAVR- MAS estimates calibrated correctness, path disagreement, downstream propagation risk, task criticality, and verification cost.  A budget-adaptive controller selects among no check, critique, cross-agent checking, evidence- grounded checking, and re-decomposition by maximizing estimated error reduction per unit cost.  Routing and final fusion use conservative posterior estimates of agent reliability rather than raw confidence or majority voting. We prove one-step Bayes optimality under correct estimators, a finite-error routing-regret bound, monotone verification workload with respect to the trigger threshold, a cumulative budget-violation bound, and consistency of nondiscounted reliability memory.   We also define an executable evaluation protocol on  GSM8K, MATH, PlanBench,  MuSiQue,  and  HotpotQA, with controlled baselines,  ablations,  calibration analysis,  and paired statistical tests.  The result is a technically complete reliability-control formulation whose empirical claims are explicitly falsifiable.

References

[1] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, vol. 35, pp. 24824–24837, 2022.

[2] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in International Conference on Learning Representations, 2023.

[3] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in International Conference on Learning Representations, 2023.

[4] S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, “Tree of Thoughts: Deliberate problem solving with large language models,” in Advances in Neural Information Processing Systems, vol. 36, 2023.

[5] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegrefe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al., “Self-Refine: Iterative refinement with self-feedback,” in Advances in Neural Information Processing Systems, vol. 36, pp. 46534–46594, 2023.

[6] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” in Advances in Neural Information Processing Systems, vol. 36, 2023.

[7] Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving factuality and reasoning in language models through multiagent debate,” in Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pp. 11733–11763, 2024.

[8] G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “CAMEL: Communicative agents for ‘mind’ exploration of large scale language model society,” in Advances in Neural Information Processing Systems, vol. 36, pp. 51991–52008, 2023.

[9] S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, et al., “MetaGPT: Meta programming for a multi-agent collaborative framework,” in International Conference on Learning Representations, 2024.

[10] W. Chen, Y. Su, J. Zuo, C. Yang, C. Yuan, C. Qian, C.-M. Chan, Y. Qin, Y. Lu, R. Xie, et al., “AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors,” in International Conference on Learning Representations, 2024.

[11] Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al., “AutoGen: Enabling next-gen LLM applications via multi-agent conversation,” arXiv:2308.08155, 2023.

[12] J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou, “Mixture-of-Agents enhances large language model capabilities,” in International Conference on Learning Representations, 2025.

[13] G. Zhang, Y. Yue, Z. Li, S. Yun, G. Wan, K. Wang, D. Cheng, J. X. Yu, and T. Chen, “Cut the Crap: An economical communication pipeline for LLM-based multi-agent systems,” in International Conference on Learning Representations, 2025.

[14] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” in International Conference on Learning Representations, 2025.

[15] J. Dekoninck, M. Baader, and M. Vechev, “A unified approach to routing and cascading for LLMs,” in Proceedings of the 42nd International Conference on Machine Learning, 2025.

[16] C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, et al., “ChatDev: Communicative agents for software development,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 15174–15186, 2024.

[17] Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.

[18] J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins, “Solving math word problems with process- and outcome-based feedback,” arXiv:2211.14275, 2022.

[19] H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” in International Conference on Learning Representations, 2024.

[20] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proceedings of the 34th International Conference on Machine Learning, PMLR 70, pp. 1321–1330, 2017.

[21] K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. Manning, “Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback,” in Proceedings of EMNLP, pp. 5433–5442, 2023, doi:10.18653/v1/2023.emnlp-main.330.

[22] L. Kuhn, Y. Gal, and S. Farquhar, “Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation,” in International Conference on Learning Representations, 2023.

[23] J. Geng, F. Cai, Y. Wang, H. Koeppl, P. Nakov, and I. Gurevych, “A survey of confidence estimation and calibration in large language models,” in Proceedings of NAACL-HLT, pp. 6577–6595, 2024.

[24] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al., “Training verifiers to solve math word problems,” arXiv:2110.14168, 2021.

[25] D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” in NeurIPS Datasets and Benchmarks Track, 2021.

[26] K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati, “PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change,” in Advances in Neural Information Processing Systems, vol. 36, pp. 38975–38987, 2023.

[27] H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “MuSiQue: Multihop questions via single-hop question composition,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 539–554, 2022, doi:10.1162/tacl__a__00475.

[28] Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning, “HotpotQA: A dataset for diverse, explainable multi-hop question answering,” in Proceedings of EMNLP, pp. 2369–2380, 2018, doi:10.18653/v1/D18- 1259.

Downloads

Published

2026-08-29

How to Cite

Adaptive Verification and Confidence-Aware Routing for Reliable Multi-Agent Large Language Model Reasoning. (2026). International Journal of Artificial Intelligence Engineering and Systems, 1(2). https://ijaies.org/index.php/home/article/view/125