DOI: 10.3390/app16199699 ISSN: 2076-3417

A Hybrid BDI + RAG Multi-Agent Architecture: Plan-Based Reasoning over Multimodal Retrieval

Halil Yesil, Baris Tekin Tezel, Moharram Challenger

Retrieval-Augmented Generation (RAG) gives language models access to external text and image collections, but it does not make the decisions taken on that content reproducible. Classical Belief–Desire–Intention (BDI) agents provide explicit plans and traceable decisions, although they normally expect beliefs in symbolic form. We connect these capabilities in a hybrid architecture in which RAG updates agent beliefs while BDI plans remain responsible for reasoning and coordination. The plans are implemented in AgentSpeak on JASON/JADE, agents communicate through FIPA-ACL, and probabilistic models are accessed as computational services through the Model Context Protocol (MCP). The orchestrator routes messages but does not derive authorization decisions. Domain agents and the fusion agent produce those decisions through plan execution. We instantiated it for multimodal access control over 167 research posters and compared it with AutoGen and LangGraph baselines on an identical service backend driven by one locally served model, qwen3.5:9b, so that only the control layer differs. Eight campaigns each vary one setting: pipeline length, tool declaration order, the prerequisite and terminal hints, the sampling seed and the fusion policy. They total 4800 baseline runs, with the plan-based arm measured in the two campaigns that have a plan-library counterpart, 300 runs each. Under a favorable configuration, both baselines schedule the workflow correctly and no verdict changes between repeats, yet only 52% and 86% of queries reproduce the same sequence. Their valid-schedule rate falls to 32.0% and 28.0% at nine agents, and presentational edits to the tool declarations move the ordering and, in one arm, the verdict. The plan-based arm returns 100.0% valid schedules and a single sequence in both campaigns.