DOI: 10.1145/3845988 ISSN: 2770-6699

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems

Reza Yousefi Maragheh, Yashar Deldjoo

Large language models (LLMs) are evolving from passive text generators into agentic systems that can plan, maintain state, invoke tools, and coordinate with other agents. This perspective paper examines what this shift means for recommender systems (RS). We define agentic recommender systems as recommendation pipelines in which one or more stateful agents observe, plan, call tools, and verify, rather than score in a single shot, while operating over users, item catalogs, candidate sets, and recommendation objectives. Their value is strongest when this added machinery measurably improves recommendation-layer outcomes such as relevance, constraint satisfaction, bundle coherence, grounding, explanation faithfulness, fairness of exposure, or user effort, and is not justified merely because a pipeline contains an LLM or several modules. To make these notions precise for recommendation rather than for agents in general, we introduce a recommender-specific formalism that models an individual recommender agent by its state (user, context, history, and candidate set), a reasoning core, recommendation-specific tools, a hierarchical memory, and explicit policy constraints, and captures a multi-agent recommender as a triple of agents, a shared environment exposing the item catalog and feedback signals, and a communication protocol. Within this framework, we develop four representative task families (interactive goal-oriented recommendation, user simulation and evaluation, contextual and multimodal recommendation, and grounded explanation) and an operational agenda that ties five recurring challenge families (communication protocols, scalability and cost, hallucination and error propagation, emergent misalignment and collusion, and brand and policy compliance) to measurable RS signals. Finally, we conduct a controlled empirical study comparing single-shot and multi-agent pipelines under shared user histories, candidate sets, prompts, and metrics. A pilot next-item ranking study on Amazon-2023 categories shows that multi-agent systems are not uniformly superior: on representative samples the single-shot baseline is Pareto-efficient, whereas decomposition and ensemble agents become useful mainly for high-diversity user histories. This supports a conditional design principle: agentic complexity should be routed to the cases where its marginal quality improvement justifies the additional latency, cost, and governance risk. To support reproduction, we release all pipeline implementations, prompts, and results here: https://github.com/RezaYM/agenticrecsys.git.