Responsible Large Language Models in Finance: A Descriptive Bibliometric Overview and Taxonomy of Responsibility
Chong Hui Tan, Qinxu DingThe rapid adoption of large language models (LLMs) in financial services has generated a growing literature on “responsible AI” in domains such as investment analysis, credit assessment, risk management, compliance, and financial advisory systems. Unlike earlier AI systems, LLMs introduce responsibility challenges that differ in important ways from those addressed by earlier responsible AI frameworks, including hallucinations, prompt injection and manipulation, generative opacity, instruction-following failures, context sensitivity, output inconsistency, and emergent capabilities that arise only at larger model scales. While responsible AI concerns extend across AI systems more broadly, this review focuses specifically on LLMs, which since 2022 have become a major technological force reshaping financial services and have generated a distinctive body of responsibility discourse that existing frameworks are only beginning to address. However, the meaning of responsibility in this literature remains heterogeneous and inconsistently operationalized across studies. This paper provides a structured review of research on responsible LLMs in finance using a broad-to-narrow screening process across Web of Science and Scopus, reported with reference to the applicable PRISMA 2020 items. We combine a descriptive bibliometric overview with qualitative content analysis to examine the corpus. The descriptive overview covers publication trends, publication venues, disciplinary orientations, geographic distributions, and institutional affiliations, while the qualitative analysis codes the corpus on two dimensions: responsibility depth—whether responsibility is constitutive of the paper’s contribution, operative within its design, or peripheral to its framing—and responsibility mode—whether that engagement is normative, technical, evaluative, or mixed. The findings reveal a field concentrated in the Operative–Technical mode. Constitutive contributions cluster in Normative and Mixed modes, while a purely Evaluative orientation remains comparatively uncommon. Addressing these limitations requires moving beyond principle-level discourse toward institution-specific accountability mechanisms, empirical evaluation across clearly documented research and operational settings, and systematic alignment with relevant financial regulatory requirements. These efforts must address the distinctive challenges posed by LLMs, rather than only concerns inherited from the broader responsible AI literature.