Failure Taxonomy of Retrieval-Augmented Generation Systems in High-Stakes Information Environments

Main Article Content

Tan Wei Syn Wendy
Hasan Mahmud
Izzati Shafa

Abstract

Retrieval-Augmented Generation (RAG) systems are increasingly adopted to improve factuality, evidence grounding, and domain adaptability in knowledge-intensive applications. However, their use in high-stakes information environments introduces failure patterns that cannot be reduced to hallucination or answer inaccuracy. This study develops a pipeline-level and risk-sensitive taxonomy of RAG failures across clinical, legal, financial, and public administration scenarios. The analysis identifies five dominant failure categories: retrieval incompleteness, evidence misalignment, unsupported synthesis, citation fabrication, and calibration failure. Legal scenarios recorded the highest retrieval incompleteness count with 31 cases, followed by public administration with 28 cases and clinical scenarios with 24 cases. Financial scenarios showed the strongest unsupported synthesis pattern with 29 cases, indicating the risk of converting conditional evidence into deterministic conclusions. Pipeline-level analysis further shows that generation failures reached a 35.2% failure share, retrieval failures reached 32.7%, attribution failures reached 29.8%, reranking failures reached 24.5%, and query parsing failures reached 18.4%. Evidence quality strongly influenced failure severity, as severe failure rates decreased from 29.5% in low-quality evidence strata to 5.1% in very high-quality strata. Citation fabrication emerged as the most critical risk category because it combines low detectability, false traceability, and high recovery burden. Validation results show that taxonomy refinement improved failure presence agreement from 0.74 to 0.91, category agreement from 0.61 to 0.84, severity agreement from 0.58 to 0.81, domain transfer from 0.64 to 0.79, and audit usefulness from 0.67 to 0.88. These findings demonstrate that high-stakes RAG evaluation requires more than final-answer scoring. The proposed taxonomy provides an operational basis for audit design, benchmark development, evidence governance, and risk-aware deployment of RAG systems.

Article Details

Section
Articles