Publication Details
Abstract
Standard large language models (LLMs) are more complicated by the intense morphological richness and lack of high-quality datasets, as well as the ever-present danger of factual hallucinations when processing and reasoning over long Arabic text. This paper presents a novel explainable multi-agent retrieval-augmented framework, called XARAG+ , for long-document comprehension in Arabic. Unlike a single language model, XARAG+ carries out operations in a specialized pipeline, comprising of document preprocessing, semantic retrieval, context verification, logical inference and explanation generation. It is a multi-agent coordination mechanism which strives to remove inconsistent and redundant retrieved chunks before the final inference step. The framework should be systematic in its organization and link all of the answers with the verified texts to make it easy to see and easily checkable. Last, XARAG+ develops a verifiable modular architecture appropriate to various domains which have a stringent requirement of factual accountability, e.g. Legal Arabic Document Analysis and Scientific Arabic Document Analysis.