| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 142 |
| Year of Publication: 2026 |
| Authors: Moqbel Tawhib Salah Abdo, Kasmi Manal |
10.5120/ijca14a09bd7b6b5
|
Moqbel Tawhib Salah Abdo, Kasmi Manal . A Hybrid Deep Learning Framework for Detecting and Mitigating Hallucinations in Large Language Models. International Journal of Computer Applications. 187, 142 ( Oct 2026), 71-77. DOI=10.5120/ijca14a09bd7b6b5
Large language models (LLMs) such as GPT-4, LLaMA-3, and Gemini 1.5 remain prone to generating fluent but factually incorrect content—a failure mode termed hallucination. This paper presents HalluGuard v2, a novel hybrid deep learning framework integrating a Self-Consistency Checker (SCC) and a Retrieval-Based Verifier (RBV) through a Query-Adaptive Gating Network (QA-GN). The QA-GN is a lightweight two-layer MLP that dynamically predicts per-query fusion weights, enabling context-sensitive combination of intra-model consistency signals and evidence-grounded verification. An Adversarial Probing Module activates as a post-hoc step for low-confidence factual decisions, challenging the model to expose fragile false confidence. A Triple-Signal OOD Handler activates dynamic web-search verification when retrieval confidence falls below a corpus-coverage threshold. A new expert-annotated benchmark, HalluDetect-3K, comprising 3,000 QA pairs spanning six domains and five hallucination categories is constructed. HalluGuard v2 achieves 90.3% accuracy and 88.1% macro F1-score on HalluDetect-3K, outperforming the strongest prior baseline by 11.4 and 11.7 percentage points, with AUC-ROC = 0.941 confirmed by five-fold cross-validation (88.9% ± 0.35%). The QA-GN contributes 5.9 pp over static logistic regression; adversarial probing adds a further 0.7 pp. A dedicated adversarial robustness evaluation demonstrates that a Suspicion Score mechanism reduces evasion from 31.4% to 12.5%. Corrected responses achieve 4.61/5 factual accuracy while preserving 91.2% BERTScore fluency. Latency optimisation via token-probability entropy approximation reduces inference from 3.6 s to 1.2 s at a 2.4 pp accuracy cost.