| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 143 |
| Year of Publication: 2026 |
| Authors: Adekemi O. Amoo, Damilare E. Bakare, Theresa O. Omodunbi, Mary T. Onifade, Samuel I. Omilo |
10.5120/ijca8d87a675ada3
|
Adekemi O. Amoo, Damilare E. Bakare, Theresa O. Omodunbi, Mary T. Onifade, Samuel I. Omilo . A Real-Time Voice Fraud Detection Framework using Whisper and Fine-Tuned Phi-3 Mini for Context-Aware Fraud Classification. International Journal of Computer Applications. 187, 143 ( Sep 2026), 23-30. DOI=10.5120/ijca8d87a675ada3
Voice fraud has recently become a critical cybersecurity concern, where fraudsters impersonate trusted individuals over phone calls to deceive victims and obtain confidential information. Conventional fraud detection systems, which rely on predetermined rules or post-call detection, are increasingly becoming ineffective against evolving fraud patterns and do not ensure user safety in real-time. This study proposes the design and implementation of a real-time voice fraud detection system that analyses real-time conversations to identify potential fraudulent activity. The proposed system integrates OpenAI’s Whisper model for speech-to-text transcription with a fine-tuned Phi-3-mini Large Language Model (LLM) for contextual fraud detection. The system architecture uses the Browser MediaStream API for audio capture and WebSocket communication to transmit streaming data to a FastAPI-based backend for real-time processing. The fine-tuned LLM was trained on a dataset comprising 447 labeled conversational samples and subsequently evaluated on 307 unseen test samples, achieving an overall accuracy of 95%. Experimental results demonstrate a precision of 97% and a recall of 98% for normal conversation, and a precision of 75% and a recall of 67% for fraudulent conversation. These findings underscore the effectiveness of leveraging context-aware language modelling and real-time audio transcription in detecting fraudulent speech patterns, thereby surpassing the limitations of traditional rule-based systems. The study contributes to the body of knowledge by demonstrating the feasibility of combining speech processing and large language models to enable proactive, real-time detection of voice fraud during phone calls.