A HYBRID FRAMEWORK FOR ZERO-SHOT SUSPICIOUS BEHAVIOR DETECTION USING LARGE MULTIMODAL MODELS AND LOCAL OBJECT PROPOSERS

Main Article Content

Ahmed Yassin Mohammad, Abdulamir Abdullah Karim

Abstract

Conventional deep learning video surveillance systems excel at identifying predetermined activities but encounter considerable constraints in scalability and adaptability, stemming from their need on large labeled datasets and intrinsic semantic inflexibility. These algorithms frequently encounter difficulties in identifying fresh or unexpected suspicious activities and generally lack the ability to deliver context-aware, interpretable outputs. This research presents a novel hybrid framework that combines the precise cognitive reasoning of Large Multimodal Models (LMMs) with the efficiency of lightweight edge computing to facilitate zero-shot identification of suspicious activity. The proposed distributed architecture operates in two primary stages: first, a YOLOv8n model deployed at the edge performs real-time, efficient detection of human-centric events, significantly reducing computational overhead. Second, a temporally coherent sequence of relevant frames, augmented with a carefully engineered textual prompt, is dispatched to a cloud-hosted LMM (InternVL) for deep semantic analysis of spatio-temporal interactions. The findings reveal that the suggested model exhibited consistently strong performance in detecting anomalous events across the three analyzed categories: Fighting, Assault, and Vandalism. The highest accuracy was achieved in the " Fighting " category at 95.6%, followed by " Assault " at 94.2%, and " Vandalism " at 93.3%. Experimental assessments performed on the standard UCF-Crime dataset indicate that this prompt-driven, zero-shot methodology markedly surpasses conventional anomaly detection benchmarks. The framework transcends basic classification by producing comprehensive, human-readable JSON outputs that elucidate the characteristics of the suspicious activity, the individuals implicated, and a quantitative confidence score, thereby directly tackling the essential issue of interpretability in security-sensitive AI applications. This study introduces a scalable and adaptable framework for advanced intelligent surveillance, which efficiently minimizes the gap between low-level visual perception and high-level semantic comprehension.

Article Details

Section
Articles