•  
  •  
 

Corresponding Author

SUDHEER REDDY BANDI

Subject Area

Computer Science

Article Type

Special Issue Original Study

Abstract

The rapid growth of urban areas has greatly heightened the need for smart video surveillance systems that can automatically process extensive amounts of CCTV footage. Traditional surveillance methods largely depend on human monitoring, which is not only inefficient but also susceptible to human mistakes, especially in intricate and crowded environments. To tackle these issues, this paper introduces a combined object detection and temporal attention for intelligent video surveillance that concurrently analyzes spatial and temporal data from video streams. The proposed system analyzes real-time CCTV footage utilising a multi-pathway frame extraction technique that includes slow, fast, and full-frame sampling to capture both immediate movements and long-term temporal relationships. A lightweight pathway combines YOLOv8 with the SlowFast network to enable real-time object detection and action recognition, prioritizing computational efficiency and quick responses. Simultaneously, a high-capacity pathway uses a 3D convolutional neural network (R3D-18) along with a Transformer-based temporal attention module to accurately model intricate spatial-temporal patterns. The outputs from both pathways are integrated via an ensemble fusion module employing weighted voting and meta-learning, allowing for robust decision-making. The experimental analysis shows that the suggested framework delivers enhanced accuracy, resilience, and real-time capabilities when compared to conventional surveillance systems, making it well-suited for extensive intelligent monitoring applications.

Keywords

CCTV Surveillance, Machine Learning, Computer Vision, Deep Learning, Real-Time Monitoring, Smart Cities

Creative Commons License

Creative Commons Attribution 4.0 License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS