Real-time face recognition system leveraging facenet and VGGFace fusion embedding, and siamese network with ArcFace and triplet loss function

Loading...
Thumbnail Image

Publisher

BRAC University

Citation

Abstract

The face recognition system can be defined as the identification of the unique biometric facial features of an individual through image or video frames. For CCTV-based surveillance the facial recognition system automates the identification of any person applied on live surveillance stream. But in CCTV surveillance system the facial recognition endures unique challenges like low resolutions, extreme poses, illumination variations, etc. There, a long span of development has been seen in this system from the geometric landmark distancing to deep-learning approaches, but a significant gap between the real-life surveillance and controlled benchmark conditions is observed. This thesis presents a comprehensive study on the development of a face recognition system designed for use with CCTV footage, with a focus on addressing challenges such as side-profile recognition and low-resolution images. Traditional face recognition models often perform poorly in real-world surveillance scenarios due to the quality of footage, low-resolution face crops, illumination variations, the varied angles of faces captured, different head orientations, occlusion causing partial face visibility, and motion effects. To overcome these limitations, this research explores the integration of a Siamese network with FaceNet and VGGFace, creating a hybrid model that leverages the unique strengths of each component. The Siamese network is employed to compare facial data, while VGGFace and FaceNet generate vigorous embeddings for accurate feature extraction, and a hybrid of Triplet Loss and ArcFace respectively enhances discriminative capabilities and combining through directing better embedding generations for similar faces, ensuring precise identification across challenging inputs. A custom dataset was constructed to simulate real-world CCTV environments, incorporating diverse facial angles and resolutions to rigorously test the model’s performance. The primary aim of the research is to evaluate the system’s ability to accurately recognize faces under non-ideal conditions, optimize computational efficiency, and assess the scalability of the model for large-scale surveillance systems. Through extensive experimentation and iterative refinement, the proposed hybrid model demonstrates improved recognition accuracy in low-quality, a variety of angled, and different expression face images, which contributes to advancements in the field of computer vision and offers practical implications for security and surveillance applications. As a dataset of 5409 face crop images from 86 individual person have been comprised here and later balanced to a total of 4300 samples and the proposed hybrid model has achieved an accuracy of 91.01% with 91.16% f1-score and ROC-AUC 0.9665, leading performance over the deepface models in comparison. Also the model have shown a competitive performance on the benchmark datasets of LFW, CelebA and VGGFace2, achieving an accuracy of 99.28%, 98.57% and 96.82% with f1-score of 98.71%, 98.46% and 96.27% respectively. Though the gap of benchmark datasets and CCTV footage data, it demonstrates a practical applicability for using in face recognition operations.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 66-70).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International