Real-time face recognition system leveraging facenet and VGGFace fusion embedding, and siamese network with ArcFace and triplet loss function
Loading...
Date
Publisher
BRAC University
Authors
Citation
Abstract
The face recognition system can be defined as the identification of the unique biometric
facial features of an individual through image or video frames. For CCTV-based
surveillance the facial recognition system automates the identification of any person
applied on live surveillance stream. But in CCTV surveillance system the facial
recognition endures unique challenges like low resolutions, extreme poses, illumination
variations, etc. There, a long span of development has been seen in this system
from the geometric landmark distancing to deep-learning approaches, but a significant
gap between the real-life surveillance and controlled benchmark conditions is
observed. This thesis presents a comprehensive study on the development of a face
recognition system designed for use with CCTV footage, with a focus on addressing
challenges such as side-profile recognition and low-resolution images. Traditional
face recognition models often perform poorly in real-world surveillance scenarios due
to the quality of footage, low-resolution face crops, illumination variations, the varied
angles of faces captured, different head orientations, occlusion causing partial face
visibility, and motion effects. To overcome these limitations, this research explores
the integration of a Siamese network with FaceNet and VGGFace, creating a hybrid
model that leverages the unique strengths of each component. The Siamese network
is employed to compare facial data, while VGGFace and FaceNet generate vigorous
embeddings for accurate feature extraction, and a hybrid of Triplet Loss and
ArcFace respectively enhances discriminative capabilities and combining through
directing better embedding generations for similar faces, ensuring precise identification
across challenging inputs. A custom dataset was constructed to simulate
real-world CCTV environments, incorporating diverse facial angles and resolutions
to rigorously test the model’s performance. The primary aim of the research is to
evaluate the system’s ability to accurately recognize faces under non-ideal conditions,
optimize computational efficiency, and assess the scalability of the model for
large-scale surveillance systems. Through extensive experimentation and iterative
refinement, the proposed hybrid model demonstrates improved recognition accuracy
in low-quality, a variety of angled, and different expression face images, which
contributes to advancements in the field of computer vision and offers practical implications
for security and surveillance applications. As a dataset of 5409 face crop
images from 86 individual person have been comprised here and later balanced to
a total of 4300 samples and the proposed hybrid model has achieved an accuracy
of 91.01% with 91.16% f1-score and ROC-AUC 0.9665, leading performance over
the deepface models in comparison. Also the model have shown a competitive performance
on the benchmark datasets of LFW, CelebA and VGGFace2, achieving
an accuracy of 99.28%, 98.57% and 96.82% with f1-score of 98.71%, 98.46% and
96.27% respectively. Though the gap of benchmark datasets and CCTV footage
data, it demonstrates a practical applicability for using in face recognition operations.
Description
This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 66-70).
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 66-70).
Publisher Link
Type
Thesis
Creative Commons license

Except where otherwise noted, this item's license is described as
Attribution-NonCommercial-NoDerivatives 4.0 International