Title: Multi-Camera Subject Tracking Through Deep Metric Embeddings and YOLO-Based Detection: A Layered System Architecture
Authors: Serif Oyindamola Oyesiji, Toussida Fatah Tanguy Minoungou, Kingsley Chinazaekpere Ndupu, Chukwudera Obumneke Anunagba.
Volume: 10
Issue: 9
Pages: 335-344
Publication Date: 2026/09/28
Abstract:
Tracking individual subjects across networks of non-overlapping cameras is a foundational capability for traffic analytics, retail and facility understanding, sports analysis, and safety monitoring, yet it remains substantially harder than either of its component problems. Single-camera multi-object tracking and appearance-based re-identification are each mature; their composition, maintaining consistent identities as subjects leave one view and appear in another under changed pose, illumination, and viewpoint, multiplies their failure modes and adds association problems neither addresses alone. This paper develops a research concept for a multi-camera subject tracking system built from three specified layers: per-camera detection and short-term tracking using YOLO-family detectors with motion-and-appearance trackers; a metric embedding layer trained with re-identification objectives to produce identity-discriminative appearance vectors under cross-view variation; and a cross-camera association layer that matches tracklets across views using appearance similarity gated by spatio-temporal feasibility derived from camera topology, managed through a gallery of identity prototypes with principled creation, update, and retirement policies. The concept specifies the embedding training regime, the association algorithm as a constrained assignment over tracklet pairs, and a real-time implementation architecture with approximate nearest neighbor indexing for gallery search at network scale. A phased evaluation plan measures per-camera tracking, embedding quality, and end-to-end cross-camera identity consistency with modern higher-order metrics on public multi-camera benchmarks, with prespecified ablations isolating each layer's contribution. Applications, failure modes, and the privacy and governance constraints that subject tracking obligates are analyzed. The concept's claim is architectural: cross-camera identity is best engineered as a composition of separately measurable layers with explicit contracts, not as a monolithic learned system.