← All projects
02 / Computer Vision · Tracking

Multi-Camera Person Tracking

A warehouse-safety system that tracks people across four CCTV cameras and shows everyone on one top-down floor map, with the same ID for a person in every view. Built as a prototype for Sentics GmbH.

4 camerassynchronised CCTV feeds
YOLOv8person detection (COCO class 0)
Homographyimage → floor-plane projection
DBSCAN + Hungariandedupe + global ID tracking
Four CCTV views with person detections above a top-down map showing Person 3, 4 and 5
REAL OUTPUT

The dashboard running on the four warehouse cameras. Green boxes are YOLOv8 person detections in each view. Underneath, every person is projected onto the shared floor map with a single global ID.

4× CCTV cam_52 · 139 140 · 142 640×360 YOLOv8 keep class 0 (person) Foot point ((x₁+x₂)/2, y₂) bottom-centre of box Homography p' = Hᵢ · p one 3×3 per camera DBSCAN eps = 25 px merge duplicates Global tracker IoU cost + Hungarian match Top-down map "Person 3" at (x, y) + 2×2 camera grid Runs every frame · all four feeds go into one shared set of detections before tracking
Pipeline in yolo_trackdown_1.py. Each detection is reduced to the point where the person's feet touch the floor, because only points on the floor plane map correctly through a homography.

Try it: four noisy cameras → one clean map

simulation

Small grey dots are raw detections from each camera after projection, one shade per camera. The larger dots are the DBSCAN-merged people with their tracker IDs. Set eps too low and one person splits into several. Set it too high and two people standing close together merge into one.

1 · Calibrating each camera (homography)

homography_calibrator.py grabs the first frame of a camera. I click four points on the floor, and they are matched to a square in world coordinates (100,100) → (700,700). cv2.findHomography then solves for the 3×3 matrix H, which is saved as h_camN.npy. There is one per camera.

# clicked image points ↔ top-down world square
world_points = [[100,100],[700,100],[700,700],[100,700]]
H, _ = cv2.findHomography(np.array(image_points), np.array(world_points))
np.save("homographies/h_cam2.npy", H)

A homography maps one plane to another. The warehouse floor is a plane, so any pixel on the floor can be moved into a shared top-down coordinate system. This lets detections from four different viewpoints be compared directly.

2 · Detect → project

Every frame is resized to 640×360 and passed through YOLOv8. Only class 0 (person) is kept. The bottom-centre of each box goes through cv2.perspectiveTransform with that camera's H. The result becomes a small 20×20 box on the floor map.

center_bottom = np.array([[[ (x1+x2)//2, y2 ]]], dtype='float32')
topdown_pos   = cv2.perspectiveTransform(center_bottom, homographies[cam_id])
bbox_topdown  = [x-10, y-10, x+10, y+10]

3 · Merge duplicates with DBSCAN

A person seen by three cameras gives three nearby points on the map. cluster_topdown_bboxes() runs DBSCAN (eps = 25, min_samples = 1) on the box centres and averages each cluster into a single detection. DBSCAN suits this job because I don't know in advance how many people are in the room.

4 · A custom global tracker

tracking/custom_tracker.py keeps a list of Trajectory objects, each holding up to 100 recent boxes in a deque. Each frame it does four things:

  • Builds an IoU cost matrix between every live trajectory's last box and every new detection (vectorised with np.repeat / np.tile).
  • Solves the assignment optimally with the Hungarian algorithm (scipy.optimize.linear_sum_assignment).
  • Smooths positions with an exponential moving average (α = 0.2) so dots don't jitter.
  • Manages lifecycles. Unmatched tracks coast on their last box and die after 100 missed frames. Short tracks (< 10 nodes) die after a single miss, which removes phantom IDs. A track only counts as "real" once it has been matched more than 5 times.
receivers, providers = linear_sum_assignment(-iou_matrix)
for r, p in zip(receivers, providers):
    traj = trajectories[alive_index[r]]
    traj.add_node(candidates[p], p)   # EMA-smoothed
    traj.lost_times = 0

5 · A second approach: per-camera DeepSORT

yolo_topdown_tracker.py tries the other design. YOLOv8s runs with DeepSORT (max_age=30) inside each camera, after an adaptive filter (confidence ≥ 0.4, width ≥ 20 px, height ≥ 35 px, aspect ratio 0.9–5.0) that throws out bad boxes. The projection point is a blend, 0.8·centre + 0.2·bottom, which holds up better when feet are hidden behind shelves. Comparing both designs showed the trade-off: appearance-based re-ID works well within one camera, while geometric fusion is what keeps IDs consistent across cameras.

Output

A single dashboard: the four camera views in a 2×2 grid with their boxes, and underneath, the top-down room with each person drawn as a coloured dot labelled "Person N".