Multi-Camera Person Tracking
A warehouse-safety system that tracks people across four CCTV cameras and shows everyone on one top-down floor map, with the same ID for a person in every view. Built as a prototype for Sentics GmbH.
The dashboard running on the four warehouse cameras. Green boxes are YOLOv8 person detections in each view. Underneath, every person is projected onto the shared floor map with a single global ID.
yolo_trackdown_1.py. Each detection is reduced to the point where the person's feet touch the floor, because only points on the floor plane map correctly through a homography.1 · Calibrating each camera (homography)
homography_calibrator.py grabs the first frame of a camera. I click four points on the floor, and they are matched to a square in world coordinates (100,100) → (700,700). cv2.findHomography then solves for the 3×3 matrix H, which is saved as h_camN.npy. There is one per camera.
# clicked image points ↔ top-down world square world_points = [[100,100],[700,100],[700,700],[100,700]] H, _ = cv2.findHomography(np.array(image_points), np.array(world_points)) np.save("homographies/h_cam2.npy", H)
A homography maps one plane to another. The warehouse floor is a plane, so any pixel on the floor can be moved into a shared top-down coordinate system. This lets detections from four different viewpoints be compared directly.
2 · Detect → project
Every frame is resized to 640×360 and passed through YOLOv8. Only class 0 (person) is kept. The bottom-centre of each box goes through cv2.perspectiveTransform with that camera's H. The result becomes a small 20×20 box on the floor map.
center_bottom = np.array([[[ (x1+x2)//2, y2 ]]], dtype='float32') topdown_pos = cv2.perspectiveTransform(center_bottom, homographies[cam_id]) bbox_topdown = [x-10, y-10, x+10, y+10]
3 · Merge duplicates with DBSCAN
A person seen by three cameras gives three nearby points on the map. cluster_topdown_bboxes() runs DBSCAN (eps = 25, min_samples = 1) on the box centres and averages each cluster into a single detection. DBSCAN suits this job because I don't know in advance how many people are in the room.
4 · A custom global tracker
tracking/custom_tracker.py keeps a list of Trajectory objects, each holding up to 100 recent boxes in a deque. Each frame it does four things:
- Builds an IoU cost matrix between every live trajectory's last box and every new detection (vectorised with
np.repeat/np.tile). - Solves the assignment optimally with the Hungarian algorithm (
scipy.optimize.linear_sum_assignment). - Smooths positions with an exponential moving average (α = 0.2) so dots don't jitter.
- Manages lifecycles. Unmatched tracks coast on their last box and die after 100 missed frames. Short tracks (< 10 nodes) die after a single miss, which removes phantom IDs. A track only counts as "real" once it has been matched more than 5 times.
receivers, providers = linear_sum_assignment(-iou_matrix) for r, p in zip(receivers, providers): traj = trajectories[alive_index[r]] traj.add_node(candidates[p], p) # EMA-smoothed traj.lost_times = 0
5 · A second approach: per-camera DeepSORT
yolo_topdown_tracker.py tries the other design. YOLOv8s runs with DeepSORT (max_age=30) inside each camera, after an adaptive filter (confidence ≥ 0.4, width ≥ 20 px, height ≥ 35 px, aspect ratio 0.9–5.0) that throws out bad boxes. The projection point is a blend, 0.8·centre + 0.2·bottom, which holds up better when feet are hidden behind shelves. Comparing both designs showed the trade-off: appearance-based re-ID works well within one camera, while geometric fusion is what keeps IDs consistent across cameras.
Output
A single dashboard: the four camera views in a 2×2 grid with their boxes, and underneath, the top-down room with each person drawn as a coloured dot labelled "Person N".