01 / Overview
Overview
Detection overlays alone can hide missed people or broken identities. The interface replays the same source and result frames side by side, with translucent overlays and crossing records.
A local Python app performs detection, tracking and counting. The public website is a viewer for recorded videos, CSV and JSON.
02 / System
System
The default app uses YOLO26n + ByteTrack. Confidence 0.1, a gate from (0.1, 0.55) to (0.9, 0.55), downward IN and an 8 px boundary band stay fixed. Comparison videos are rendered from saved coordinates without rerunning the model.
Detector comparisons share ByteTrack and the counter. Tracker comparisons reuse boxes and scores from the first YOLO run. No retraining or threshold tuning was performed on the evaluation clips.
- 01Input
Source frames
- 02Detect
YOLO26n
- 03Track
ByteTrack
- 04Count
Finite gate · IN / OUT
- 05Compare
Video · run evidence
2026.09.21 / Stress tests
Stress tests
Three new 15-second clips were generated with Higgsfield Seedance 2.5. Each has 361 frames, totaling 1,083 unique frames. Four detectors and three cached-tracker settings ran three times each. Fifteen crossing intervals were locked from source-only AI review before inference.
The run fixes FP32, TF32 OFF, 640×640 and batch 1, with the same ByteTrack and counter. YOLO26n uses its one-to-many head with external NMS. YOLO letterboxing, RT-DETRv2 warping and model capacities still differ.
All four detectors produced 5 IN / 0 OUT for perspective, 0 / 0 for occlusion and 5 / 5 for reappearance. Strict source intervals matched 13 / 15 events; the predefined ±0.5 s analysis matched 15 / 15. This is direction/time agreement, not identity accuracy.
| Detector | Perspective FPS | Occlusion FPS | Reappearance FPS |
|---|
| YOLO26n | 89.91 (86.27–90.54) | 93.63 (92.73–94.84) | 91.03 (88.71–92.00) |
| YOLOv8n | 105.11 (104.85–109.41) | 116.71 (116.42–120.73) | 113.61 (112.70–117.24) |
| YOLO11n | 94.14 (93.39–94.74) | 99.65 (96.28–101.62) | 100.22 (94.94–100.63) |
| RT-DETRv2-S | 39.08 (38.22–40.03) | 38.73 (38.21–41.30) | 39.17 (38.98–39.74) |
- FPS: median and min–max of three RTX 4090 runs. Reading, detection, tracking, counting and JSONL are included; initialization, five warm-ups, rendering and encoding are excluded. GPU jobs ran sequentially on a desktop in use.
- The requested bidirectional crossing failed to generate: five people approach in one direction. The occlusion clip has zero visible crossings, which does not prove no crossing occurred behind the wall. Requested motion and actual footage are distinguished.
- YOLOv8n had the highest throughput in these conditions. Equal counts do not establish overall detection quality or a ranking across environments. These runs are separate from the 09.15 FP16/native-input comparisons.
2026.09.21 / Tracking
Missed crossing
ByteTrack, BoT-SORT and TrackTrack receive the same ordered boxes and scores from the first YOLO26n run. ReID and camera-motion compensation are off, and the lost buffer is 24 frames. Each tracker retains its native association thresholds.
TrackTrack missed one of five reviewed crossings in the perspective clip in all three repeats. IDs decreased from 12 to 10 to 5 while counts changed from 5 to 5 to 4; fewer IDs did not establish better performance.
The missed subject was the leading gray-shirt person. At F135, ByteTrack and BoT-SORT counted IN with ID 4 while TrackTrack left the same box unassigned. The maximum score among foot-covering candidates in F0–138 was 0.674920, below the 0.7 new-track threshold. Only ID-assigned points enter the counter, so the missing new ID excluded this input. This observation involves mixed/duplicate boxes in this configuration; no threshold-change experiment was run.
In the first-repeat fixed-anchor review, none of the seven conditions confirmed the same ID across wall occlusion at F18→192 or exit/reappearance at F120→240: IDs changed or detections/IDs were absent. All seven retained the same ID at the two short-passing anchors F48→72, which does not establish continuity in every intervening frame. Source identity remains an appearance-based inference, and the long absences exceed the 24-frame retention window.
| Tracker | Perspective IN / OUT | Track IDs | p50 / p95 ms |
|---|
| ByteTrack | 5 / 0 | 12 | 0.528 / 0.699 |
| BoT-SORT | 5 / 0 | 10 | 0.511 / 0.688 |
| TrackTrack | 4 / 0 | 5 | 0.599 / 0.788 |
- These times are the three-run medians of tracker-stage latency, excluding detection. They are not directly comparable with pipeline FPS.
- All three settings produced 0 / 0 for occlusion and 5 / 5 for reappearance. Long occlusions and exits exceed the 24-frame retention window; identity recovery with ReID off is not guaranteed.
- Source reviews, per-repeat events, frame coordinates and video-alignment evidence are published. The default app remains on ByteTrack.
03 / 09.15 Detection
Earlier detectors
Each of three detectors ran three times over 602 unique source frames. Every repeat produced 4 IN / 8 OUT for the crowd clip and 1 IN / 1 OUT for the two-person clip. Repeated footage is not an independent sample.
FPS below shows the median and min–max of three runs on RTX 4090. It includes reading, detection, tracking, counting and JSONL logging; initialization, five first-frame warm-ups, rendering and video encoding are excluded. GPU runs were sequential and synchronized through CUDA completion.
| Detector | Input · precision | Crowd FPS | Two-person FPS |
|---|
| YOLO26n | 640 · FP16 | 60.89 (58.61–61.23) | 67.01 (59.29–74.07) |
| RF-DETR Small | 512 · FP32 eager | 31.66 (31.04–33.40) | 31.07 (30.81–33.31) |
| DEIMv2-S | 640 · FP32 deploy | 18.29 (15.73–20.88) | 21.57 (20.40–23.01) |
- Native input size, preprocessing and precision differ by model. This is not an equal-compute comparison or a general model ranking.
- The earlier demo’s 24.08 FPS included solid dashboard rendering, storage and first inference. Its scope differs from the new FPS values, which exclude rendering.
- Measurements were taken on a desktop in use, not an isolated benchmark machine. Capture-to-display latency was not measured.
04 / 09.15 Tracking
Earlier trackers
Across 602 frames × three settings × three repeats, all 5,418 frame records preserved the same YOLO cache order, timestamps, boxes and scores. All settings produced crowd counts of 4 / 8 and two-person counts of 1 / 1.
These are tracker-stage latencies for the crowd clip, excluding detection. Each value is the median of the three per-run p50 or p95 values. Cached playback FPS is not full inference FPS.
| Tracker | Track IDs | p50 ms | p95 ms |
|---|
| ByteTrack | 22 | 0.720 | 1.607 |
| TrackTrack | 18 | 1.307 | 3.004 |
| TrackTrack + ReID | 18 | 16.333 | 98.002 |
- ReID ON and OFF produced identical ID assignments in every frame. No additional ReID benefit was observed on these two clips.
- The coat-region ByteTrack boxes had duplicate IDs 16/35 before continuing as 35. TrackTrack recovered ID 12 but left frames 291–293 unassigned. A reduction from 22 to 18 IDs does not establish a lower overall fragmentation rate.
- This compares the Ultralytics TrackTrack implementation. ReID uses the official ONNX, CUDAExecutionProvider and 224×224 crops; ON/OFF settings differ only in with_reid. This is not a full reproduction of the paper.
05 / 09.15 Review
Earlier review
The 14 intervals from the existing AI source review stayed fixed. YOLO26n, DEIMv2-S and both TrackTrack settings matched 14 / 14; RF-DETR matched 13 / 14. One RF-DETR OUT was 0.083 s beyond its interval; an auxiliary ±0.5 s comparison matched 14 / 14.
Matching uses clip, direction and time one to one. It does not verify person identity and is not human-consensus ground truth.
- At frame 0, YOLO missed the lower-left partial person; RF-DETR and DEIM detected it but also emitted overlapping boxes. At frame 291, all three retained two overlapping coat-region boxes.
- These two observations do not establish full-clip detection recall. Frame numbers, coordinates and scores are available in the review record.
06 / Server
Model server
A separate local HTTP server runs the actual YOLO26n weights on CPU. Recorded evidence covers 20 normal inferences, deadline handling and automatic recovery after inference-worker termination, with actual boxes, HTTP timings and error logs.
The run fixes CPU FP32, two threads and one inference slot. The website is an evidence viewer and the rehearsal server was stopped. GPU serving, long-running operation, HTTP-server/host recovery and model-version rollback were not tested.
| Check | Observed |
|---|
| Actual model inference | 20 / 20 |
| Admission limit / deadline | 429 / 504 |
| Restart after worker exit | 503 → 200 |
| Inflight after recovery | 0 |
ENGINEERING DECISIONS
Decisions
01Count evidence
The lower TrackTrack ID count on 09.15 was insufficient to choose it. The added perspective clip exposed one missed crossing while ByteTrack and BoT-SORT retained five. The default stays ByteTrack pending identity-grounded evaluation on real footage.
02Review the source
Requested people and paths in the generation prompt were not ground truth. Crossings in the completed source were reviewed before comparison with the run.
03Occupancy is unknown
With unknown initial occupancy, the calculation baseline of zero is separate from actual people present. Neither net change −4 nor the clamped zero establishes absolute occupancy.
04Masking and comparison
The local default is solid masking. The website’s 12% fill is an inspection overlay with privacy_protection=false; undetected people also remain unmasked in solid mode.
SCOPE & LIMITS
Limits
- The experiments cover two earlier AI-generated clips with 602 frames and three added clips with 1,083 frames. Repeated runs are not new samples. Full-frame box/ID/mask ground truth and human-consensus labels are absent; HOTA, IDF1 and field accuracy remain unmeasured.
- Generated footage has unnatural strides, edge cropping and similar appearances. Initial occupancy is unknown, and fragmented IDs are not additional unique people.
- Real CCTV, webcams, RTSP, long-running operation and capture-to-display latency remain unverified. The added comparison matches FP32 and input dimensions, while preprocessing and model size differ. Dense repeated crossings were not adequately generated.
- SAM 3.1 was not run because its official checkpoint requires access approval (HTTP 401).