Measured Benchmarks
Every number on this page was measured on the project's own test bench, on one machine, with one evaluation protocol applied identically to all three models. Nothing here is quoted from a vendor page, and the methodology is published in full underneath the table.
Results
fotonete is compared against two models chosen because they occupy the same practical niche: a small CNN detector and a small transformer detector. All three run at 640×640 on the same hardware, in the same process, through their own official code paths.
| Model | Arch | Params | GFLOPs | E2E | B1 GPU | B8 GPU | VRAM | COCO mAP |
|---|---|---|---|---|---|---|---|---|
| fotonete | CNN | 1.72 M | 5.42 | 8.3 ms | 3.43 ms | 10.51 ms | 201 MB | 0.171 |
| YOLO26n | CNN | 2.41 M | 5.36 | 10.4 ms | 6.95 ms | 16.73 ms | 235 MB | 0.395 |
| D-FINE N | Transformer | 3.75 M | 7.07 | 22.1 ms | 14.39 ms | 38.39 ms | 353 MB | 0.428 |
E2E = full predict path including preprocessing and postprocessing. B1/B8 = pure GPU forward pass at batch 1 and batch 8. VRAM = peak allocated during the run.
Reading the table honestly
fotonete is not the most accurate model here — it sits well below both peers on mAP. That is the deliberate trade this project makes. It is the smallest model in the comparison by a wide margin, roughly a third of D-FINE N's parameters, and like the best modern detectors it needs no NMS stage. If you are optimising for accuracy alone, a peer will beat it. If you are optimising for latency per megabyte of weights on hardware you already own, that is the axis fotonete is built for.
It is also worth stating plainly: fotonete is still mid-training. The row shows its latest published checkpoint, not a finished model. Accuracy will move, and the aim is a model that lands in the same league as the peers above — that is the hope, not a measurement.
Methodology
Reproducible numbers require stating the conditions, so here they are in full.
| Setting | Value |
|---|---|
| CPU | 13th Gen Intel Core i3-13100F (4C / 8T) |
| GPU | NVIDIA GeForce RTX 4060, 8 GB |
| Driver / CUDA | 610.74 / 12.8 |
| OS | Windows 11 |
| PyTorch | 2.11.0+cu128 |
| Input size | 640 × 640 |
| Latency iterations | 80 end-to-end, 150 forward (median reported) |
| Evaluation set | COCO val2017, all 5,000 images |
| mAP protocol | pycocotools, conf 0.01, maxDets [1, 10, 100] |
| Compute measurement | thop |
- Each model runs in its own subprocess, so cold start and peak VRAM are honest.
- Peer models are evaluated through their own official packages and source, not a reimplementation.
- mAP is computed once over the full validation set with identical settings for every model, which is deterministic.
- Timing figures are medians; the p90 end-to-end for fotonete is 8.8 ms against a 8.3 ms median.
Caveat: the GPU is shared with a live training run during measurement. That makes the published latencies conservative — a quiet GPU would be faster, never slower. Accuracy is unaffected.
In the browser
The same weights run client-side on this site through ONNX Runtime Web, on WebGPU where the device supports it and WASM everywhere else. The backend is chosen once on load by running a real inference on a known reference frame, then cached, so a device that advertises WebGPU but cannot actually run it falls back silently instead of showing an empty canvas.
The live inference demo is the fastest way to see what the model actually does on a given machine — and the only honest way to judge whether it fits your hardware.