Benchmarks

Measured Benchmarks

Every number on this page was measured on the project's own test bench, on one machine, with one evaluation protocol applied identically to all three models. Nothing here is quoted from a vendor page, and the methodology is published in full underneath the table.

Results

fotonete is compared against two models chosen because they occupy the same practical niche: a small CNN detector and a small transformer detector. All three run at 640×640 on the same hardware, in the same process, through their own official code paths.

ModelArchParamsGFLOPs E2EB1 GPUB8 GPU VRAMCOCO mAP
fotoneteCNN1.72 M5.42 8.3 ms3.43 ms10.51 ms 201 MB0.171
YOLO26nCNN2.41 M5.36 10.4 ms6.95 ms16.73 ms 235 MB0.395
D-FINE NTransformer3.75 M7.07 22.1 ms14.39 ms38.39 ms 353 MB0.428

E2E = full predict path including preprocessing and postprocessing. B1/B8 = pure GPU forward pass at batch 1 and batch 8. VRAM = peak allocated during the run.

Reading the table honestly

fotonete is not the most accurate model here — it sits well below both peers on mAP. That is the deliberate trade this project makes. It is the smallest model in the comparison by a wide margin, roughly a third of D-FINE N's parameters, and like the best modern detectors it needs no NMS stage. If you are optimising for accuracy alone, a peer will beat it. If you are optimising for latency per megabyte of weights on hardware you already own, that is the axis fotonete is built for.

It is also worth stating plainly: fotonete is still mid-training. The row shows its latest published checkpoint, not a finished model. Accuracy will move, and the aim is a model that lands in the same league as the peers above — that is the hope, not a measurement.

Methodology

Reproducible numbers require stating the conditions, so here they are in full.

SettingValue
CPU13th Gen Intel Core i3-13100F (4C / 8T)
GPUNVIDIA GeForce RTX 4060, 8 GB
Driver / CUDA610.74 / 12.8
OSWindows 11
PyTorch2.11.0+cu128
Input size640 × 640
Latency iterations80 end-to-end, 150 forward (median reported)
Evaluation setCOCO val2017, all 5,000 images
mAP protocolpycocotools, conf 0.01, maxDets [1, 10, 100]
Compute measurementthop
  • Each model runs in its own subprocess, so cold start and peak VRAM are honest.
  • Peer models are evaluated through their own official packages and source, not a reimplementation.
  • mAP is computed once over the full validation set with identical settings for every model, which is deterministic.
  • Timing figures are medians; the p90 end-to-end for fotonete is 8.8 ms against a 8.3 ms median.

Caveat: the GPU is shared with a live training run during measurement. That makes the published latencies conservative — a quiet GPU would be faster, never slower. Accuracy is unaffected.

In the browser

The same weights run client-side on this site through ONNX Runtime Web, on WebGPU where the device supports it and WASM everywhere else. The backend is chosen once on load by running a real inference on a known reference frame, then cached, so a device that advertises WebGPU but cannot actually run it falls back silently instead of showing an empty canvas.

The live inference demo is the fastest way to see what the model actually does on a given machine — and the only honest way to judge whether it fits your hardware.