Architecture

NMS-Free Object Detection: How fotonet Works

Most detectors end inference with Non-Maximum Suppression — a heuristic that throws away duplicate boxes after the model has already decided what it saw. fotonet removes that step entirely by training the head so that one prediction per object is the correct answer, not a duplicate to be filtered out later.

The problem with NMS

A standard detection head is trained one-to-many: every ground-truth object is matched to many candidate predictions so the loss keeps pushing useful gradient signal early in training. That is good for learning and bad for deployment, because the model emits many overlapping boxes for one object and something has to collapse them.

Non-Maximum Suppression does that collapse: sort by confidence, keep the best box, delete everything overlapping it above a threshold. It works, but it is a post-hoc heuristic sitting between the network and the output. It has its own IoU threshold to tune, it can delete a genuine neighbouring object, and its behaviour changes with object density — the same weights give different results on a crowded street than on an empty one.

One-to-one assignment instead

fotonet adds a second branch trained with one-to-one matching. During training, each ground-truth box is assigned to exactly one prediction through Hungarian bipartite matching on the cost matrix. A prediction that is not assigned to anything is trained directly toward the background class, so a duplicate is not merely out-ranked — it is actively wrong.

The result is that the best-scoring box is the object. There is nothing left to suppress, so the inference path carries no NMS step and no IoU threshold to tune.

Dual-branch training, single-branch deployment

One-to-one matching alone converges slowly, because early in training most predictions are poor and few of them win their assignment. fotonet therefore trains both branches at once:

  • Auxiliary one-to-many branch — supplies dense gradient supervision during early epochs so the backbone converges quickly.
  • Deploy one-to-one branch — produces the NMS-free predictions that are actually shipped.

At export the auxiliary branch is discarded entirely, at zero cost in parameters and zero cost in latency. You pay for the fast convergence during training and keep only the clean head at inference. The full assignment contract is documented on the documentation page.

The all-dense nano graph

The published model is fotonete: 1.72M parameters built from a single design law — few, wide, dense convolutions. There are no depthwise convolutions anywhere in the graph. Every backbone stage is a plain residual 3×3 stack, each neck junction is a single pointwise reduction followed by one spatial 3×3, and each head level starts from one dense 3×3 stem.

Depthwise convolutions are popular because they are cheap in parameter count, but on ordinary GPU and NPU hardware they are memory-bound and rarely reach their theoretical throughput. Wide dense layers are the shape that general-purpose accelerators actually like, which is why the model sustains real-time latency at 1.72M parameters.

StageOutput channelsStrideGrid @ 640Anchors
P364880 × 806,400
P4961640 × 401,600
P51603220 × 20400
Total———8,400

Anchor count is the deployed output length, 8,400 per image at a 640×640 input.

The output contract

Every exported graph emits a single tensor of shape [B, 8400, nc + 4]. For the 80-class COCO model that is [B, 8400, 84]: 80 class logits followed by 4 normalised box values in xywh form.

Boxes are predicted with distribution focal loss (reg_max: 12), so each side is represented as a discrete distribution rather than a single regression value. That is what lets the model express genuine uncertainty about an edge instead of averaging two plausible answers into one blurry box.

Because the graph stops there, the postprocess is short and explicit: sigmoid on the logits, pick the best class per anchor, threshold, clip to the image, cap the count. Fotonet(...).predict() applies exactly that, and any runtime that can execute ONNX or TorchScript can do the same. See export & deployment.

Checkpoint compatibility

Every configuration resolves to a stable architecture_fingerprint covering the channel profile, the strides, the regression width and the class count. Loading cross-validates that identity before any tensor is read, so a checkpoint built against a different graph is rejected with a clear message instead of being partially loaded into a mismatched network.

Why this matters in practice: existing checkpoints resume and run without conversion. There is nothing to migrate, because there is only one accepted graph.