Deployment

Export & Deployment

fotonet exports to ONNX, TorchScript, TensorRT and calibrated INT8 ONNX. Every artifact carries a metadata sidecar describing its exact preprocessing and coordinate contract, so a runtime that has never heard of fotonet can still decode the output correctly.

ONNX

The default deployment format, and what the browser demo on this site runs.

from fotonet import Fotonet

model = Fotonet("fotonete")
artifact = model.export(format="onnx", path="exports/fotonete.onnx", imgsz=640)
print(artifact["artifact"], artifact["metadata"])

Export writes a .metadata.json sidecar next to the artifact and checks native-vs-ONNX Runtime numerical parity before returning. The metadata records the preprocessing contract: RGB, NCHW, values in [0, 1], the class names, the raw output layout, and the stride-padding rule for mapping boxes back to the caller's image.

FP16 for the browser and mobile

The model powering this site's live demo is an FP16 ONNX export, roughly half the size of the FP32 graph and the reason it loads quickly on a phone:

model.export(
    format="onnx",
    path="exports/fotonete_fp16.onnx",
    imgsz=640,
    half=True,      # FP16 weights, halved download and memory
)

Variable input sizes

Static artifacts accept only the shape they were exported with. When deployment needs variable input, export dynamic — the graph pads right and bottom to its maximum stride internally and maps normalised boxes back to the unpadded input:

model.export(
    format="onnx",
    path="exports/fotonete_dynamic.onnx",
    imgsz=(640, 960),
    dynamic=True,
)

INT8

INT8 export is calibrated static-QDQ ONNX, not a dtype cast pretending to be quantisation. Supply representative preprocessed batches in [0, 1] NCHW:

artifact = model.export(
    format="onnx",
    path="exports/fotonete_int8.onnx",
    imgsz=640,
    int8=True,
    calibration_data=calibration_batches,
)

TorchScript and TensorRT

TorchScript archives embed the same metadata contract as the sidecar, so copying the single artifact preserves the class count and preprocessing rules:

model.export(format="torchscript", path="fotonet.torchscript", imgsz=640)

TensorRT export requires trtexc in PATH:

model.export(format="tensorrt", path="fotonet.engine", imgsz=640, half=True)

Verify on your hardware. Engine building validates the build command and shape profile only. Perform platform-specific numerical and latency validation before production deployment — the benchmarks page documents how this project does it.

The output contract

Exported graphs emit [B, N, nc + 4]: class logits followed by normalised xywh. For the COCO model that is [B, 8400, 84]. The postprocess is short and is not baked into the graph:

  1. Sigmoid on the class logits.
  2. Select the best class per anchor.
  3. Threshold on confidence.
  4. Map normalised boxes through the letterbox transform and clip to the image.
  5. Cap the number of returned detections.

Fotonet("artifact.onnx").predict(...) applies exactly that, so an artifact can be reloaded through the normal API. Any other runtime can do the same from the documented layout — there is no NMS step to reimplement, which is the practical payoff of the one-to-one head.