Documentation / Getting Started / Installation & Setup
Installation & Setup
fotonet is packaged as a pure PyTorch computer vision library with zero external dependencies beyond standard scientific Python packages.
PyTorch Compatibility
fotonet requires PyTorch 2.2+ (tested through PyTorch 2.11 with CUDA 12.8). No C++ compiler or external CUDA SDK is needed for standard Python deployment.
Install for Development
git clone https://github.com/hazegreleases/fotonet.git
cd fotonet
pip install -e .
Documentation / Getting Started / Quickstart Guide
Quickstart Guide
Load pre-trained weights and run deterministic, NMS-free object detection in under 5 lines of Python code.
Basic Object Detection
from fotonet import Fotonet
import cv2
model = Fotonet("fotonete")
frame = cv2.imread("sample.jpg")
results = model.predict_bgr(frame, imgsz=640, conf=0.25)
for box in results[0].boxes:
print(f"Class: {box.cls} | Score: {box.conf:.3f} | BBox: {box.xyxy}")
Try it visually
You can explore recorded predictions and confidence thresholds on our
Inference demo.
Documentation / Getting Started / Pre-trained Checkpoints
Pre-trained Checkpoints
Official weights staged under strict release protocols. Checkpoints are never committed directly to git source repositories.
Official Checkpoint Lineup
Every published weight uses the same fixed graph (1,723,672 training parameters; 1,698,564 deployment parameters):
| Model File |
Version |
Epoch |
Parameters |
Release verification |
Status |
| fotonete.pt |
v1.0.0 |
80 / 300 |
1.73 M |
Published with release asset |
Release candidate |
Checkpoints are written continuously during training, so the fingerprint above refers to the snapshot used for the currently published benchmark figures. Weights are distributed through GitHub Releases and never committed to the repository.
Compatibility Notice
Checkpoints built against a retired graph are rejected by design rather than silently mis-loaded. For the exact contract, see
Invariance Contract.
Documentation / Architecture / Invariance Contract
Invariance Contract
The architectural contract governing fotonete graph definition, tensor shapes, and state_dict backward compatibility.
Strict Mathematical Contract
fotonet ships exactly one accepted graph. Every proposed code refactor must guarantee:
- Output tensor shape strictly preserved:
[Batch, 8400, nc + 4] (where nc = 80, 4 = xywh coordinates).
- Exact integer channel counts across stages: Backbone [64, 96, 160], Neck [64, 96, 160].
- No conversion scripts required when resuming training or loading public weights.
- Unvarying
architecture_fingerprint hash across all platforms.
Documentation / Architecture / All-Dense Residual Stacks
All-Dense Residual Stacks
Memory-efficient feature propagation inspired by dense connectivity without cross-stage buffer thrashing.
Design Principles
Traditional transformer and depthwise architectures incur significant latency overhead on edge memory busses. fotonete utilizes dense residual blocks operating with explicit integer channels, keeping compute density high while maintaining a small GPU footprint.
Hardware Alignment
Channel dimensions (64, 96, 160) are aligned to 16/32-byte boundaries, ensuring optimal warps on NVIDIA Tensor Cores and ARM Neon vectors.
Documentation / Architecture / Dual-Branch NMS-Free Head
Dual-Branch NMS-Free Head
Eliminating Non-Maximum Suppression at inference time through one-to-one matched loss during training.
One-to-One vs One-to-Many Dual Training
During training, two branches operate simultaneously:
- Auxiliary One-to-Many Branch: Provides rich gradient supervision during early epochs for rapid convergence.
- Deploy One-to-One Branch: Matches exactly 1 prediction per ground-truth bounding box via Hungarian bipartite matching.
At deployment time, the auxiliary branch is completely discarded with zero cost. Predictions from the one-to-one head are 100% deterministic and require no NMS heuristics.
Documentation / Inference Runtimes / Python SDK Reference
Python SDK Reference
High-level Python API for loading models, batch processing, camera streaming, and coordinate decoding.
Fotonet(weights_path)
Initializes the deploy graph and auto-detects CUDA / CPU availability.
Methods:
model.predict_bgr(image_or_path, imgsz=640, conf=0.25): Primary entrypoint for OpenCV BGR images or file paths.
FastPredictor(model, device="cuda", half=True): High-throughput path with pinned memory and a reusable staging buffer; predict(frame) returns an (N, 6) array of x1, y1, x2, y2, score, class.
model.export(format="onnx", **kwargs): Export engine supporting ONNX, TensorRT, and CoreML.
Documentation / Inference Runtimes / Zero-Copy CUDA Path
Zero-Copy CUDA Fast-Path
Eliminate CPU-GPU memory roundtrips for maximum FPS in camera streaming pipelines.
Pinned Memory & CUDA Stream Binding
By pinning host memory and reusing one staging buffer and one GPU tensor per stream, per-frame allocation and host-device roundtrips disappear, bringing end-to-end latency down toward the pure GPU-pass time (see the Benchmarks section for current measured figures).
Documentation / Inference Runtimes / TensorRT Export
Fotonete TensorRT
Fotonete TensorRT FP16 is the deployment path measured on our RTX 4060 test bench. The engines are compiled from the current root fotonete.pt checkpoint at 640 x 640 with static batch shapes.
What the measured pipeline covers
The figures below are TensorRT stage timings, not PyTorch and not browser ONNX timings. Each engine was measured in a fresh trtexec process with a 1-second warmup, a 10-second timed window, and 10-run averages. TensorRT reports GPU compute, host latency, enqueue, H2D, and D2H separately; these stages can overlap and must not be added together. For this like-for-like comparison, the Fotonete benchmark wrapper applies the existing two-stage top-300 selector and emits [B,300,6] rows in x1,y1,x2,y2,score,class_id order. The underlying Fotonete checkpoint, architecture, and raw production output contract remain unchanged. YOLO26n is an end-to-end checkpoint and its TensorRT output is explicitly verified as [B,300,6]; Ultralytics ignores the generic NMS flag for this end-to-end model because the 300-row output is already part of the graph.
Scope
These are processing-stage measurements for Fotonete. We do not claim individual convolution-kernel timings because those were not measured for this release.
Fotonete final-row TensorRT stage speeds
Measured on the RTX 4060 with TensorRT 10.16.1. Values below are the median and 90th-percentile latency reported by trtexec for the timed window.
| Stage |
Batch 1 median |
Batch 1 p90 |
Batch 8 median |
Batch 8 p90 |
| GPU compute |
0.7670 ms |
0.8096 ms |
3.8289 ms |
4.3126 ms |
| Host latency |
1.1592 ms |
1.2041 ms |
6.9395 ms |
7.4087 ms |
| Enqueue time |
0.2856 ms |
0.3926 ms |
0.4443 ms |
1.0000 ms |
| Input H2D transfer |
0.3887 ms |
0.3926 ms |
3.0840 ms |
3.1108 ms |
| Output D2H transfer |
0.0031 ms |
0.0034 ms |
0.0070 ms |
0.0073 ms |
The timed engine used random input bindings and measured TensorRT execution only. The Fotonete selector is included in the benchmark engine; image decoding and preprocessing are not.
Fotonete vs YOLO26n
Both systems were compiled as static FP16 TensorRT engines at 640 x 640 and measured on the same RTX 4060 using the same trtexec protocol. YOLO26n was exported with max_det=300; both ONNX outputs were shape-checked as [B,300,6]. The comparison is runtime-only and does not change or imply an accuracy claim.
| Metric |
Fotonete |
YOLO26n |
Fotonete difference |
| Batch 1 GPU compute median |
0.7670 ms |
1.5615 ms |
2.04x faster |
| Batch 1 host latency median |
1.1592 ms |
1.7616 ms |
1.52x faster |
| Batch 8 GPU compute median |
3.8289 ms |
4.9224 ms |
1.29x faster |
| Batch 8 host latency median |
6.9395 ms |
6.4883 ms |
0.93x |
| Batch 1 throughput |
810.8 queries/s |
527.8 queries/s |
+53.6% |
| Batch 8 throughput |
1,112.0 images/s |
1,198.8 images/s |
-7.2% |
| Batch 8 output contract |
[8,300,6] |
[8,300,6] |
two-stage top-300 selector |
1-Line CLI Export
fotonet export model=fotonete.pt format=onnx path=fotonete.onnx opset=17
Documentation / Training / Hungarian Dual-Assignment
Hungarian Dual-Assignment Loss
Cost matrix formulation balancing SIoU localization error and focal classification probabilities.
Cost Matching Equation
The bipartite matching cost between prediction $i$ and ground truth $j$ is formulated as:
Cost(i, j) = λ_cls · FL(p_i, c_j) + λ_box · L1(b_i, b_j) + λ_siou · SIoU(b_i, b_j)
Documentation / Training / Resumable Recovery
Resumable Training Protocol
Deterministic interruption-safe recovery capturing optimizer moments, EMA state, LR schedule, and RNG seeds.
Automatic Resume
fotonet train model=fotonete data=my_dataset.yaml resume=fotonet_last.pt
Documentation / Training / Hardware Telemetry Rig
Hardware Telemetry Rig
Empirical hardware specifications and concurrent load disclosures for reproducibility.
Test Bench Specifications
- CPU: 13th Gen Intel Core i3-13100F · 4 cores / 8 threads
- GPU: NVIDIA GeForce RTX 4060 (8 GB VRAM) · driver 610.74
- Memory: 2133 MHz configured memory speed
- Storage: WDC WDS480G2G0C SSD · 2,229 MB/s measured sequential read
- Runtime: Windows 11 · CUDA 12.8 · PyTorch 2.11.0+cu128
- Load: Benchmarks are recorded in a clean process with no training job running. PyTorch latency uses 80 end-to-end and 150 forward iterations per worker; mAP uses the full COCO val2017 validation set. TensorRT uses a separate 10-second
trtexec window after a 1-second warmup. Peer models (YOLO26n and D-FINE N) are evaluated on the same machine; nothing is quoted from vendor pages.