Object detection running in the browser
Training YOLOv5 on 521 self-labelled images across six object classes, then running inference in the browser two ways: LiteRT.js and ONNX Runtime Web.
- YOLOv5
- PyTorch
- ONNX Runtime Web
- LiteRT.js
- React
I photographed and labelled my own set of desk objects, trained YOLOv5s on it, then took
the model and ran it entirely in the browser. The repo is split into three parts —
train, yolov5-litert, yolov5-onnxruntime-web — because the real point was never the
model, it was comparing two web runtimes on the same model.
The self-labelled dataset#
521 phone photos in total, hand-labelled in YOLO format across six classes: ban phim,
but long, cay keo, con chuot, doi dua, tai nghe (keyboard, marker, scissors,
mouse, chopsticks, headphones). A small script shuffles the list with random.seed(42)
and splits 80/20 into 416 training and 105 validation images.
A dataset you build yourself always carries junk. Running val.py, YOLOv5 reported eight
corrupt JPEGs it had to repair and dropped one image entirely because a label
coordinate fell outside the image (1.0912 instead of ≤ 1) — so the real evaluation ran on
104 images and 270 instances.
Training and exporting to two formats#
Training from the yolov5s.pt checkpoint on Colab with a single Tesla T4: 640px images,
batch 8, up to 300 epochs, with --patience 30 to stop once it stops improving.
!python train.py --img 640 --batch 8 --epochs 300 \
--data ../data.yml --weights yolov5s.pt \
--cache images --patience 30
!python export.py --weights runs/train/exp/weights/best.pt --include tflite --device cpu
# simplify folds away redundant operators, keeping the file lean for onnx runtime web
!python export.py --weights runs/train/exp/weights/best.pt --include onnx --simplifyOn the validation set the model reached precision 0.810, recall 0.827, mAP50 0.843 and
mAP50-95 0.508. cay keo scored best at mAP50 0.910, but long worst at 0.724 — which
makes sense, since a thin marker blends into the desk behind it.
LiteRT.js makes me move tensors between devices myself#
The LiteRT build initialises in two steps: load the WASM runtime first, then compile the
model with the webgpu accelerator. The biggest difference from ONNX is that a tensor
does not land where it is needed — I have to push it to WebGPU before the run and pull it
back to WASM afterwards.
await loadLiteRt(WASM_PATH);
const model = await loadAndCompile(MODEL_PATH, { accelerator: "webgpu" });
// TFLite expects NHWC layout, not the NCHW the ONNX build uses.
const tensor = await new Tensor(inputData, [1, IMG_SIZE, IMG_SIZE, 3]).moveTo("webgpu");
const output = model.run([tensor])[0];
// The result sits on the GPU; it has to come back to WASM before JS can read it.
const cpuOutput = await output.moveTo("wasm");
const rawBoxes = parseYoloOutput(cpuOutput.toTypedArray(), classes);Preprocessing here is plain JavaScript: draw the frame into a 640×640 canvas, call
getImageData, then loop by hand to drop the alpha channel and divide by 255. No external
dependency at all — but also no letterboxing, so the image is stretched straight into a
square.
ONNX Runtime Web is heavier but preprocesses correctly#
The ONNX build runs on the wasm execution provider with simd and threads enabled,
plus one warmup pass on an empty tensor so the first real inference does not stall. The
.onnx file after --simplify is 27.2 MB, so I fetch it with an XMLHttpRequest that
reports onprogress as a percentage instead of leaving the user on a blank screen. Craco
also needs copy-webpack-plugin to get onnxruntime-web's .wasm files into the build
output.
In exchange, preprocessing is clearly better: OpenCV.js pads the image to a square with
copyMakeBorder and only then resizes to 640×640 and scales by 1/255 through
blobFromImage. The source aspect ratio survives, and that same pad ratio maps the boxes
back onto real image coordinates.
NMS still lives in JavaScript#
Both models were exported without NMS, so the output is a raw (1, 25200, 11) tensor
and both web builds walk all 25,200 anchors and run a hand-written NMS. Their thresholds
have drifted apart too: LiteRT filters at 0.45 and suppresses per class, while ONNX
filters at 0.25 with a 0.2 class threshold, suppresses class-agnostically and caps at
topk 100. In camera mode I throttle inference to 250 ms so the requestAnimationFrame
loop cannot stack runs on top of each other.
Outcome#
- The six-class model reaches mAP50 0.843 on 104 validation images — good enough for a demo, but still a number produced by a 521-image dataset
- LiteRT.js is lighter on dependencies: only
@litertjs/core, straight to WebGPU, no bundler config — at the cost of managing tensor placement and writing every preprocessing step by hand - ONNX Runtime Web is the safer bet since it only relies on WASM and gets letterboxing
from OpenCV.js, but it drags along a 27.2 MB download plus webpack config for
.wasm - I have not measured in-browser inference time on either build; the only timings I have are 9.0 ms inference and 5.3 ms NMS per image on the T4, so I cannot yet say which runtime is faster on the web