Runs entirely client-side: ONNX via onnxruntime-web (WASM) or TensorFlow.js -- no server, CPU only. Checkpoints are fetched on demand from azemel/retinaface-xs.