Face verification that runs on the device
A Next.js PWA wrapped for Android with Capacitor: face detection and cropping happen on the phone, while a FastAPI backend matches with ArcFace and Qdrant.
- Next.js
- Capacitor
- MediaPipe
- FastAPI
- Python
A face verification app written in Next.js, running as a PWA on the web and shipped as an Android APK through Capacitor, with its own FastAPI backend doing the matching.
What stays on the phone, what goes to the server#
The first question I had to answer was where the boundary sits. I ended up with two quite different flows.
The home screen does 1-to-1 comparison and runs entirely on the device: pick a reference photo, take a live one, let face-api.js extract a descriptor from each and compare the distance. Nothing leaves the phone.
Enrolment and identification are 1-to-N, so they need a server to search the whole set of
registered people. Here I send images, not embeddings: multipart JPEGs to
/api/register (five photos per person) and /api/identify (one photo). The reason is
plain — an embedding is only comparable to embeddings from the same model, and the model
running in the browser is not the ONNX model on the server. Uploading a face-api.js
descriptor to compare against ArcFace would be meaningless.
The device still does the part that matters for data, though: what gets uploaded is a tight crop of the face, not the whole frame.
Frame first with MediaPipe, shoot second#
The camera screen uses the FaceDetector from @mediapipe/tasks-vision (BlazeFace
short-range, GPU delegate) in VIDEO mode, driven by requestAnimationFrame. On each frame
I check whether the face bounding box sits fully inside the oval drawn on screen; if it
does, a three-second countdown starts, and the shot only fires if the user holds still for
all three seconds.
const detector = await FaceDetector.createFromOptions(vision, {
baseOptions: { modelAssetPath: BLAZE_FACE_SHORT_RANGE, delegate: 'GPU' },
runningMode: 'VIDEO',
})
const detectFace = () => {
const { detections } = faceDetector.detectForVideo(video, performance.now())
const box = detections[0]?.boundingBox
/* The face must sit inside the oval and stay there for COUNTDOWN_TIME seconds. */
if (box && isInsideFrame(box)) {
alignStartTimeRef.current ??= Date.now()
if (Date.now() - alignStartTimeRef.current >= COUNTDOWN_TIME * 1000) takePhoto()
} else {
alignStartTimeRef.current = null
}
requestAnimationFrame(detectFace)
}Two feature spaces, two metrics#
On device, face-api.js gives a 128-dimension descriptor and I compare with Euclidean distance against a 0.4 threshold. A few gates run before that: exactly one face per image, a detection score of at least 0.75, and a bounding box whose shorter side is at least 100px. A blurry photo or a face too far away is rejected early with a reason, instead of producing a distance nobody should trust.
The detector is also chosen per platform via Capacitor.isNativePlatform(): MTCNN on the
web, SSD MobileNet v1 on native. MTCNN picks up small faces better but costs noticeably
more, and inside a webview that cost is visible.
Server side: ArcFace and Qdrant#
The FastAPI backend has exactly two working endpoints. An incoming image is decoded with OpenCV, resized to 112×112, normalised the way ArcFace expects, then run through onnxruntime on CPU.
def extract_face_embedding(self, image_data: bytes):
img = cv2.imdecode(np.frombuffer(image_data, np.uint8), cv2.IMREAD_COLOR)
if img is None:
return None
# ArcFace preprocessing
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
img = cv2.resize(img, (112, 112))
img = (img.astype(np.float32) - 127.5) / 128.0
embedding = self.session.run(
[self.output_name], {self.input_name: img[np.newaxis, ...]}
)[0][0]
# L2 normalise so Qdrant's cosine score means something
norm = np.linalg.norm(embedding)
return (embedding / norm).tolist() if norm > 1e-6 else embedding.tolist()The 512-dimension vector goes into a Qdrant collection configured with
Distance.COSINE. Enrolment writes five separate points for one person rather than
averaging them: each angle is its own point and the search only needs the nearest one.
Identification calls query_points with limit=1 and a score_threshold the client
sends (0.8 by default); below the threshold it returns "unknown" instead of guessing.
Packaging for Android#
next.config.js exports a static build into out/, which is the directory Capacitor
reads. The interesting bit is the service worker: next-pwa is disabled outright for the
Android build, because inside a native webview the assets already ship in the APK and a
second caching layer only creates version drift. The face-api.js models are copied from
public/models into out/models by an afterEmit webpack hook, so they travel with the
bundle. Camera permission is requested through @capacitor/camera, and the native path
uses @capacitor-community/camera-preview for a real preview surface instead of a
<video> tag.
Outcome#
- The 1-to-1 flow runs fully on the device; the 1-to-N flow uploads only a tight face crop
- A backend of exactly two endpoints, 512-dimension embeddings in Qdrant, cosine matching with a client-adjustable threshold
- No anti-spoofing yet: a printed photo or a phone screen held up to the oval still passes, because the three-second hold checks framing stability, not a live person
- MediaPipe's wasm and tflite assets are still fetched from a CDN, so the first camera launch needs network — unlike the face-api.js models, which ship inside the bundle
- The two endpoints have no authentication, and the 0.4 / 0.8 thresholds are hand-tuned: I have no evaluation set yet to say which numbers are right