Skip to content

/lab/perception

03

Perception

MACHINE LEARNING / COMPUTER VISION

Teach the system to see.

HYPOTHESIS

A real hand-landmark model can run entirely inside a browser tab, fast enough to feel like direct manipulation, without any frame ever leaving the device.

HOW TO OPERATE IT

With camera access explicitly granted, a hand-landmark model tracks 21 points on one hand in real time and drives an on-screen target. Camera access is never requested until this experiment is opened, and a pointer/touch-driven mode with the same interaction is always available, camera or not.

WHAT WAS DIFFICULT

Designing a fallback path that isn't a downgrade in dignity — the pointer mode had to be a first-class version of the same idea, not an apology screen, since a meaningful number of visitors will never grant camera access.

WHAT IT TAUGHT ME

Inference confidence is itself worth surfacing honestly — showing the model's actual detection confidence, rather than hiding uncertainty behind a smooth-looking cursor, is what makes the experiment read as a real ML system instead of a magic trick.

TECHNOLOGY

  • MediaPipe Tasks Vision (WASM, in-browser)
  • getUserMedia
  • Canvas 2D overlay

REAL SKILLS THIS DRAWS ON

  • Applied ML in the browser
  • privacy-respecting sensor design
  • graceful degradation
BUILD SOMETHING