A demo of what’s possible.
Computer vision helps computers extract useful information from images and video.
This isn’t a production solution to one specific problem. It’s a playground for seeing what computers can recognize, learning a few core ideas, and asking a more useful question:
What would be valuable for a computer to notice?
The footage is recorded. The detections are live.
Nothing is scripted to match the demo videos. Object detection is happening as they play.
Switch to Camera and point it at the world around you. The same system will analyze the live camera feed and recognize supported objects in real time.
Once everything is loaded, that work happens directly on your device and can continue offline.
Not everything needs a giant server.
Modern devices contain hardware built to perform huge amounts of math in parallel.
Graphics processors, or GPUs, became powerful because graphics and games require many calculations to happen at the same time. That same capability happens to be extremely useful for machine learning, including computer vision and generative AI. Many newer devices also include specialized chips designed specifically for this kind of work.
That creates an important option: sometimes the device connected to the camera can do the work itself.
Local processing can mean faster responses, less network traffic, better privacy, offline operation and less dependence on centralized computing.
Cloud computing is still the right answer for plenty of problems. But “Can this happen where the data is created?” should be a serious design question.
Good design can require less compute.
A useful experiment in this demo is hiding in the moving boxes.
Many screens refresh around 60 times per second, but this demo performs object detection only 8 times per second by default. The interface smoothly transitions the boxes between those observations instead of running another expensive detection pass for every displayed frame.
That cuts the number of detection passes by more than 85% compared with analyzing every displayed frame, while the experience still feels fluid.
The larger lesson is simple: the model is only one part of the system.
Choices about hardware, software, timing and presentation can sometimes make the difference between needing constant cloud connectivity and making a small, self-contained, offline-capable setup practical.
What it recognizes is only the starting point.
This demo uses a general-purpose model that already recognizes common objects, so we cherry-picked a few that demonstrate well: produce, people and cars.
Other computer vision systems can be specialized around a particular need—a product, defect, gesture, shelf condition, piece of equipment or countless other visual signals.
Recognition can then become one input among many: tracking, business rules, sensors, other data or generative AI.
What is here?
What changed?
Does it matter?
What should happen next?
Seeing is only the beginning.
Technical details
- Task
- Real-time object detection
- Model
- SSDLite MobileNet V2
- Training dataset
- COCO
- Model size
- ~18.6 MB
- Runtime
- TensorFlow.js 4.22
- Execution
- Browser-local
- Acceleration
- WebGL / GPU
- Internal model input
- 300 × 300 px
- Default inference rate
- 8/sec
- Adjustable inference rate
- 2–20/sec
- Presentation
- DOM overlays rendered independently from inference
- Motion
- Box position and size interpolated between detections
- Alignment
- Adjustable playback delay for visual synchronization
- Demo classes
- Apples, bananas, oranges, people and cars
- Connectivity
- No server inference; offline-capable after assets are cached
- Camera
- Live front/rear camera support where available
This is a learning demo, not a production vision system. General-purpose models can miss objects, confuse them or struggle when things are small, moving, obscured or outside what they were trained to recognize.
Media and software credits