Local LLMs on device
Quantized models running on phones and edge hardware with embedded vector stores. Private, offline, instant.
I'm Rich Olson. I run language models on phones, agents on persistent loops, and all of it on hardware I rack myself — and this site teaches you to do the same. Live exhibits you can touch, tools you can download, field notes from the edge of local and frontier AI. Nothing for sale.
Six tools I reach for, end to end — from a quantized model running offline on a phone to a patented live-video pipeline.
Quantized models running on phones and edge hardware with embedded vector stores. Private, offline, instant.
Multiple LLMs on persistent task loops with custom tooling — in production apps and in test rigs for features still in development.
Custom embeddings over your data, retrieval that cites its sources, tuned to run where the data lives.
Patented live-video technology, real-time rooms, RTSP pipelines and low-latency delivery that holds up.
Web-view apps calling native via the Capacitor bridge — background downloads, full offline, Socket.IO across apps.
My own prompt-to-production platform — describe the site, and it's built, hosted, and live on infrastructure I run myself. Build it, host it, ship it.
A private, browser-based conferencing platform — Zoom-scale in scope, friction-free in practice. Built on streaming techniques I hold a US patent on — granted 2026, sole inventor.
Tools leave the shop when they're ready — documented, tested, and worth your time. Three are out now. More are on the bench.
A native Socket.IO client plugin for Capacitor — robust, CORS-free real-time comms on iOS and Android.
Tools and a Codex skill for keeping long agent threads healthy — health checks, handoffs, recovery.
Async check for whether an RTSP stream is live, using ffprobe to read codec and duration info.
My favorite work: local LLMs running on phones and devices, answering from custom data embedded in vector databases. Private, offline, instant.
Try it — this harness runs entirely in your browser. No network, retrieving from a tiny store of facts about me.
I run my own DNS, mail, video, hosting, and inference — on hardware I rack myself — because intelligence you depend on should be something you control.
If it's on this site, it's running. Nothing here is a mockup.
I'm heads-down on my own AI projects — a few you can see in the exhibits here, and a few I'm not talking about yet. I'm not taking client work, with one exception: a problem interesting enough to earn a spot on the bench.
If yours is that problem: the person you email is the person who architects it, builds it, and ships it — no account managers, no hand-offs, no juniors learning on your dime.
Not right for me? I'll tell you quickly — and where I can, point you to someone who is.
The bench is full and I'm not looking for work — but a genuinely interesting problem can still earn the exception. Tell me why yours is the one. It comes straight to me, and I'll tell you quickly either way.