Offline AI: Camera & Chat — my own iOS product, submitted to the App Store

Offline AI — An Assistant That Never Leaves the Phone

What changes when the model lives in your pocket?

Client
Offline AI: Camera & Chat — my own iOS product, submitted to the App Store
My role
Sole designer and developer — product, UX, SwiftUI, on-device model integration
Timeline
2026
Impact
6 on-device language and vision models · 1,215 object labels · 246 seeded knowledge categories · works in airplane mode

The challenge

A cloud assistant hides its constraints: memory, heat and latency are somebody else’s problem. On a phone they are the product. A model that fits in 230 MB is not as capable as one that needs 1.6 GB, the camera and the language model compete for the same chip, and a download that fails at 90% costs the user a gigabyte and their trust. I wanted an assistant that chats, sees, listens and talks with the network switched off, with no account and nothing sent anywhere. The question was not whether a small model can answer. It was what the interface has to do so that a small model, finite memory and a warm phone still feel like a capable assistant.

What I did

1. Answered “will it fit?” before the download. The model manager reads the phone’s memory and free storage and labels every model Fits well, Heavy, Too large or No space before a byte is downloaded. iOS lets one app use roughly half the memory before shutting it down, so a model that technically loads can still be marked Heavy. People pick a model their phone can actually run.

2. Designed when the agent speaks, not just what it says. Agent Vision checks the camera once a second and stays quiet until the same object holds for three seconds. Then it says one sentence of at most fifteen words, with Skip always on screen, and does not narrate the same object again for a minute. A narrator that talks constantly is a narrator people switch off.

3. Gave vision and language their own turns. Recognition and the language model share one chip, and running both at once froze the camera feed in testing. Camera analysis now pauses while the model thinks and speaks, and resumes when the sentence ends. The scheduling became part of the interaction: the agent looks, then talks.

4. Let people teach it. When an object is stable but unfamiliar, the agent asks by voice what to call it and listens for the answer. The name is stored on the phone with the object’s visual signature and a node in the knowledge graph, so next time it recognises this particular mug, not just “a mug”. It asks once, then leaves that object alone for three minutes.

The solution

A multimodal assistant designed around the device. Chat streams from a local LFM2 model through Liquid AI’s LEAP SDK. MobileCLIP matches camera frames against 1,215 labels. Agent Vision narrates and learns names, and a knowledge graph starts with 246 seeded categories and grows with what you teach it. The interface is native SwiftUI with system controls. After a model is downloaded, airplane mode changes nothing. The code is open source and version 1.0 is in App Store review, so there are no adoption numbers to report yet.

What I’d do differently

I started from what the models could do and treated the device as an engineering detail. Fit labels, one-sentence narration and turn-taking between vision and language all arrived later, as fixes for memory, latency and heat. They are the product, not patches. Next time I would put the device limits into the first sketch and test the whole flow on a 4 GB iPhone before designing anything for a Pro.

More workAll case studies →