← All work

Open-source app · Built by Bytesicht

opnlocal

A private AI app for people who have never heard of VRAM or quantization. It checks the device, suggests a model that fits, measures how fast it runs, and keeps every conversation on the device.

Product design · Desktop & mobile app · On-device AI runtime

Visit opnlocal Source code
opnlocal suggesting two models for everyday help on this device, each marked as running comfortably, with download size and free space needed
Model suggestions · Windows PC with a 16 GB graphics cardView full size

Private AI,
without the jargon.

Running a language model on your own computer usually starts with choices most people cannot make: which model, which size, which quantization, and whether it will fit at all.

opnlocal checks the memory, graphics, processor, and free space first. It asks what the person wants to do, suggests up to three models in plain words, downloads the one they choose, measures how fast it really runs, and opens a chat. Conversations, documents, and images stay on the device.

Bytesicht designed and built the app, the Rust engine behind it, the signed model catalog, and the releases for Windows, macOS, Linux, Android, and iPhone.

opnlocal benchmark result measured on the device: replies at about 14 words per second, with prompt speed, context size and memory use in the details
Speed measured on this device · Qwen3.5 9BView full size

How the app is put together

One engine, five platforms.

  1. Svelte interface

    One interface for desktop and phone, shown in the system web view through Tauri 2.

  2. Rust engine

    Detection, fit estimates, downloads, benchmark, chat, and storage in one crate with no UI dependency.

  3. llama.cpp in the app

    The model runs inside the app process: Vulkan on Windows and Linux, Metal on Apple devices, the processor on Android.

  4. Signed model catalog

    An ed25519-signed model list, bundled with the app and refreshed at most once a day.

The interface reaches the engine through Tauri's in-process calls. There is no local server and no open port, so there is no local API to secure. This is a component overview, not a deployment topology.

Engineering decisions

Honest numbers.
Safe defaults.

Fit before speed.

Before a model runs, the app only says whether it fits and where it runs: graphics chip or processor. Speed appears after a short benchmark on the device, labelled as measured, and counted from the words the model actually produced.

Every file is checked.

Models download from a pinned Hugging Face commit. An interrupted download resumes where it stopped, and the file is kept only if its sha256 matches the catalog. The catalog itself is accepted only with a valid signature and a newer version than the one installed.

A failed graphics load falls back.

A marker file is written before any model is loaded onto the graphics chip. If the app stops during that load, the next launch uses the processor and says why, instead of failing the same way again.

Long documents are cut down, not refused.

The context grows with the conversation. A long document makes the engine reload the model with the smallest context that holds it and still fits. Beyond that, plain word matching (BM25) keeps the parts most relevant to the question, and the chat says so.

opnlocal chat with a reply generated on the device, noting that nothing is sent anywhere
Chat running on the device · Unedited model replyView full size

Built with

Rust / llama.cpp / Tauri 2 / Svelte 5 / TypeScript / Vulkan / Metal

Current status

Open-source app under the Apache-2.0 license, version 0.2.2, released in September 2026. The full journey has been tested on a Windows PC and an Android emulator; macOS, Linux, and iOS builds are verified on build machines and simulators so far. No usage figures are claimed.

Build with Bytesicht

Put AI where the constraints are.

From choosing a model to running it on the device and shipping it on several platforms, we can build AI features around the limits they actually run under.

Discuss your project