oMNI Chat
A local-first, multi-user AI chat application that runs against any OpenAI-compatible backend. This manual covers everyday use and the configuration that operators need to run it.
Chapter 1
Introduction
Omni Chat is a lightweight chat application built around a single Go core. There is no
database server to install and no external service to manage — conversations, projects,
documents, and memory all live in a local SQLite file. The core talks to any endpoint that
speaks the OpenAI API (/v1/chat/completions, /v1/embeddings,
/v1/models), so it works with local engines like llama.cpp and vLLM, a LiteLLM
proxy, or hosted providers.
Local-first
One binary, local SQLite storage with vector search. Your data stays on the machine running the core.
Multi-user
OS-login style accounts with private workspaces. Projects can be marked shared for a household or team.
Any backend
Point it at one endpoint, or combine several backends so all their models appear in one picker.
Memory & sources
Optional durable memory of facts and decisions, plus retrieval over documents you upload.
Native Apple apps automatically select AppleFoundation as the auxiliary model on iOS 26 / macOS 26 and later, including OS 27. Apple Intelligence must be enabled and its model downloaded in system Settings. This works independently of the Qwen chat-model download. Under Settings → Models → Auxiliary, choose AppleFoundation, Managed (the Local services llama.cpp), a connection + model, or Disabled; use Test to check availability. Existing custom connections and explicit disabled settings are preserved. A dedicated router classifier takes precedence. On-device Qwen 3.5 2B displays reasoning in the Thinking block and allows opt-in Python/JavaScript code execution through the client sandbox.
Two shells, one core
The same Go core powers two front ends:
- Web app — runs in any browser over HTTP. This is the supported way to run Omni Chat today.
- Desktop app — a native Tauri shell that bundles the core as a sidecar and opens it in its own window. Apple Silicon Macs running macOS 26 or later; see Chapter 12.
- iPad app — the Tauri shell links the same Go core into the application (iPadOS cannot spawn a sidecar), serves the same API on loopback, and can run a verified Qwen 3.5 2B model through MLX on iPadOS 26 or later on M-series iPads with at least 8 GB RAM. Its context is 16,384 tokens, room for web-search results and a longer conversation. Simulator is remote-only. How much the model can write in one reply scales with the device — roughly 4,000 tokens on an 8 GB machine, 6,000 on a large Mac — and is shared between its thinking and its answer, so a model that thinks for too long is stopped and asked again without thinking rather than running out of room mid-reply. Answers stream as they are written, including on turns that can use tools, and the answer after a web search starts quickly because the model does not re-read the whole conversation. The model loads from the downloaded snapshot; it does not compile on the device. The 2B's thinking is capped so a turn where it keeps second-guessing itself is answered within seconds. Apple Silicon Macs on macOS 26 or later use the same model controls and run Qwen 3.5 4B (16 GB of memory or more) or Qwen 3.5 2B (8 GB) with a 32,768-token context. Weights from earlier versions — the GGUF download or the Mac's retired Core AI bundle — are listed as old weights; use Remove old weights to reclaim their space. Close any main-window on-device model notice with the × in its top-right corner; this is remembered for your account on this device. The notice returns if what it says changes — a different reason, download state or error — and model controls appear when a verified bundle becomes available. Open Settings → Models → On-device model at any time to download, resume, pause, or remove the model, even after dismissing the notice. These actions apply immediately; Save & Restart is not needed.
This manual describes how to use and configure the app. Build and process notes live in
AGENTS.md.
Chapter 2
Getting Started
The fastest way to run Omni Chat is Docker. This chapter walks through everything end-to-end: getting an AI backend running if you don't have one yet, picking how much of the optional stack (web search, URL fetch, the auxiliary model, smart routing) fits your hardware, and starting the app. Building from source instead? Skip to Local Development Setup.
What you need
Docker (Docker Desktop, or Docker Engine + the Compose plugin on Linux) and
access to an OpenAI-compatible AI backend — something that answers
/v1/chat/completions. That can be:
- Something you already have running (llama.cpp, vLLM, a LiteLLM proxy, an existing hosted provider) — skip ahead to Running Omni Chat.
- A model you set up locally right now with Ollama or LM Studio — see below.
- A hosted/cloud endpoint with no local model at all — see Or use a hosted endpoint instead.
Setting up a local AI backend
Ollama
Install: curl -fsSL https://ollama.com/install.sh | sh (Linux),
brew install ollama or the installer from
ollama.com/download (macOS), or the Windows
installer from the same page.
ollama pull llama3.2
ollama run <model> also starts the server if it isn't already running.
Ollama serves an OpenAI-compatible API on port 11434 with no
/v1 in the base URL:
# Docker
CHAT_ENDPOINT_URL=http://host.docker.internal:11434
# Host-run (make dev)
CHAT_ENDPOINT_URL=http://localhost:11434 make dev
LM Studio
Download from lmstudio.ai, search for and download a model, load it, then open the Developer (Local Server) tab and click Start Server. LM Studio's default port is 1234:
# Docker
CHAT_ENDPOINT_URL=http://host.docker.internal:1234
# Host-run (make dev)
CHAT_ENDPOINT_URL=http://localhost:1234 make dev
Either way, the model you pulled/loaded shows up in Omni Chat's Model dropdown automatically
— it's fetched live from the backend's /v1/models.
Or use a hosted endpoint instead
No GPU, or you'd rather not run a local model at all: point Omni Chat at a cloud endpoint. This is also the recommended path at 8GB of memory or less — see How much memory do you have? below.
- Ollama Cloud — an OpenAI-compatible hosted endpoint with no local
install. Create a key at
ollama.com/settings/keys (a free tier
exists), then:
Pick from cloud-hosted models likeCHAT_ENDPOINT_URL=https://ollama.com CHAT_API_KEY=<your ollama.com key>gpt-oss:120b,qwen3-coder:480b, ordeepseek-v3.2in the Model dropdown — no local GPU or RAM spent on inference. - RouteLLM — an open-source router
(lm-sys/RouteLLM) that fronts a cheap/fast
model and a strong/expensive one and exposes its own OpenAI-compatible endpoint, routing each
request to control cost. Self-host it and point
CHAT_ENDPOINT_URLat wherever it listens (hosted RouteLLM-compatible options also exist). This is a different routing axis than Omni Chat's own Auto (Smart Router) — RouteLLM chooses between two backends for cost; Auto chooses among your configured backends for task fit — and the two can be combined. - Any other hosted OpenAI-compatible provider you already use — just set
CHAT_ENDPOINT_URLandCHAT_API_KEY.
How much memory do you have?
Search and Fetch are lightweight: SearXNG is CPU/RAM only, and the Fetch reader sidecar is a
headless-browser process — neither touches VRAM or unified memory, so they're safe to add
regardless of hardware. The chat model itself, plus the optional Auxiliary
model (recommended: LiquidAI/LFM2.5-1.2B-Instruct, ~730MB GGUF), the
Router classifier (recommended: katanemo/Arch-Router-1.5B,
~1GB GGUF), and the Image generation sidecar (see the note below the table)
all compete for the same VRAM/unified-memory pool — size your setup by how many of
those you plan to run at once.
| Available VRAM / unified memory | What fits | Recommended path |
|---|---|---|
| ≤ 8 GB | One small chat model (1–3B, Q4) or no local model at all | Barebones Compose + a small local model (e.g. Llama 3.2 3B, Qwen2.5 3B), or skip local inference and use Ollama Cloud / RouteLLM / a hosted provider above. Search and Fetch overlays are still fine to add. |
| 8–16 GB | A 7–8B chat model (Q4) comfortably | Barebones, or + Search/Fetch freely. Add the Auxiliary-model sidecar once you're past ~12GB. |
| 16–32 GB | A 7–14B chat model and the Auxiliary-model sidecar together | + Search + Fetch + Aux. |
| 32 GB+ | 14B+ chat model plus both the Auxiliary-model and Router-classifier sidecars | The full stack: + Search + Fetch + Aux + Router. |
These are starting points. Actual usage depends on quantization and context length — watch your system monitor the first time you load a new model.
The image sidecar's recommended model, Tongyi-MAI/Z-Image-Turbo (a fast
8-step 6B model, ~12 GB download), wants roughly 16 GB of VRAM on
its own in bf16 — and setting IMAGEGEN_EDIT_MODEL_ID for precise instruction
edits keeps a second, typically larger model resident
(Qwen/Qwen-Image-Edit-2509, the reference open editor, is ~20B — plan for
40–60 GB on top). Plan for it the way
you'd plan for a second large chat model: run it on a dedicated GPU box with
compose.imagegen-only.yaml when your main machine is already busy with chat
inference, or point CHAT_IMAGE_GEN_URL at a hosted OpenAI-compatible images
endpoint (e.g. gpt-image-1) and spend no local memory at all. All of these
model choices ship as ready-to-uncomment examples in .env.example.
Running Omni Chat: pick your combination
Each step below just adds one more -f flag to the previous command — start
barebones and layer on whichever of these your memory tier and needs call for. What each
overlay does in depth (healthchecks, image pinning, reader vs. direct fetch, running any sidecar
without Compose) is covered in Administration & Deployment;
this is just the fast path.
# Once, before any of the combinations below
cp .env.example .env
a. Barebones — chat only, no web search:
docker compose up -d
b. + Search — adds self-hosted web search (SearXNG):
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.yaml -f compose.search.yaml up -d
c. + Fetch — adds JS-page rendering and YouTube-transcript support for the
fetch_url tool:
docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml up -d
d. + Auxiliary model — adds a small dedicated model for chat titles, memory jobs, and (with no dedicated classifier configured) smart-router classification:
docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml \
-f compose.aux.yaml up -d
e. + Auxiliary model + Router classifier — the full stack; adds a dedicated routing model that outranks the auxiliary model for smart-router classification:
docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml \
-f compose.aux.yaml -f compose.router.yaml up -d
f. + Agent handoff & SFTP — lets the assistant delegate real coding tasks to a containerized agent (each conversation gets a private workspace) and gives users SFTP access to their files. The core runs each approved delegation as a sibling container through the host's Docker socket, so this overlay needs a one-time setup before the first run.
1. Build the agent image. This is the container each delegation runs in
(Pi plus git):
docker build -t omni-agent-pi docker/agent-pi
2. Create the shared workspace directory. Managed workspaces live here, and
the same absolute path must exist on the host, in .env, and in
agents.json — because the host Docker daemon resolves the sibling container's
mounts against host paths. On Linux:
sudo mkdir -p /srv/omni-agent-work
sudo chown 10001:10001 /srv/omni-agent-work # 10001 is the core container's user
On macOS (Docker Desktop) the /srv path is on the sealed
read-only system volume, so put the directory under your home instead — and skip the
chown, since Docker Desktop maps ownership automatically:
mkdir -p ~/omni-agent-work
3. Add the agent config. Copy the example and set managed_root
to the directory from step 2 (a literal absolute path — no ~, which the container
cannot expand):
cp docs/agents.compose-example.json config/agents.json
The example is managed-mode with a Docker-runtime Pi agent and SFTP enabled. On macOS, edit
managed_root to e.g. /Users/you/omni-agent-work.
4. Pre-generate the SFTP host key. The core runs as uid 10001, so a key
you create yourself must be owned by that user (no chown needed on macOS).
When ./config is writable by uid 10001 the server can also create
sftp_host_key on first start:
ssh-keygen -t ed25519 -f config/sftp_host_key -N ""
sudo chown 10001 config/sftp_host_key # Linux only
5. Point .env at the workspace — this value must match
managed_root exactly:
echo 'OMNI_AGENT_WORK_DIR=/srv/omni-agent-work' >> .env # macOS: /Users/you/omni-agent-work
6. Bring the stack up with the overlay (stackable with the others). On
Linux, pass the group that owns the Docker socket so the non-root core may use it; on macOS
Docker Desktop, omit DOCKER_GID entirely:
DOCKER_GID=$(getent group docker | cut -d: -f3) \
docker compose -f compose.yaml -f compose.agents.yaml up --build -d
docker compose -f compose.yaml -f compose.agents.yaml up --build -d # macOS Docker Desktop
SFTP publishes on 127.0.0.1:2222 by default; set
OMNI_SFTP_PUBLISH_ADDRESS=0.0.0.0 in .env to reach it from the LAN.
Users sign in with their Omni Chat username and password and see only their own chats'
folders — sftp -P 2222 you@host, or the VS Code
SSH FS extension pointed at sftp://you@host:2222.
See Getting at your files (SFTP) above for client details.
g. + Image generation — lets the assistant create images ("generate an image of a cyberpunk cat riding a skateboard") with a bundled Z-Image-Turbo sidecar. It requires an NVIDIA GPU (with the NVIDIA Container Toolkit installed; ~16 GB of VRAM) and downloads ~12 GB of model weights on first start:
docker compose -f compose.yaml -f compose.imagegen.yaml up --build -d
Watch docker compose logs -f imagegen on the first run — the app waits for the
sidecar's healthcheck, which stays starting until the weights are downloaded and
loaded. The default model is Tongyi-MAI/Z-Image-Turbo; image edits fall
back to img2img re-imagining with those same weights unless you set
IMAGEGEN_EDIT_MODEL_ID in .env to a dedicated instruction-edit model
— see the verified choices and their (large) sizes in .env.example and
the memory note above. No NVIDIA GPU on this machine? Run
compose.imagegen-only.yaml on a GPU box
and set CHAT_IMAGE_GEN_URL=http://gpu-host:8001 in .env
instead — or point it at any other server that speaks the OpenAI Images API. See
Generating images for how it behaves in chat.
Jetson/L4T or arm64 CUDA hosts: the default base image is x86-64 CUDA, so
override IMAGEGEN_BASE_IMAGE in .env with a CUDA-enabled PyTorch image
matching your host's driver stack — torch is not reinstalled by the build, so it must
come from the base — and set IMAGEGEN_RUNTIME=nvidia to expose the GPU. For example
IMAGEGEN_BASE_IMAGE=dustynv/pytorch:2.7-r36.4.0 on a Jetson (match your JetPack),
or nvcr.io/nvidia/pytorch:25.09-py3 for arm64 CUDA (DGX Spark / Grace-Hopper) —
the GB10 (sm_121) needs a CUDA ≥ 12.9 base, or edit-model JIT kernels fail with
nvrtc: invalid value for --gpu-architecture.
Containers on macOS never see the GPU (the overlay's NVIDIA reservation fails, and CPU
diffusion is unusably slow), but the sidecar runs fine directly on the Mac's GPU via PyTorch
MPS. From the repo: cd docker/imagegen && python3 -m venv .venv &&
source .venv/bin/activate && pip install -r requirements.txt torch && uvicorn app:app
--port 8001 — it auto-detects MPS and downloads the weights on first start. Then set
CHAT_IMAGE_GEN_URL=http://localhost:8001 for a host-run core, or
http://host.docker.internal:8001 in .env when the core itself runs
under Compose. All the IMAGEGEN_* knobs from .env.example apply as
plain process environment variables when running natively — .env only feeds
Docker Compose — e.g.
IMAGEGEN_EDIT_MODEL_ID=Qwen/Qwen-Image-Edit-2509 uvicorn app:app --port 8001.
h. + Read aloud (text-to-speech) — lets the assistant's responses be
read out loud with a bundled OmniVoice sidecar (docker/tts/,
Omni Chat's own OpenAI-compatible wrapper around the k2-fsa OmniVoice model).
No GPU required — the sidecar auto-detects its device: on CPU it drops to
8 inference steps (usable, though slower than real-time), and a CUDA GPU runs full
32-step quality much faster than real-time. On a Mac, run the sidecar natively
instead of in Docker — containers cannot use the Mac's GPU, but the native run
picks Apple-Silicon MPS and delivers full quality at nearly real-time speed (from the
repo: cd docker/tts && python3 -m venv .venv && source .venv/bin/activate &&
pip install -r requirements.txt torch && uvicorn app:app --port 8880, then point
CHAT_TTS_URL at it):
docker compose -f compose.yaml -f compose.tts.yaml up --build -d
The first start downloads the OmniVoice model weights into the tts-models
volume — watch docker compose logs -f tts. For CUDA set
TTS_BASE_IMAGE to a torch-CUDA image that also ships torchaudio
(omnivoice needs it) and TTS_RUNTIME=nvidia in .env — the same
switch on every CUDA host: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime
on x86-64, dustynv/pytorch:2.7-r36.4.0 on Jetson/L4T (match your JetPack), or
nvcr.io/nvidia/pytorch:25.09-py3 on arm64 NGC (DGX Spark / Grace-Hopper; GB10
needs CUDA ≥ 12.9). The overlay passes NVIDIA_VISIBLE_DEVICES
through so the GPU is exposed on all three with no per-host deploy.devices block
(TTS_DEVICE is optional — auto-detect picks cuda). To run the sidecar on
another machine (or beside a host-run core) use compose.tts-only.yaml and set
CHAT_TTS_URL=http://host:8880 plus
CHAT_TTS_MODEL=omnivoice; any other OpenAI-compatible speech server works too —
e.g. the hosted OpenAI API with CHAT_TTS_URL=https://api.openai.com,
CHAT_TTS_MODEL=gpt-4o-mini-tts, and CHAT_TTS_API_KEY. See
Read aloud for how it behaves in chat.
i. + Voice input (dictation) — lets you dictate a message instead of
typing it, with a bundled whisper sidecar (docker/stt/, Omni
Chat's own OpenAI-compatible wrapper around OpenAI's Whisper model). No GPU
required — the sidecar runs the faster-whisper engine and auto-detects CPU int8
compute; a CUDA GPU is much faster than real-time. On a Mac, run the sidecar
natively instead of in Docker — containers cannot use the Mac's GPU, but the
native run picks the Apple-Silicon MLX engine (from the repo:
cd docker/stt && python3 -m venv .venv && source .venv/bin/activate &&
pip install -r requirements.txt mlx-whisper && uvicorn app:app --port 8881, then
point CHAT_STT_URL at it):
docker compose -f compose.yaml -f compose.stt.yaml up --build -d
The first start downloads the whisper model weights into the stt-models
volume — watch docker compose logs -f stt. For CUDA set
STT_BASE_IMAGE to a CUDA-capable image, STT_DEVICE=cuda, and on
Jetson/L4T STT_RUNTIME=nvidia in .env. To run the sidecar on
another machine (or beside a host-run core) use compose.stt-only.yaml and set
CHAT_STT_URL=http://host:8881 plus
CHAT_STT_MODEL=large-v3-turbo; any other OpenAI-compatible transcription server
works too — e.g. the hosted OpenAI API with CHAT_STT_URL=https://api.openai.com,
CHAT_STT_MODEL=whisper-1, and CHAT_STT_API_KEY. With no server
backend configured at all, dictation still works via the browser's own on-device speech
recognition where supported (Chrome, Safari). See Voice input for
how it behaves in chat.
Whichever combination you run, open http://127.0.0.1:8080. The default AI
endpoint is http://host.docker.internal:8000; edit CHAT_ENDPOINT_URL
in .env to point at the backend you set up above. See
Chapter 12 for running without Compose, pinning an image tag,
backups, and upgrades.
First run: create the owner
The first time you open a fresh instance, no users exist yet. You'll be taken to a setup screen to create the owner account (username, display name, password). After that, the instance shows a login screen on every visit.
Logging in sets an opaque session token in an HttpOnly, SameSite
cookie. Your identity always comes from that session — never from anything the browser
sends — which is how the app keeps each user's data isolated.
Add to iPhone Home Screen
When the web app is reachable from your phone, open it in Safari, tap Share, then tap Add to Home Screen. iOS uses the bundled Omni app icon and launches Omni Chat in a standalone browser window from that Home Screen icon.
Building from source instead of using Docker? See Local Development Setup. Ready to combine multiple backends or try the Auto Smart Router? See Models & Backends.
Chapter 3
Local Development Setup
For contributors, or anyone who'd rather build from source than use Docker. If you just want to run Omni Chat, see Getting Started instead.
Prerequisites
Building Omni Chat from source needs the following toolchain. Everything runs through the
Makefile, so make is required too.
| Tool | Minimum | Why |
|---|---|---|
| Go | 1.25+ | Builds the core (matches core/go.mod). |
| Node.js | 20 or 22 LTS | Builds the Svelte/Vite web app. |
| npm | bundled with Node | Installs and runs the web build. |
| make | any recent | Runs every documented command. |
| git | any recent | Clones the repository. |
Storage uses the pure-Go github.com/ncruces/go-sqlite3 (no CGO), so there is
nothing else to install. Rust and the Tauri CLI are needed only to build the
desktop app — the core and the web app build without them.
To actually run the app you also need an OpenAI-compatible endpoint already listening — see Getting Started for setting one up with Ollama, LM Studio, or a hosted provider if you don't have one. That's a runtime dependency, not a build dependency.
The README has copy-paste install commands for macOS (Homebrew) and Ubuntu/Debian.
Quick launch
Point the app at a running endpoint and start the development build. This compiles the web app and runs the core with the SPA embedded:
# Ollama defaults to port 11434, LM Studio to 1234, llama.cpp/vLLM commonly to 8000/8080
CHAT_ENDPOINT_URL=http://localhost:11434 make dev
make dev serves the app on port 8080, so open
http://127.0.0.1:8080 in your browser. The core also prints one handshake line to
stdout when it's ready:
READY port=8080
Add CHAT_PORT to use a different port, CHAT_API_KEY if your endpoint
needs a token, and CHAT_DB_PATH to choose where the SQLite file lives — see
Chapter 11 for the full list. Running the compiled binary directly
(Chapter 12) defaults CHAT_PORT to 0, which picks an ephemeral port and
reports it in the READY line. First run creates the owner account the same way as
the Docker path — see Getting Started.
Running modes during development
| Command | What it does |
|---|---|
make dev | Build the web SPA and run the Go core with it embedded. Simplest single-process workflow. |
make dev-core | Run only the Go core. Pair with dev-web for hot reload. |
make dev-web | Run the Vite dev server for the Svelte app (auto-reloads on changes), proxying API calls to the core. |
make build-cli | Build the standalone omni-chat wrapper at .cache/bin/omni-chat. |
make build-cli-sidecar | Build the same CLI into desktop/binaries/omni-chat so the desktop bundle can ship it. |
make dev-desktop | Build the SPA and core sidecar, then run the Tauri desktop shell. Needs Rust and the Tauri CLI. |
make check-ipad-prereqs | Verify full Xcode, the Rust iOS device target, CMake, CocoaPods, and Tauri CLI before generating the mobile project. |
make check-ipad-simulator-prereqs | Additionally verify the ARM64 iOS Simulator SDK, the Rust target, and that Simulator.app is present in the active Xcode (Tauri boots the device by launching it). |
make init-ios | Intentionally regenerate the tracked Tauri Xcode project when its mobile scaffolding changes. |
make dev-ipad | Build the on-device model framework and embedded Go core, then install/run on an attached iPad. |
make build-ipad | Create an iPadOS device build and open it in Xcode for signing. |
make dev-ipad-simulator | Build and run on a selected ARM64 iPad Simulator without device signing. |
make build-ipad-simulator | Create an unsigned ARM64 iPad Simulator bundle without launching it. |
See Chapter 12 for production builds and running the compiled binary.
Chapter 4
The Interface
Omni Chat uses a three-column, adaptive layout. The side columns collapse to give the conversation more room; the right Settings panel starts collapsed by default, and each panel remembers its expand/collapse state per user.
Left sidebar — Library
Your projects and chats in a nested tree. Search by title or tag from the Search chats or #tags box — press ⌘⇧P on macOS or Ctrl+Shift+P on Windows/Linux to jump there (it expands the sidebar if it was collapsed). Filter by hashtag from the tag chips. Collapse the sidebar to widen the conversation.
Center — Conversation
Message history plus the composer. The composer grows as you type and shrinks after you send, so answers get the space. It also tucks into a slim peek bar when an answer starts streaming (or when you hide it with the chevron); press ⌘⇧L on macOS or Ctrl+Shift+L on Windows/Linux to expand it and put the cursor in the input for a follow-up. Press ⌘F on macOS or Ctrl+F on Windows/Linux to find text in the open conversation — a small bar highlights matches in visible message text; Enter goes to the next hit and Escape closes it. Those shortcuts do not re-collapse the composer, and they are ignored while a dialog is open. Send becomes Stop while a response is streaming, and a thin orange glow sweeps around the composer border (with the collapse chevron gently breathing) for as long as the model is generating — the hint shows on the collapsed peek bar too, so a live turn stays visible even with the composer tucked away. With reduced motion enabled in your OS, the sweep becomes a static accent ring instead.
Right sidebar — Settings
Per-session response controls: profile, tone, length, creativity, format, sources, and the editable session instructions. It starts collapsed by default to give the conversation more room — open it from the gear button on its rail, and your choice is remembered per user on that device. Expand Profile to reach its style controls, nested Session Instructions, and the selected profile's Skill when one is attached. Both text viewers start collapsed and have their own eye icons.
The top-right user button opens the account menu. From there you can edit your username, display name, and local avatar image, toggle Memory for the current chat, manage saved memories, or sign off.
The composer row also carries lean Model and Profile dropdowns so you can switch either without opening the Settings panel. The rest of this manual walks through each control.
Light & dark theme
The top bar carries a sun/moon button beside the switch-user button that toggles between the light and dark themes. The theme defaults to your operating system's preference and is remembered per user in the browser, so each account keeps its own choice on that device.
Make the interface larger or smaller
In the desktop app, press Cmd++ or Cmd+- on macOS (Ctrl instead of Cmd on other platforms) to zoom the whole interface. Cmd/Ctrl+0 returns to 100%. Zoom steps through 80%, 90%, 100%, 110%, 125%, 150%, 175%, and 200%, and the selected level is remembered for each user on that device. The web app uses the browser's native zoom shortcuts and limits; touch and trackpad pinch zoom remains available.
Chapter 5
Chatting
Starting and organizing chats
- New chat — start a conversation inside a project or leave it unassigned. After the first answer completes, Omni Chat may replace the initial first-message title with a short generated summary of that first exchange.
- Projects — group related chats. A project can hold shared instructions that are added to every chat inside it (see Chapter 8).
- Hashtags — type
#to create or attach a tag (for example#researchor#implementation). Tags appear as chips and filter the left sidebar.
Chat Details
The Chat Details control in the top bar lets you rename a chat, move it to another project, add or remove hashtags, archive it, or delete it.
Archiving hides a chat from the list without deleting it. Turn on Show archived in the sidebar filters to bring archived chats back into view; Chat Details then offers Unarchive. Deleting is permanent, and both actions ask first.
Sharing & export
The share icon in the top bar exports the current conversation as a standalone document. Both options render what's already on screen into one self-contained HTML page — no server round-trip.
- Download HTML saves a single
.htmlfile with the chat title, every message, rendered markdown, code-run outputs, and web-source citations. Image attachments are inlined as data URLs so the page opens fully offline; other attachments are listed by filename. A chat containing math also carries the KaTeX stylesheet and fonts inside the file. Assistant thinking is not included. - Save as PDF opens that same document in a print view — choose your browser's Save as PDF destination to keep crisp, selectable text without any extra tooling. (Allow pop-ups for the app if the print view doesn't open.)
- Download Markdown saves a
.mdfile with each assistant answer kept as Markdown, code-run outputs as fenced blocks, and web sources as links. Image attachments are referenced by filename rather than embedded, so the file stays small and portable.
Working with messages
- Copy — copy a message to the clipboard. It is on your messages as well as the assistant's, and copies the text as you typed it (including any code blocks).
- Edit & resend — change one of your earlier messages and re-run from that point.
- Edit response — revise an assistant response from a short edit instruction, such as adding sound to a generated game or rephrasing one paragraph. The original stays visible while the edit runs; revision thinking can be expanded, and the answer is replaced only after the edit succeeds. The instruction is not saved as a chat message.
- Regenerate — re-run the assistant's answer from any prior turn (handy after switching model or profile).
- Continue — when a response stops short (it reached the length limit, or the connection dropped mid-stream), the newest message shows a "cut off / interrupted" note with a Continue button. It resumes the same message in place — the model picks up where it left off instead of starting over — so long code or documents can be finished across a few continuations. Raise the Response length style setting to reduce how often this happens.
- Background generation — a turn keeps generating even if you click into
another chat, reload the page, or briefly close the lid: generation runs server-side,
detached from the browser connection. The generating chat pulses gently in the sidebar so
you can find your way back; opening it shows everything streamed so far — thinking trace
included — and continues live. On reload the app re-attaches automatically. Only pressing
Stop (or leaving no browser attached for the detach grace,
CHAT_STREAM_DETACH_GRACE, default 5 minutes) ends a turn early. While a turn runs, other chats are read-only — their Send button explains where the model is busy. - Stuck-thinking recovery — a reasoning model that loops in its thinking
or thinks past its reasoning budget is stopped automatically instead of burning the whole
output window. Running out of thinking budget no longer costs you the turn: the model is
asked once more with thinking off, handed the notes it already wrote, and finishes the
answer — the message then carries a short "its thinking ran long, so it wrote this answer
from its own notes" line. A repetition loop, or an overrun on a model whose thinking cannot
be switched off, still shows "The model got stuck thinking, so it was
stopped" with two one-click actions: Retry without thinking (re-runs the turn
with thinking off, for that turn only) and Regenerate. The partial reasoning
stays viewable in the collapsed Thinking block. While a model is thinking, the Thinking label
shows a live elapsed timer so long reasoning reads as progress, not a hang. Known local model
families (Qwen3, Gemma, gpt-oss, DeepSeek-R1, Nemotron, GLM-4.5+) also automatically get their
publishers' recommended anti-repetition sampling settings on self-hosted backends — see
sampling_defaultsin the endpoints reference. - Small models on modest hardware — tool schemas are not free. Shown a wide
tool surface, a very small model stops answering and starts routing: asked to pull some fields
out of a message as JSON, a 1.2B model replies "I don't have a tool that can do that — would
you like me to search?", even though it answers the same question correctly with no tools
attached. Testing showed the cause is the presence of tool schemas, not prompt length
or tool count — a shorter prompt, an explicit "just answer directly" instruction, and cutting
ten tools to two all failed to help. So Omni Chat sizes the tool surface to the model. Below 2B
parameters a tool is offered only on a turn that actually calls for it (you pasted a link, asked
for a picture, told it to remember something, or asked about something current); below 10B only
the external-agent handoff is withheld. Once the model uses a tool the full set returns, so
multi-step work like search-then-read still chains. Models whose size is unknown are treated as
large and behave exactly as before. Tools a model is too small for are dimmed in the Tools menu
with the reason. Override any of this with
prompt_profileper endpoint or per model inendpoints.json(auto/full/terse/minimal), orCHAT_PROMPT_PROFILEwithout a config file. To see how a model behaves under this framing, runmake probe ENDPOINT=http://localhost:8000 MODEL=<id>: it replays a fixed set of prompts at the resolved profile and again forced tofull, and prints a pass/fail table. It needs a live backend and is not part ofmake check. - Delete with undo — remove a message with a short undo window.
- Stop — cancel a streaming response; the partial answer is kept.
- Generation details — hover or focus the info icon for a quick view, or click it to pin the panel until you close it, click elsewhere, or press Escape.
Writing code in the composer
Type three backticks on an otherwise empty composer line to open a monospace code block. Paste or type code there without syntax highlighting; Enter and Shift+Enter both add another line and never submit. Press Arrow Down at the end of the block, or click the blank line below it, to continue with normal text. A populated code block disappears only after all of its contents are deleted or cut; Backspace dismisses a newly opened empty block. After sending, the user message keeps the visual code block while ordinary user text remains literal.
The Generation details panel shows the profile used (or No profile), the exact model and endpoint that handled the response, the SOUL source/hash, whether private identity context was included, memory retrieval mode/fallback and the selected memories with reasons, thinking mode, and the tool outcomes for the response — web access, and, when those features were in play, code execution and agent delegation — plus labeled input/output counts, prefill and generation speed, time to first token, thinking time when available, and total request-to-completion time. Omni Chat saves these details with each new assistant response, so changing a chat's settings later does not rewrite its history. Older responses show only values that can be proved from their saved data and label the rest Not recorded.
Attaching files (images & documents)
Use the paperclip button in the composer — or drop files anywhere on the conversation (the transcript, welcome screen, or composer; Finder drops work in the desktop app too), or paste an image from the clipboard — to attach files to your next message. A drop expands a hidden composer so you can see the chips. Off-the-record chats and an in-flight response refuse the drop, same as the paperclip. Supported types: images (PNG, JPEG, WebP, GIF) and documents (PDF, plain text, Markdown, HTML), up to 20 MB per file and 8 files per message. Attached files appear as chips; remove one with its ✕ before sending. Sent images show as thumbnails in the chat history and documents as clickable chips.
- Images are sent to the model as image content on that turn, so you can
ask "what's in this photo?". This needs a vision-capable model (for example a local
Qwen2.5-VL served by llama.cpp or Ollama); the composer warns you when the selected model
cannot view images, and Auto (Smart Router) picks a vision model automatically. Original
source pixels are preserved and each current turn is resized once for the resolved endpoint's
Low VRAM, Balanced, or High detail budget. Use the bounded High image detail
option when small text or diagram detail matters; it never sends unrestricted original
dimensions. To keep follow-up turns
cheap on local hardware, images are only sent in full on the turn you attach them —
later turns reference them by name.
Vision support is detected from, strongest first:model_overridesinendpoints.json; the endpoint's own reported modalities (the llama.cpp router exposesinput_modalitiesper model — allama-serverstarted without--mmprojreports text-only and is treated as such, so load the mmproj file to enable images); and a built-in catalog of known multimodal families (Gemma 3/4, Qwen 3.5/3.6, Qwen-VL, LLaVA, Pixtral, GLM-4.5V/4.1V, GPT-4o/5, Claude, Gemini). If a vision model is still misdetected, add{"model_pattern": "…", "vision": true}to that endpoint'smodel_overrides. That override is also the answer for a model whose vision support depends on the build rather than the name — GLM-5.2 ships as both a text-only and an image-capable deployment under the same model id, so the catalog claims no vision for it and you opt in on the endpoint serving the multimodal build. - Documents have their text extracted and given to the model as source context, on the turn you attach them and on follow-up questions in the same chat (within the model's context budget).
- Scanned PDFs with no text layer are handled by extracting their page images and sending those through the vision path — chat about a scanned document exactly like a photo. If nothing can be extracted, the upload is rejected with a clear message.
Reasoning / thinking traces
When a model emits a separate thinking or reasoning block, Omni Chat keeps it apart from the answer and shows it in a collapsible region. Reasoning is stored separately and is never fed back into later prompts — only the answer content is. The compact brain switch beside the composer Profile selector and the labeled switch in Settings are synchronized and update the same per-chat state. When the leading reasoning phase ends, its folded title shows the elapsed time from the first reasoning token to the first visible answer token; this duration remains available after reloading the chat. If the selected model has no known reasoning channel or adapter, the Thinking switch is shown as unavailable. OMLX models are detected from their chat-template capability metadata rather than a model-name allowlist.
Math rendering
Assistant answers render LaTeX with KaTeX, bundled into the app rather than loaded from a
CDN, so formulas work offline and look the same in the browser and the desktop app.
$$…$$ and \[…\] become a centred display equation — a long one
scrolls inside the message instead of stretching it — while $…$ and
\(…\) render inline on the text baseline.
Money in ordinary prose is left alone: "it costs $5, so $10 total" stays plain text. A bare
$…$ counts as math only under the strict rules TeX itself uses — no space after
the opening $, no space before the closing one, no digit immediately after it,
and no line break in between — and anything money-shaped fails at least one of them.
Because models routinely emit slightly invalid LaTeX, the renderer is deliberately
forgiving. A literal % (as in \mathbf{77.18%}, which TeX would read
as a comment and use to swallow the rest of the line) and a literal $ (as in
\frac{$140}{$860}, which would otherwise be a parse error) both render as
written; the trade-off is that TeX comments cannot be used inside math. Genuinely broken
LaTeX shows its source in red rather than failing the whole message, and math inside code
spans or fenced code blocks is never rendered. Exported chats carry the KaTeX stylesheet and
its font faces inside the file, so math survives offline in a downloaded HTML page.
Code preview
Assistant code blocks include toolbar controls to copy or save the displayed code. The save
action opens a Save As picker when the browser supports it, otherwise it falls back to a normal
local download. Named blocks keep their filename and unnamed blocks use a language-based
snippet.* name; both include an ordering suffix such as main-01.py or
snippet-02.js. HTML code blocks in an assistant message can be rendered in a
sandboxed preview that runs inline CSS and JavaScript, including dynamically evaluated code used
by apps such as calculators, and may load resources over HTTPS from a CDN. Adjacent
css and js blocks in the same message are bundled into the same
preview. Use the popout control on an HTML preview to open the same sandbox on its own — a new
browser tab in the web app, a separate app window in the desktop app; the inline message
returns to code view while keyboard-driven demos such as games keep focus in the popout.
The preview runs without allow-same-origin, so previewed code cannot read
your app cookies or call the authenticated /api/* endpoints. A popped-out
preview in the desktop app gets the same sandbox on an app-internal address of its own, and
reaches none of the app's own commands.
Running code blocks
JavaScript, TypeScript, Python, Go, and bash code blocks show a Run button that executes the snippet locally, in a sandboxed WebAssembly runtime (QuickJS for JS/TS, Pyodide/CPython for Python, a Yaegi interpreter for Go, WASIX GNU bash for bash/sh) inside a dedicated Web Worker — nothing is sent to a server. Output (stdout and stderr, the exit code, and how long it took) streams into a panel below the code, with Stop, Re-run, and Close controls. The runtime downloads once on first use and is then cached (the Python runtime is about 6 MB, Go about 8 MB, bash about 4 MB).
Interactive programs: a JavaScript or TypeScript program that reads input
with Node readline or process.stdin, a Python program that calls
input(), or a bash block runs in a small terminal below the block — type into it and
the run keeps going. For JS/TS the run pauses while it waits for a line (idle time doesn't count
against the 30-second limit; a runaway loop after you answer still does) and works in any browser
with no special setup. Python input() and bash need a cross-origin-isolated
context (a SharedArrayBuffer): they work in Chrome and Firefox and show a
"not supported" notice in Safari. bash runs your block as a real script (bash -c)
with coreutils (ls, cat, and friends); if the script calls
read, the terminal stays live for input. It runs without the 30-second wall-clock
limit (Stop and the 1 MiB output cap are the guards). Because bash runs in WebAssembly
without full POSIX signals (no SIGPIPE), a filter reading an unbounded
stream piped into another command — e.g. tr < /dev/urandom | head -c 16 —
produces no output; bound the source first (head -c 64 /dev/urandom | …).
Multi-file programs: when code blocks in one answer carry filenames —
on the fence (```python main.py) or as a bold/heading line right above the block
(**utils.py**) — they form a workspace. Running a named block mounts its named
siblings into the sandbox's in-memory filesystem, so import utils, a TypeScript
import './greet', or a multi-file Go package works. The files exist only for that
run and are discarded with it; blocks without filenames run alone exactly as before. If the
same filename appears twice in an answer, the later block wins.
Because the code is often AI-generated, the sandbox has no network, no DOM, and no
filesystem: fetch, XMLHttpRequest, WebSockets, and browser
storage are all removed before the code runs, so it cannot reach your session or the network.
Each run is capped at 30 seconds of wall-clock time and
1 MiB of output, memory is capped at 128 MiB, and nothing runs
until you click Run. CPU use is bounded only by the timeout — a busy loop simply runs until it
is killed.
A curated set of packages is bundled offline and loads automatically when your code imports
it: numpy and pandas (plus their dependencies — a one-time ~8 MB
download, fetched from the app itself, never the internet). Anything else fails with a
ModuleNotFoundError followed by an honest note: package installation
(pip/micropip) is not possible because the sandbox has no network access — that is the
security boundary, not a missing feature. Graphical packages (pygame, tkinter) wouldn't work
in the sandbox anyway: there is no display. If a runtime takes longer than two minutes to
download, the run is stopped with “runtime failed to load in time”.
The bash sandbox ships GNU bash + coreutils only. When a script calls a missing command
(pip, python, node, apt…), the normal
command not found error is followed once by a note that only bash + coreutils are
available. If an AI answer pairs a Python program with pip install … /
python file.py shell instructions, skip the shell block and click Run on the
Python code block itself.
Go support runs a Go interpreter (Yaegi) compiled to WebAssembly: the standard library is
available, but networking (net, net/http), os/exec,
syscall/syscall/js, and unsafe are withheld from the
interpreter as a security boundary, and third-party modules and go get are not
supported. Interpreted Go is slower than native and has no goroutine preemption — a runaway
program is stopped by the 30-second limit rather than interrupted mid-computation.
Letting the model run code itself (Code Execution)
With a tool-capable model, the assistant can execute code as part of answering — for
example "sort these numbers and tell me the median, use a tool". Turn on the
Code Execution toggle in the composer — the </> tray button
or its Tools-menu switch, which outlines orange when active (off by default). The model is then offered a
run_code tool for Python and JavaScript: it writes a program,
your browser executes it in the same sandbox as the Run button — no network, no filesystem, no
input, 30-second cap — and the printed output flows back into the model's answer. The code
never leaves your machine for execution; the server only relays it to your own browser tab.
While it runs, the response shows a Running Python… chip that turns into a collapsed card such as Ran Python · ok · 0.4 s. Expand the card to see exactly what code executed and its stdout/stderr; the card is saved with the message, so you can audit it after a reload too. The model gets at most three runs per response. If the run fails or times out, the model is told so and answers as best it can — check the card when a computed answer looks off. Go and bash are not offered to the model (they stay Run-button-only), and with the toggle off nothing ever executes without your explicit click.
Per-chat tools live in the composer as a tray of up to four icon toggles
plus a Tools menu (the grid button) that lists every tool with an on/off
switch; hover a row for its description. Tools whose backing service was never configured don't
appear anywhere — a bare install shows only Code Execution, Thinking, Memory, and (when the
browser supports speech) Read aloud. A tool whose service is configured but currently down
stays in the menu dimmed with the reason and re-enables by itself when the service returns.
Enabled tools claim the tray slots first, so whatever you switched on stays one click away;
the rest waits in the menu. All of these apply to the chat you're in right now. To set what
a fresh chat starts with, use the New chat defaults switches in the right
settings sidebar. Those defaults are saved to your account and sync across devices;
changing one only affects chats you create afterward, never a conversation already open.
The Thinking default is the exception for omni-chat launch:
a launch has no chat, so it uses that account switch (on/off) plus Thinking effort.
(Agent handoff and MCP Tools are not new-chat defaults — they're situational and start off
in every chat until you turn them on.)
Off the record (temporary chats)
The privacy-shield toggle next to the theme switch in the top bar starts a temporary chat: while it is on, nothing about the conversation is saved anywhere. No chat appears in History, no messages are written to the database, and no memory, rolling summaries, or chat title are generated. The transcript lives only in your browser for the session and disappears when you reload the page or open another chat — you'll be asked to confirm first, because it can't be recovered.
Because nothing is recorded, some features are turned off in this mode: Memory is forced off and its switch is locked, and file attachments, image generation, and the Agent handoff are unavailable (they would each leave something on disk). Web search, URL fetch, Code Execution, and MCP tools still work — none of them persist once the turn isn't saved. You can only switch a chat to off-the-record while it is brand new and empty; once you've sent the first message it stays on until you start or open a different chat. A banner across the top of the conversation reminds you the chat isn't being saved.
Opening interactive Pi with Omni Chat models
If Pi is installed on your computer, the standalone wrapper can launch your normal Pi
interface with every text model that the Omni Chat model picker can currently see.
The macOS one-line installer puts omni-chat on your PATH.
From source, build it first:
make build-cli
.cache/bin/omni-chat launch pi
A packaged desktop app already contains the same binary at
Omni Chat.app/Contents/Resources/omni-chat.
Launch uses your New chat defaults → Thinking switch
(and Thinking effort), not the composer toggle of the chat you have open.
Restart the launch after changing that default.
The command checks that the backend and its database are healthy, asks for your Omni Chat
username and password, then opens Pi. Use Pi’s /model picker at any time; model
calls travel through Omni Chat, so OAuth subscription endpoints and Responses-protocol
endpoints work without copying their credentials into Pi. Your password and endpoint keys are
never written to disk. A short-lived launch token lives only in Pi’s environment, is renewed
while Pi runs, and is revoked when it exits. If an endpoint connection drops before producing
any model output, the core retries it twice with a short bounded backoff; after any model event
arrives, the request is never replayed. When a model does not report an output cap, the
temporary Pi provider uses the proxy’s 32,768-token file-writing fallback; an explicit
endpoint/model cap still wins, and the proxy clamps the request to the context room left after
Pi’s transcript and tool schemas. Reasoning may use that entire resolved output allowance:
the proxy does not apply the smaller chat-response thinking budget and does not need to
recognize the model family or inference backend. Repetition-loop protection remains active;
if a backend streams past the full allowance, the proxy returns a normal
finish_reason: "length" and completed stream rather than disconnecting Pi.
The wrapper looks for the server in this order: --server,
OMNI_CHAT_URL, the desktop app’s remembered port, then
http://127.0.0.1:8080. Remote servers should use HTTPS; non-loopback HTTP requires
an explicit --allow-insecure-http. You can choose the initial model and pass normal
Pi arguments after --:
omni-chat launch pi --server https://chat.example.com --username alice
omni-chat launch pi --model local/qwen3-coder -- --continue
Your regular Pi sessions, skills, themes, and extensions remain available. The wrapper owns
the provider/model routing flags, and its v1 catalog is text-only: it excludes the virtual Auto
Router and image-generation backends. Trusted Pi extensions can still make their own unrelated
network calls. This command runs Pi on your computer and does not need
agents.json; it is separate from Agent handoff.
Handing a task to an external agent (Agent)
Code Execution runs in a sealed sandbox — no files, no network. When a task needs the real
thing (edit a project, run its tests, make a git commit), the assistant can hand it to an
external coding agent running on the machine that hosts Omni Chat — either
installed there directly or inside a Docker container the operator configures (see
Running Omni Chat, combination f, for the Compose
setup). Built-in adapters support Pi and
Hola's
hola-coder tail (0.6 or newer). This only works if whoever runs your instance has enabled it
(by creating an agents.json that lists which folders agents may touch — or, in the
desktop app, filling in Settings → Agent handoff); if you
don't see the Agent toggle do anything, it isn't configured on your
instance.
Turn on the Agent icon (the robot button) and ask for something that needs a real workspace — "in ~/Work/myproj, add a failing test for the parser and make it pass — delegate it." The assistant writes a self-contained brief and a confirmation box appears showing the folder and the brief. Nothing runs until you approve, and you can edit the brief first — the agent sees exactly what you approve. The agent uses the same model your chat is using — its model calls are relayed through Omni Chat itself, so it works with every configured backend (including subscription ones like the ChatGPT Codex endpoint) and the agent never sees your endpoint's API key. It also inherits the effective Thinking choice for that turn: switching Thinking off applies the endpoint/model’s off-body and non-thinking sampling policy to every Pi or Hola model call, leaving the agent’s own harness to manage planning and convergence.
While it works you'll see a live activity card with an animated working indicator, a live elapsed timer, and a Stop task button — so a quiet agent still looks active. When it finishes, a collapsed card records the brief, what it did, a short summary, and a list of the files it changed — saved with the message so it survives a reload. That file list is a git diff when the workspace is a repository and a plain created/modified/deleted list when it is not, so a fresh per-chat folder reports its work like anything else. If you deny it, stop it, or it times out, the assistant answers with what it has — including the files the agent had already finished, so asking again continues that work instead of starting it over.
When the task succeeds the assistant tells you where the files are rather than printing them into the chat. That is deliberate: the assistant cannot read the agent's workspace, so anything it wrote out would be its own second version of work that is already done, not the file on your disk. Open the file from the workspace folder to see the real result.
If the code is already in this conversation, the assistant can pass saved code blocks
directly to the agent. The approval dialog lists the selected blocks. After approval, Omni
copies their exact bytes into separate input files in a .omni-handoff-* folder;
the transfer does not overwrite project files. Up to eight blocks (1 MiB total) can be passed,
separately from the task brief. Their references remain in the saved run card. This includes
code in your current or edited message, and the original answer during an assistant revision.
Each turn permits an initial handoff and one corrective handoff, with a new approval for each. A workspace picked or reset during approval also applies to the corrective handoff. Denying or stopping a task ends further handoffs for that turn. Run finished means the agent process finished; check its summary and validation results to see whether the requested task succeeded. After three consecutive failed tool calls, Omni stops tool attempts and asks the assistant to explain the remaining blocker.
Image tools are for pictures relevant to your request, not for retrieving code from chat history. Omni rejects invented image links before downloading: links must come from your messages, image-search results, a fetched page, or a supported live-chart endpoint. Revising an answer does not turn assistant-invented links into user-supplied links; links in your actual edit request are still eligible.
A finished Pi task leaves its session behind, so you can open a terminal, cd into
the workspace folder and run pi --continue to pick the work up by hand. Give it a
provider and model when you do — pi --continue --provider <yours> --model
<yours> — because the session records the temporary routing the app set up for that
one run, which no longer exists once the run ends. A task you stop mid-run leaves the
session folder but no transcript: the agent is killed rather than asked to exit, so there is
nothing for it to flush. The agent can only ever work inside the
folders your instance allows, one task at a time.
On a shared (multi-user) instance the operator usually configures managed workspaces: each conversation gets its own private folder, created the first time you delegate and reused for follow-ups in the same chat — so "now add sound effects" continues where the last task left off. Other users can never see your folders. In this mode you don't pick a directory; the confirmation box shows the one assigned to the chat.
In the desktop app the confirmation box also lets you change that folder: press Change… to browse, or create a new subfolder from the file dialog, and the task runs there instead. The choice sticks for the rest of that conversation — follow-ups reuse it without asking again — while a new chat starts fresh with its own folder. Use this chat’s own folder undoes it. You can only pick somewhere inside your own workspace area, so a chosen folder is still private to you and still reachable over SFTP; anywhere else is refused with the allowed folder named, and the box stays open so you can pick again. If a folder you chose is later renamed or on a drive that is not mounted, the task falls back to the chat's own folder and says so, and your choice is remembered for when it comes back.
Getting at your files (SFTP)
If your instance has SFTP enabled (managed workspaces only), you can browse and copy
everything the agent built — and drop files in for it to work on next — using any SFTP client.
Sign in with your Omni Chat username and password; you only ever see your own folders, one per
conversation. Each run card shows that conversation's folder name
(cht_…) next to two copy buttons: one copies the folder's path on the
server, the other (the server icon, shown only when SFTP is on) copies a ready-to-paste
sftp://you@your-server:2222/cht_… address — drop it into any of the
clients below and you land straight in that conversation's folder. Paste it into an SFTP
client or the Terminal, not a web browser — browsers can't open sftp:// links,
so Safari or Chrome will show an "invalid address" error.
- VS Code: install the SSH FS extension and add a
connection to
sftp://you@your-server:2222— your workspaces mount like a local folder. Note it must be SSH FS, not Microsoft's Remote - SSH: Remote - SSH needs to run a server program on the remote machine, and this port deliberately allows file access only — no commands. - FileZilla / Cyberduck: protocol SFTP, host
your-server, port2222, password login. - Terminal:
sftp -P 2222 you@your-server.
The port may differ on your instance — ask whoever runs it. There is no shell login on this port, just file access.
Calling tools on MCP servers (MCP Tools)
Beyond the built-in tools, the assistant can call tools from MCP servers
(Model Context Protocol) that the operator has configured — remote services or local tool
programs. Like Agent handoff, this only exists if your instance enables it (by creating an
mcp.json); when configured, an MCP Tools icon (the server button)
appears in the composer and the right settings sidebar lists the configured servers. It is off
in every chat until you turn it on, and it is not a new-chat default.
With the toggle on, the model may call the servers' tools while answering. Every call asks you first: a confirmation box names the tool and server in plain language, explains what the tool does in the server's own words, and lists every argument on its own labelled row — so you can read what will happen instead of decoding JSON. Hover the ⓘ beside an argument for the server's description of it. A long value (a file's contents, say) is clamped to a readable preview with Show all N lines to see the rest, and View raw JSON shows the exact payload whenever you want to check it character for character. Nothing runs until you approve. For a tool you trust (say, a read-only search), tick "always allow this tool in this chat" and it stops asking — the grant applies only to that tool in that chat, never globally. Denying a call is fine: the assistant is told and answers with what it has. The response's info popover shows whether MCP tools were used, and each call (server, tool, outcome, duration) is stored with the message.
Some servers never let you skip that prompt. If the operator marked a server always ask — because its tools do something in the real world that shouldn't repeat unattended, like placing a trade — you'll see a shield note where the always allow checkbox usually sits, and a shield beside the server in the settings sidebar. Approving one call on such a server never carries over to the next one.
Each configured server has an on/off switch in the settings sidebar's MCP servers list. Turn one off and it offers no tools in any of your chats until you turn it back on — a quick way to silence a single server (say, one you're not using right now) without affecting the others. This is your own setting and applies across all your chats; it's separate from the per-chat MCP Tools switch in the composer, which turns every MCP server on or off for just that one chat.
Some servers require signing in with your own account — they show Needs connection in the settings sidebar's MCP servers list. Click Connect: the provider's sign-in page opens — in the desktop app, in your default browser, since many providers refuse to sign in inside an embedded window. Approve it there and close the tab; the list checks for a couple of minutes and updates itself, or use Refresh status. Your sign-in is yours alone — other users on the instance connect their own accounts — and Disconnect removes it at any time. Until you connect, such a server simply offers no tools. Setup for operators is in Chapter 11 and the README's "MCP tools" section.
Trading on Robinhood. Robinhood publishes an MCP server, so if your
operator has added it you can ask about your portfolio and place orders from a chat. Connect it
the same way — Connect next to robinhood in the settings sidebar,
then sign in on Robinhood's page (this needs a desktop browser; Robinhood sets
up your Agentic account during that first sign-in). Two things to know: Robinhood only
lets an agent place trades in that separate, separately funded Agentic account — reads cover
your other accounts, orders don't — and every single call asks you first, showing the symbol,
side, and quantity as labelled rows. Approving one order never pre-approves another. Read each
prompt: a model can misread a request or act on a stale quote, and denying is always the safe
answer.
Operators offering several servers: the model is offered at most
max_offered_tools MCP tools per turn (a top-level field in mcp.json,
default 128, max 512), filled in server order. If you connect several large servers whose
combined tools exceed that, the later servers' tools are silently dropped (the core logs
mcp: offered tool cap reached) — raise max_offered_tools, reorder the
servers, or narrow a big one with tool_allowlist.
Operators running under Docker: the omni-chat container ships without
Node, so MCP servers launched with npx cannot run inside it. Use the
compose.mcp.yaml overlay instead — it runs each npx-installed plugin (a Figma
server and a filesystem server come ready-made) in its own Node sidecar bridged to Streamable
HTTP, and config/mcp.json points at them with ordinary
"transport": "http" entries (e.g.
http://mcp-figma:8000/mcp). Secrets like FIGMA_API_KEY go in
.env and reach only the sidecar. Details and the add-another-plugin recipe are in
the README's "Docker: running npx MCP plugins" section.
Generating images
If your instance has an image backend configured, just ask: "Generate an image of a cyberpunk cat riding a skateboard." A tool-capable model writes a detailed prompt (you can ask for square, portrait, or landscape) and an animated placeholder appears in the response the moment generation starts, already shaped like the image to come. With the bundled backend you then watch the image form: a blurred preview of the actual picture appears part-way through and sharpens as the percentage climbs, until the finished image replaces it in place, with the prompt as its caption. Generation happens on the server's configured backend and can take from seconds to a couple of minutes depending on the hardware — the response simply continues once it lands.
Each image has three buttons: Download saves it (named after the prompt), Variation drops a ready-made variation request into the composer — edit it if you like, then send, and the model generates a fresh take — and Edit starts an image-to-image edit of that exact picture (see below). Images are stored with the conversation (they survive reload and appear in chat exports) and are private to your account. The assistant can produce at most two images per response; if the backend is offline the assistant says so and answers without one. The image server draws one picture at a time — if someone else's generation is running, the placeholder reads "Waiting for the image server… (N ahead)" until it's your turn (waiting never counts against the timeout), and when the wait queue is full the assistant simply reports that the server is busy and suggests trying again shortly. If nothing happens when you ask for images, the instance has no image backend configured — see Running Omni Chat, combination g.
Editing images (image-to-image). The assistant can also transform an
existing picture: attach a photo and say "make this look like a watercolor", say
"now make it night time" about an image it just generated, or click
Edit on any generated image — it prefills the message with that image's
reference so you just type the change. While an edit runs, the placeholder starts from a
blurred copy of the source image, and the result carries an Edited caption. This
works with any tool-capable chat model — the chat model does not need vision;
it hands your instruction and an image reference to the image backend, which does the actual
transformation. (If your instance's image backend runs without a dedicated edit model, edits
are re-imaginings guided by your instruction rather than surgical changes — style changes
land better than element edits, and describing the full desired result works better than a
terse instruction. Ask whoever runs it about IMAGEGEN_EDIT_STRENGTH for a
stronger restyle, or IMAGEGEN_EDIT_MODEL_ID for precise instruction edits.)
Read aloud (text-to-speech)
If your instance has a speech backend configured, every assistant response grows a speaker button in its action row. Click it to hear the response read aloud in your chosen voice — the icon becomes a stop button while the audio is prepared and while it plays; click again to stop. Only one message speaks at a time, and replaying the same message in the same session starts instantly (the audio is kept in memory, not re-synthesized). Code blocks are skipped with a short spoken note, links read their text, and tables are read row by row.
For hands-free use, the composer has a Read aloud toggle (in the tray or the Tools menu): while it's on for a chat, each response is read aloud automatically the moment it finishes streaming. Responses that error out or that you stop are never spoken. A matching switch under New chat defaults in the settings sidebar makes auto-speak the default for your new chats.
Expand the default-collapsed Voice & Dictation section in Settings and use its Voice Engine group: it contains the engine, a voice list (from the instance's speech backend, or several grouped lists when more than one is configured), a speed control (0.5×–2×), and a Preview button that speaks a sample sentence. With the bundled OmniVoice backend, a voice that isn't one of the listed presets is used verbatim as a voice description — save your own, e.g. male, elderly, low pitch, british accent (attributes: gender, age, pitch, accent, whisper style). The preference is saved to your account, so it follows you across devices. Nothing about read-aloud is stored on the server — audio is synthesized on demand and streamed to your browser, and it works in off-the-record chats too.
No speech backend? Your device speaks. When the instance has no speech backend configured — or every configured one is down — read-aloud automatically falls back to your device's own OS voice through the browser's built-in speech synthesis: instant, free, and offline. The Engine select in Voice Engine (shown when both options exist) lets you force it: Auto prefers the server backend when it's healthy, This device always uses the OS voice. The device engine speaks with the voice configured at the OS level (macOS: System Settings → Spoken Content) — there is deliberately no in-app list of the hundred-plus system voices — and it honors your speed preference. To get the richer OmniVoice/hosted voices instead, see Running Omni Chat, combination h.
Voice input (dictation)
The composer has a mic button before Send. Click it (or press Ctrl+Shift+D) to start recording; click it again (or press the shortcut again) to stop — the recording is transcribed and the text lands at your cursor. Press Escape while recording to cancel and discard the audio instead. A recording auto-stops after 3 minutes.
Dictation uses one of two engines: the instance's server transcription backend when one is configured, or your browser's own on-device speech recognition (Chrome, Safari) when it isn't — or when you'd rather keep the audio on-device. Pick which under Dictation in the default-collapsed settings sidebar Voice & Dictation section: an engine select (Auto / Server / This device — Auto prefers the server backend when it's configured and healthy) and a language select (auto-detect, or pin a language for more accurate transcription). The preference is saved to your account, so it follows you across devices.
Privacy: audio is transcribed and discarded, never stored in the server's database. The core holds it in memory only; the bundled whisper sidecar writes a transient temp file it needs for transcription and deletes it immediately afterward, so nothing is retained on either engine. It works in off-the-record chats too. If the mic button doesn't appear at all, the instance has no server transcription backend configured and your browser doesn't support on-device speech recognition either — see Running Omni Chat, combination i.
Chapter 6
Models & Backends
Choosing a model
Pick a model from the Model dropdown in the composer row. The list is
fetched from your backend's /v1/models endpoint. Each chat
remembers both its model and its backend, so reopening a
chat restores the right one. For a fresh chat, the picker defaults to the model you used
last — preferring it only when it is online (loaded, loading, or from a backend that
doesn't report residency). If it isn't, the next online model on the same backend is
chosen, else an online model from another backend; and if a backend disappears while
you're on the welcome screen, the selection heals itself the same way on the next
refresh — returning to your remembered model once its backend is back.
You can't switch models mid-stream — the picker is disabled while a response is in flight. Use Stop (or regenerate afterwards) to change models.
A small status LED sits beside each real model inside the composer Model picker, and the closed picker repeats the selected model's LED next to its name. Green means the selected backend confirms the model is loaded and ready; amber means it is loading; neutral means it is unloaded, sleeping, or the backend does not expose a trustworthy signal. The Auto (Smart Router) entry and the backend divider rows carry no LED. The sidebar's Image generation dropdown uses the same dots for availability instead: green = connected and healthy, amber = needs connecting (e.g. a Codex sign-in), red = configured but failing, grey = off — matching the Services panel. Omni Chat refreshes the state while the app is visible, and polls more often during the first local-network scan so a server found after the window opens still appears in the picker. It auto-detects Ollama, LM Studio, llama.cpp (single-server and router modes), vLLM, and oMLX status APIs. A failed optional status check never removes an otherwise available model.
Combining multiple backends
Omni Chat can present models from several OpenAI-compatible backends in one picker — for example a local vLLM/LiteLLM proxy alongside a hosted provider. The dropdown groups models under a divider per backend; choosing a divider selects that backend's first model. A backend that fails to load at startup is marked offline and skipped — the others keep working.
This is configured with an endpoints.json file, covered in
Chapter 11. With no such file, the app uses
CHAT_ENDPOINT_URL, which may itself be a single URL or a comma-separated list.
Subscription backends (Sign in with ChatGPT)
Some backends authenticate with a subscription sign-in instead of an API key. The
first supported one is Codex ("Sign in with ChatGPT"): declare it in
endpoints.json with "auth": {"type": "oauth", "provider": "openai-codex"}
and "protocol": "responses" (see the README's "Multiple backends" section for the
full snippet). The instance owner then connects it once from the settings
sidebar's Services section — like an API key, the connection powers everyone's
chats on that backend. On a desktop or homelab install the sign-in completes automatically —
and under Docker Compose too, once you set OMNI_OAUTH_PUBLISH_PORT=1455 in
.env, which forwards the provider's fixed localhost:1455 callback
into the container when your browser runs on the Docker host. That mapping is
off by default: claiming a fixed host port for an optional feature would stop
the whole stack from starting on a machine where another app already listens on 1455. Without
it — or from any other machine — the sign-in tab ends on an unreachable or unrelated
localhost:1455 page; that's expected, not an error: copy that
tab's full address into the paste the callback URL field the panel offers. Until it's connected,
the backend's row shows an amber dot with a Connect button (image backends sharing the
same sign-in show the same state), and its models simply don't appear in the picker. The
sign-in key is created in the config directory, or beside the database when the config
directory is read-only (e.g. an :ro Docker bind mount) — no setup needed either
way.
Auto (Smart Router)
Whenever at least one real model is available, the Model dropdown also offers a distinct ✨ Auto (Smart Router) entry pinned above the per-backend groups. Select it and every message you send is automatically classified and routed to the best available model for that turn — weighing the task type (coding, reasoning, a quick question), the capabilities the turn needs (tool calling for web search, a thinking channel), and the request's complexity and size, so heavier or larger-context turns favor bigger models and trivial ones favor faster models. No manual model-switching needed. Auto is opt-in: it is never the default for a new chat.
While a chat is on Auto, the composer shows Auto → <model> once a response
has routed, and the assistant message's Generation details panel shows which model and backend
handled it, the routing reason, and how long the routing decision took. Because the concrete
model — and therefore its context window — is chosen per message, the context meter is hidden
while Auto is selected.
Routing also weighs each model's resolved capabilities — vision, tool calling, parameter
size, and vendor/family — taken from a built-in catalog of common model ids, from id
inference, or from a per-endpoint model_overrides array in
endpoints.json. A message with an image attachment is routed to a vision-capable
model whenever one is available.
Loaded models first. On servers that report which models are in memory
(Ollama, LM Studio, the llama.cpp router and oMLX — the same state the picker's status LEDs
show), Auto picks only among models that are already loaded, so a quick question never makes
the server unload one model and load another. It loads a different model only when no loaded
one can handle the message: an image with no loaded vision model, a prompt longer than their
context, or web search with no loaded tool-calling model. The routing reason then reads
no loaded model fits; loads <model>. Models on your own servers whose load
state can't be read count as not loaded; hosted cloud APIs always count as ready.
Jev by TypeSafe AI. For better judgement of each message, open
Settings → Models → Smart Router, turn on Use Jev and paste a
TypeSafe API key. Jev is a decision model rather than a chat model: one call per message
(about 0.2 s) rates how demanding your latest message is on five levels — trivial,
simple, moderate, hard, expert — whether it needs specialist knowledge, and whether you asked
for a stronger model. Auto then picks the model, so no local model is tied up classifying.
Press Test Jev to check the key and connection before saving; it makes one
tiny call and reports the Jev version that answered and how long it took.
Your message and a little recent context are sent to TypeSafe. If Jev errors or takes longer
than its timeout (default 1500 ms), the built-in router decides instead — a message is
never held up. Routing reasons judged by Jev start with jev:, and
Smart Router (Jev) appears in the service status panel. On the desktop the key lives
in the Keychain as OMNI_KEY_ROUTING; a web install saves it in
endpoints.json's routing section.
Cloud models in Auto. If you have both your own models and a cloud subscription (such as ChatGPT), Settings → Models → Cloud models in Auto decides when Auto may use the cloud. Only for expert requests (the default) sends expert-level messages, explicit requests like "use your strongest model", and follow-ups in a thread a cloud model already answered; everything else stays on your own models. Always lets cloud models compete on every message; Never keeps Auto local whenever a local model can answer.
Follow-ups. A substantive follow-up stays within one band of the model that answered the thread, and prefers that same model, so a derivation that started on a strong model is not continued by a weak one. A simple follow-up — "thanks", "sum that up in one line" — drops to the fastest able model. A reply that continues the previous request — answering the assistant's question, pointing it to something earlier in the chat ("the code is above"), or saying "yes, do that" — stays with the model that was handling it, however short it is.
Auto works out of the box with a built-in policy. Power users can customize model tagging
and per-task preferences with a router.json file — see
Chapter 11 for the full schema and how tagging and routing
decisions work.
Service availability & self-healing
Optional integrations — the agent runner, the image-generation sidecar, web search, web
fetch, embeddings, and MCP servers — are tracked live via GET /api/status. While
the app is visible the core re-checks each configured service with a cheap probe (the image
sidecar's /health, a docker version for containerized agents, the
search/reader base URL, the embeddings server's /v1/models), and every real tool
call also feeds the same status. Nothing is probed while no browser tab is open.
The same holds at startup. If an optional service is configured but its environment is not ready — the agent workspace folder cannot be created, Docker is not running, an agent's command or image is missing — Omni Chat prints a warning saying exactly which service and why, disables just that feature, and carries on serving chat. Only a genuinely broken config file stops it from starting. The agent workspace is re-checked on every status poll, so once the folder exists (or its permissions are fixed) agents come back on their own, without a restart.
When a configured service is unreachable the UI degrades instead of pretending: the
composer's Agent toggle dims with the reason in its tooltip, and a small
alert triangle appears in the composer row. Clicking the triangle lists what is degraded
right now (failed backends included) plus the recent offline/online history. The Settings
panel's Services section shows everything the instance can talk to in one
list, grouped as Chat backends, Image generation, Speech,
Tools, and MCP servers, with one status language: green = working,
amber = needs your action (e.g. a subscription sign-in to connect — the
button is right on the row), red = failing (hover the dot for the reason),
grey = not set up. A subscription credential is never shown as a separate
line item; its state appears on every backend that uses it. Everything recovers
automatically: within one status poll (~15 s) of a service coming back, toggles
re-enable, dots turn green, and the alert clears. Each transition is also logged by the core
(service down / service recovered) for debugging.
Chapter 7
Profiles
A profile is a one-click starting point. Selecting one applies a persona prompt plus a bundle of style settings (tone, length, creativity, and format) for the current session.
Built-in profiles
| Profile | For |
|---|---|
| Brainstorm Ideas | Explore options and tradeoffs. Attaches a diverge-then-converge skill. |
| Draft Content | Create reusable content. Attaches an audience-first drafting skill. |
| Write Code | Build, debug, and refactor. Attaches a run_code vs delegate_task skill. |
| Learn Topic | Learn a subject in the chat; offers a how-to or mini-book if you want a PDF. |
| Write a How-to | A short how-to (one skill, one sitting) compiled to a PDF. |
| Write a Mini-book | A focused mini-book on one subject, chapter by chapter. |
Selecting a profile
The welcome screen shows up to eight profile cards total; every configured profile is also in the Profile dropdowns. The synthetic ✨ Auto entry described below is shown first when present and counts as one of the eight cards. Pick a card to apply its persona and style. In Settings, the Profile dropdown remains visible while its accordion is collapsed; selecting a profile expands it. The expanded area contains the style controls, the nested Session Instructions editor, and a read-only Skill viewer when the selected concrete profile has an attached procedure.
Each assistant response stores the canonical profile ID and a label snapshot for that
turn. Reopening a chat selects the profile used by its newest visible assistant response.
A newest No profile response restores No profile. If a custom profile was removed, its old
responses keep the saved label but the selector falls back to No profile. If that response
was Auto-routed (labeled Auto → <Profile Label>, or a bare
Auto), reopening restores the Auto profile selection instead of the concrete
persona it happened to route to that turn. The persona, style bundle, generated system
prompt, and prompt edits are not persisted as profile settings.
Every built-in welcome card attaches a matching skill (a procedure, not extra persona text). Write a How-to and Write a Mini-book are the writing skills: the assistant authors the teaching, an agent only typesets, and generated cover images from the chat are copied into the workspace so they can actually land in the PDF. After a compile, the PDF and page previews are attached to the assistant message so you can open them in the chat. Pick one of those cards before asking for a document.
You can define your own profiles to extend or override the built-ins, and attach a
skill (a Markdown procedure) with an optional skill field or by
using the same id — see
Chapter 11 (Custom profiles and Skills). Starter overlay
skills for the example catalog live in
docs/skills.example/.
Auto profile
The Profile dropdown also offers a synthetic ✨ Auto entry. auto
is a reserved id — you cannot define your own profile with that id. Picking Auto is an
autopilot: it also switches the chat to the Auto (Smart Router)
model (when that model is available), so one choice hands both the backend model and the
per-turn persona to the router. Auto only picks a persona when the chat's model is
Auto (Smart Router). On such a turn the profile catalog is
offered to the router: the semantic router classifier, when
configured in schema mode and answering in time, picks the best-fitting profile;
otherwise (Arch-Router, or the classifier disabled or too slow) a deterministic task-affinity
heuristic matches the turn's task to a profile whose description reads like that kind of work.
Either way that profile's persona and style are used and the response is labeled
Auto → <Profile Label>, just like a manually-picked profile. When no profile
fits (e.g. a trivial greeting) or the chat uses a manually-selected model, Auto behaves exactly
like No profile and the response is simply labeled Auto.
Chapter 8
Response Controls
The Settings panel shapes each response. These are session-level controls; a profile simply sets several of them at once.
Tone
Single-select: Default, Professional, Friendly, Direct, Executive. Controls tone only. Default adds no tone instruction.
Response Length
Single-select target for the output budget:
| Setting | Target output tokens |
|---|---|
| Short | 2000 |
| Medium | 8000 |
| Detailed | 12000 |
Creativity
Single-select: Precise, Balanced,
Exploratory. This always changes the style prompt. It sends
temperature and top_p sampling parameters only when the
backend is configured to honor sampling (CHAT_HONORS_SAMPLING=true, or
honors_sampling on the endpoint). Otherwise those parameters are intentionally
omitted.
Format
Single-select shape for the output: Natural, Structured, Bullets, Step-by-step, Table-first.
Session instructions
Session Instructions is nested inside the expanded Profile controls. It keeps an independent eye button for showing or hiding the editor and a reset button for restoring the generated instructions. Omni Chat assembles a single upstream system message from, in order: app safety invariants, hidden real-time context, assistant soul, editable session instructions, retrieved memory, retrieved sources, and summaries. The panel shows only the editable session instructions generated from project/profile/style settings; managed safety and SOUL are hidden.
Session-instruction edits apply to the current session and are not written to SQLite. Reloading discards those edits and regenerates the defaults for the restored current profile. App safety and assistant soul are enforced separately by the core and cannot be removed from the Settings textarea.
Skill
When the selected concrete profile has an attached skill, a Skill row appears after Session Instructions. Its independent eye button reveals the exact managed Markdown in a read-only panel. It starts hidden after reload and keeps its disclosure state while you switch profiles during that page session. Default and Auto show no Skill row because neither has one stable attached procedure; Auto's concrete profile is chosen per turn.
Chapter 9
Memory
Memory lets Omni Chat remember durable facts, preferences, and decisions across turns and chats — separate from the raw message history.
Your private assistant context
On first sign-in, an optional, skippable form lets you tell Omni Chat the name it should use, pronouns, location/timezone, role, organization, interests, goals, and a short “about” note. This small private profile is included on every normal turn independently of Memory, so basic identity does not depend on search ranking. Edit it later from Edit Profile. It is visible only to your own session and never in the instance member list.
The Memory toggle
- On (default) — the assistant can save durable memories and recalls the relevant ones on later turns, injected into its prompt.
- Off — chat history is still stored, but nothing is written to or read from durable memory for that chat.
The toggle lives in the composer's Tools menu and is remembered per chat.
Asking the assistant to remember
With Memory on, just tell the assistant — for example, "Remember that I'm vegetarian" — and it saves a durable memory; a "Memory · Saved" notification in the top-right corner confirms it for about 15 seconds, with Undo in case it saved something you didn't want kept. In a later chat, ask "What kind of food do I eat?" and it answers from that memory without being told again. Say "Forget that I'm vegetarian" to remove it.
Memory kinds
preference, fact, decision, constraint.
Scopes
| Scope | Visible to |
|---|---|
user | You, across all your projects (private). |
project | Every chat within the project. |
workspace | Members of a shared group (arrives with sharing). |
Managing memories
Open Manage Memories from the top-right user menu to see everything the assistant remembers about you. From there you can filter by scope, edit a memory's text or kind, pin the ones that should always be eligible for context, add a memory manually, delete any of them, and see each one's provenance — the chat and exact user-message evidence it came from. Use the status filter to audit quarantined automatic memories and restore only the ones you trust. Your memories are private to your account.
Smarter recall (optional)
If the operator has configured an embeddings endpoint (CHAT_EMBEDDINGS_URL and
friends), memory recall becomes semantic: the assistant surfaces the memories most
relevant to what you're currently asking, and saving something you already mentioned — even in
different words — updates the existing memory instead of creating a near-duplicate. Without an
embeddings endpoint, recall still uses lexical, entity, and dotted-acronym matching; it never
fills the context with unrelated recent rows. Either way, memory works.
Automatic memories
If the operator has configured the auxiliary model (CHAT_AUX_URL and
CHAT_AUX_MODEL), the assistant quietly captures durable facts you mention —
lasting preferences, your ongoing projects, decisions, and constraints — without you having to
say "remember this". This runs in the background after a reply, so it never slows a
conversation, and it uses a separate small model rather than your main chat model. It is
conservative by design: it saves only clearly-stated, lasting facts, skips sensitive data
(passwords, card or account numbers, secrets), and requires an exact quote from one of your
messages as evidence. The assistant's own answer can never become a fact about you. Every
automatic memory is announced with a brief "Memory · Remembered" notification in the
top-right corner a few seconds after the reply — click Undo to drop it on the spot, or View to open
the Memories panel (hovering the chip keeps it open). Every
automatic memory shows up in the Memories panel with its evidence, so you can review, edit,
quarantine, restore, or delete it. Existing unsupported automatic memories are quarantined on
upgrade. Auto-capture is off unless the operator turns it on, and you can disable it per chat
with the memory save policy.
Staying coherent in long chats (automatic context compaction)
When a conversation gets very long, it eventually exceeds the model's context window and the oldest messages would normally fall out of view. Instead, the assistant keeps a short rolling recap of the earlier part of the chat, refreshed quietly in the background as the conversation grows — each refresh folds the new turns into the previous recap, so even the very beginning of a long chat is never forgotten. When old messages no longer fit, the assistant sees the "Earlier in this conversation…" recap in their place, so it stays on topic instead of losing the thread. This works even without the optional auxiliary model: the recap is then generated by the chat's own model between turns, and it never slows down a reply.
Your transcript is never changed — you can always scroll back through the full conversation. When compaction was active on the latest response, a subtle dashed divider marks the boundary: messages above it reached the model only as the recap. Click View summary on the divider to read exactly what the model is told, and the response's info popup shows a "Context: Compacted" row. You can still edit, regenerate, or delete any message, including ones above the boundary — doing so simply discards the now-stale recap and rebuilds it in the background a turn later.
Tidying up over time
As your memories accumulate, the assistant occasionally tidies them in the background (with the same optional memory model, rate-limited so it runs at most once in a while). It merges duplicates, retires memories that a newer one has contradicted or made obsolete, and distills recurring themes into higher-level “insight” candidates. Because those candidates are model inferences rather than direct user quotes, they stay quarantined until you restore them. Your pinned memories are never touched. Nothing is ever permanently erased: a retired memory keeps a link to the one that replaced it, so the whole history stays auditable and reversible. The result is that a long-lived memory set gets sharper and less cluttered over time instead of filling up with near-duplicates.
Remembering across conversations
Sometimes the thing you need was discussed in a different chat entirely. When cross-chat recall is available (it needs the embeddings model configured), the assistant can pull in a short recap of your other conversations that are relevant to what you're asking now — shown to it as "From earlier conversations…". So if you start a fresh chat and ask about a project, decision, or topic you worked through days ago, it can pick up the thread instead of starting from nothing. It only ever draws on your own chats, surfaces just the few most relevant ones, and follows the per-chat Memory switch — turn Memory off for a chat and no past-conversation context is pulled in.
Chapter 10
Sources & Retrieval
Retrieval grounds answers in documents you've added, attaching source cards that show which documents and chunks were used.
The Sources section
Any answer that used sources — indexed documents, web search results, or fetched pages — shows a collapsed Sources · N row directly below it, with up to three of the source domains in the header. Click it to expand the source cards; it starts collapsed every time and there is nothing to switch on.
Search
The composer's Search toggle is on by default. Turning it off keeps that chat away from the web entirely: the model is not offered web search, image search, URL fetching, or image fetching, a link you paste is not opened, and the model is told that you turned web access off. It is a lighter privacy option than going off the record — the chat is still saved. Delegated agents and MCP tools have their own toggles and are not affected. Set whether new chats start with Search on under New chat defaults in the right settings panel. The toggle only appears when web search or web fetch is configured.
Document retrieval runs when the question looks source-worthy and the corpus has
indexed documents (the configuration-level sources.mode, default
auto); it has no switch in the UI.
Adding documents
Attach documents via upload or drag-and-drop. Once indexed, they become available to retrieval. Retrieval requires an embeddings endpoint to be configured (see Chapter 11).
If the corpus is empty, Omni Chat will not retrieve or fabricate citations. Source cards appear only when retrieval actually ran and returned context.
Chapter 11
Configuration Reference
Omni Chat is configured through environment variables set at startup, plus two optional JSON files for backends and profiles.
Environment variables — network & storage
| Variable | Default | Purpose |
|---|---|---|
CHAT_BIND | 127.0.0.1 | Bind address. Use 0.0.0.0 only when intentionally sharing the instance on a LAN. |
CHAT_PORT | 0 | 0 chooses an ephemeral port and prints it in the READY line. |
CHAT_DB_PATH | core/data/app.db via make dev | SQLite database file path. |
Environment variables — backend & model
Configurable local network discovery. Native apps default to lan; standalone servers default to local. Optional "discovery": {"mode": "lan"} in endpoints.json accepts only lan, local, or off. Explicit CHAT_ENDPOINT_DISCOVERY wins over the file, then the platform default applies. Server operators edit this file or environment and restart. The admin-only Settings → AI Servers → Network discovery auto toggle maps on to lan, off to local, and applies with Save & Restart. A saved off survives until the toggle is explicitly changed. Environment overrides disable the toggle with an explanation. Saving without manual connections is valid; older clients omitting discovery preserve it. Manual connections remain usable in every mode and discovered servers remain ephemeral. The same AI Servers page also has a Scan now button: an on-demand full pass (loopback, named local hosts, and the attached network) regardless of the auto toggle, which also surfaces servers locked behind an API key as Locked · needs API key instead of hiding them. Both the auto toggle and Scan now sweep only this machine's own private-network interfaces and mDNS-named .local hosts — a server reachable only over an overlay network such as Tailscale (its 100.64.0.0/10 range isn't "private" here, and it isn't mDNS-advertised) will not appear in either; add it by its reachable address with Add server instead.
First-run network ask. A fresh desktop or iPad install does not search your network while you create the owner account, so the system's Local Network prompt doesn't appear out of nowhere. Right after setup, a card explains what the search is for and that your device will ask next. Look for servers turns on network discovery at once (the prompt follows); Not now keeps Omni Chat to this device — turn Network discovery on later in Settings → AI Servers, or press Scan now.
On iPad, Apple prompts for Local Network access on first network use. After denial, enable Omni Chat in Settings → Privacy & Security → Local Network and return to the app. Taking a photo prompts for Camera access before the camera opens. Denying it leaves the app open, with a note and a way to attach an existing file. Dictation prompts for Microphone and Speech Recognition. Those descriptions have to be in the app: iOS quits Omni Chat instead of asking when they are missing. Explicit native denial is surfaced beside the toggle; a routing error alone is not evidence of denial. Discovery starts on first active launch, cancels LAN work while inactive, and refreshes interface targets on resume. Scans retry every five seconds during the first active minute, then every 60 seconds, coalescing without overlap. Permission dialogs never block startup or the local model. Apple DNS-SD browses _ssh._tcp, _sftp-ssh._tcp, _http._tcp, and _workstation._tcp; verified address results supply names. iOS does no raw multicast or reverse-PTR and needs no restricted multicast entitlement. Non-Bonjour servers remain discoverable by IP via the bounded IPv4 subnet scan.
| Variable | Default | Purpose |
|---|---|---|
CHAT_ENDPOINT_URL | unset | OpenAI-compatible base URL. Accepts a comma-separated list to combine backends (each shares CHAT_API_KEY). Ignored when endpoints.json is present. When unset, Omni Chat also scans for keyless local OpenAI-compatible servers and offers their models in the picker. |
CHAT_ENDPOINT_DISCOVERY | native: lan; server: local | lan (this machine, mDNS .local names, and each attached IPv4 /24), local (loopback and host.docker.internal), or off. Verified aliases of the same host and port are grouped when their complete model catalogs match, keeping all models together. DNS aliases such as host.docker.internal are recognized; the Docker host and container remain separate. Configured chat backends take precedence and saved interface selections continue to resolve. Matching model IDs on different machines do not cause merging; addresses without a discoverable host relationship remain separate. An auxiliary model on the same URL (for example oMLX on port 8000) does not hide the rest of that server's models. Hostname labels use Apple DNS-SD on iOS and Go mDNS reverse-PTR elsewhere. Native apps default to lan; standalone servers default to local. The environment overrides endpoints.json → discovery.mode, which overrides the platform default. |
CHAT_ENDPOINT_DISCOVERY_PORTS | built-in popular list | Optional comma-separated port list (max 32) replacing the defaults (Ollama 11434, LM Studio 1234, vLLM 8000, llama.cpp 8080, and other well-known OpenAI-compatible ports). |
CHAT_API_KEY | unset | Optional bearer token. Kept out of SQLite — keep secrets in env/keychain, not plaintext storage. |
CHAT_MAX_TOKENS_FIELD | max_tokens | Set to max_completion_tokens for endpoints/models that require it. |
CHAT_CONTEXT_WINDOW | 8192 | Approximate input+output context window used for preflight budgeting. |
CHAT_VISION_BUDGET | balanced | Automatic image-input budget for an environment-backed endpoint: low, balanced, or high. |
CHAT_MAX_OUTPUT_TOKENS | unset | Optional hard cap on generated tokens; response-length presets and model metadata are clamped to it. |
CHAT_CONTEXT_SAFETY_MARGIN | 512 | Tokens reserved (subtracted) from the input budget. |
CHAT_RESPONSE_HEADER_TIMEOUT | 2m | How long to wait for the endpoint's first response headers. On streaming backends these arrive with the first token, so this is effectively a time-to-first-token limit — raise it (e.g. 5m) for slow self-hosted routers that cold-load large models, or a turn errors with a timeout while the model is still loading. Accepts a Go duration (5m, 90s) or a bare number of seconds (300). |
CHAT_STREAM_IDLE_TIMEOUT | 60s | Maximum silence between streamed tokens before a turn is treated as a stalled upstream and ended (finish_reason=error, retryable). The streaming body has no read deadline otherwise, so a backend that wedges mid-response without closing the socket would hang the turn. A healthy stream never trips it; raise it only for extremely slow token generators. Accepts a Go duration or a bare number of seconds. An interrupted response can be resumed in place with Continue. |
CHAT_STREAM_DETACH_GRACE | 5m | How long a generation keeps running with no browser attached before it is cancelled. Generation is detached from the connection: navigating around the app, reloading the page, or briefly closing the laptop never kills a turn — the client re-attaches and replays what it missed, and the generating chat pulses in the sidebar. A genuinely closed browser keeps the backend busy for up to this grace, so tune it to taste. Accepts a Go duration or a bare number of seconds; 0 restores the old cancel-on-disconnect behaviour. The Stop button always cancels immediately. |
CHAT_HONORS_SAMPLING | false | When true, Creativity sends temperature and top_p. |
CHAT_EXTRA_BODY | unset | JSON object merged into the top-level chat request body for static endpoint parameters. |
CHAT_THINKING_ON_BODY | unset | Advanced adapter override: JSON object merged into the top-level chat request body when the per-chat Thinking switch is on/auto. |
CHAT_THINKING_OFF_BODY | unset | Advanced adapter override: JSON object merged into the top-level chat request body when the per-chat Thinking switch is off. |
CHAT_THINKING_BUDGET_TOKENS | unset | Reasoning-guard token cap per response. A reasoning model that thinks past this budget — or starts repeating itself — is stopped gracefully with a "got stuck thinking" note offering Retry without thinking and Regenerate, instead of burning the whole output window. The budget is added on top of the response-length preset rather than taken out of it, so thinking never starves the reply, and an overrun is retried once with thinking off (seeded with the model's own notes) instead of losing the turn. Unset applies the framing profile's default — 1024 for sub-2B models, 2048 up to 10B, 16384 above that and for models of unknown size. Also settable per endpoint (thinking.budget_tokens) or per model (thinking_overrides[].budget_tokens) in endpoints.json. |
CHAT_PARSE_THINK_TAGS | false | Advanced adapter override: parse streamed <think>...</think> answer text into the reasoning channel and strip it from the answer. |
Reasoning models often have endpoint-specific parameters. Omni Chat infers common
adapters such as Qwen served by llama.cpp and DeepSeek V4 served by ds4.c.
For OMLX endpoints, Omni Chat reads the optional /v1/models/status metadata and
enables the switch when thinking_default is either true or
false; both values mean the model template supports
enable_thinking. Custom endpoints can still define thinking request bodies in
endpoints.json; the switch sends on_body for auto and
off_body for off. Models with no inferred or configured adapter keep
the switch disabled.
Thinking effort (Settings → Thinking effort) is an account-wide Low /
Balanced / High control, default Balanced. It changes how thoroughly supporting models
reason — it is not a token cap, and it is not per-chat. Qwen 3.8 thinks at its maximum
(xhigh) unless this is set; Omni Chat maps Balanced to medium and
High to xhigh (the template rejects the name high). The same
preference is applied to agent handoff and to omni-chat launch. Models without
a native effort knob are left unchanged. Operators can opt a model in with
thinking_overrides[].effort: "qwen38".
ChatGPT (Codex) models reason through the Responses API's reasoning block.
Omni Chat sends it for every Responses connection that has no thinking settings
of its own: Thinking on asks for a reasoning summary (so the thinking streams) at the effort
above — Low → low, Balanced → medium, High → high.
These models have no "none" effort, so Thinking off drops to low and hides the
summary.
GLM-4.5 and later (GLM-4.5, GLM-4.6, GLM-5.x) are hybrid reasoners toggled through the
chat template, so a self-hosted GLM gets a fully controllable switch backed by
chat_template_kwargs.enable_thinking. Served with
--reasoning-parser glm45 the trace arrives on the native reasoning channel;
served without it, Omni Chat sniffs the leading <think> tag and routes the
trace anyway. This inference is limited to self-hosted backends on purpose — Z.ai's hosted
API toggles the same models with a different shape
(thinking: {"type": "enabled"}), so an endpoint pointed at it should set
thinking explicitly. GLM-4 and older have no thinking mode and correctly show no
switch.
Environment variables — embeddings
| Variable | Purpose |
|---|---|
CHAT_EMBEDDINGS_URL | Embeddings endpoint URL. |
CHAT_EMBEDDINGS_MODEL | Embedding model name. |
CHAT_EMBEDDINGS_DIM | Embedding dimension; must match the model. |
CHAT_EMBEDDINGS_DIM must match your embedding model. Changing it after
documents are indexed invalidates the stored vectors and forces a reindex.
Environment variables — auxiliary model
One optional small, fast, non-thinking model powers the background intelligence:
chat-title generation, memory extraction, rolling summaries, consolidation, and — when no
dedicated classifier is set — smart-router classification. It runs off the response path and
never uses the main chat model. (Rolling chat summaries are the one exception when no
auxiliary model is configured: they then fall back to the chat's own model between turns, so
long-chat compaction always works.) A 1–2B non-reasoning instruct model is ideal — the
recommended default is LFM2.5-1.2B-Instruct (Liquid AI, ~731 MB at Q4_K_M),
served with llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:Q4_K_M; see
.env.example for smaller and larger alternatives.
| Variable | Default | Purpose |
|---|---|---|
CHAT_AUX_URL | unset | OpenAI-compatible chat endpoint for the auxiliary model. Unset disables every aux-powered feature. |
CHAT_AUX_MODEL | unset | Model name (small, fast, non-thinking). Required when the URL is set. |
CHAT_AUX_API_KEY | unset | Optional bearer token for the endpoint. |
CHAT_AUX_ROUTER | true | Set false/0 to stop the aux model from serving as the smart-router classifier. |
CHAT_AUX_SHOW_IN_MODEL_DROPDOWN | true | Set false/0 to hide the aux model from manual chat selection while keeping its background jobs enabled. |
CHAT_MEMORY_*
The auxiliary model replaces the former CHAT_MEMORY_URL/CHAT_MEMORY_MODEL/CHAT_MEMORY_API_KEY
variables, which are no longer read. Point CHAT_AUX_URL/CHAT_AUX_MODEL at the same
small model to restore the memory jobs; setting CHAT_MEMORY_* now only logs a startup warning.
Environment variables — config locations
| Variable | Default | Purpose |
|---|---|---|
CHAT_CONFIG_DIR | OS config dir | App config directory. Defaults to ~/.config/omni-chat on macOS/Linux (unless Linux XDG_CONFIG_HOME is set); Windows uses the user config directory. |
CHAT_PROFILES_PATH | <config-dir>/profiles.json | Override only the profiles file. |
CHAT_SOUL_PATH | <config-dir>/SOUL.md | Override only the assistant personality Markdown file. |
CHAT_SKILLS_PATH | <config-dir>/skills | Override the overlay directory of *.md skills. Missing default dir is fine. Starter: docs/skills.example/. |
CHAT_ENDPOINTS_PATH | <config-dir>/endpoints.json | Override only the multi-backend file. |
CHAT_ROUTER_PATH | <config-dir>/router.json | Override only the Auto (Smart Router) policy file. |
CHAT_AGENTS_PATH | <config-dir>/agents.json | Override only the external-agent file that enables the Agent handoff (delegate_task) tool. |
CHAT_MCP_PATH | <config-dir>/mcp.json | Override only the MCP-server file that enables MCP tools (see the README's "MCP tools" section and docs/mcp.example.json). |
CHAT_OAUTH_CALLBACK_BIND | 127.0.0.1 | Bind host for the temporary OAuth callback listener used by subscription sign-ins (port 1455 for Codex). Containers set 0.0.0.0 so the published port reaches it — the bundled compose.yaml does this. Publishing the host side is opt-in: set OMNI_OAUTH_PUBLISH_PORT=1455 in .env (default: an ephemeral port, so the automatic callback is off and sign-in uses the paste fallback). |
Environment variables — semantic router classifier
| Variable | Default | Purpose |
|---|---|---|
CHAT_ROUTER_CLASSIFIER_URL | unset | OpenAI-compatible endpoint of a dedicated router-classifier model (e.g. Arch-Router). Unset falls back to the auxiliary model (CHAT_AUX_*) when configured, else heuristics only. |
CHAT_ROUTER_CLASSIFIER_MODEL | unset | Required when the URL is set; the one explicit classifier model id. The router overlay defaults it to arch-router-1.5b. |
CHAT_ROUTER_CLASSIFIER_FORMAT | auto | Classifier protocol adapter: auto (detect arch-router in the model id), arch (Arch-Router native routes), or schema (JSON-schema classification for general instruct models). |
CHAT_ROUTER_CLASSIFIER_API_KEY | unset | Optional bearer token for the classifier endpoint. |
CHAT_ROUTER_CLASSIFIER_TIMEOUT | 3000ms | Race deadline; accepts 100ms–10s. Requires a Go duration unit, e.g. 3000ms — a bare number like 3000 fails to parse and is fatal at startup. |
Environment variables — web search
| Variable | Default | Purpose |
|---|---|---|
CHAT_WEB_SEARCH_PROVIDER | searxng | Provider implementation. |
CHAT_WEB_SEARCH_URL | unset | Provider base URL (SearXNG); setting it enables the model tool. Ignored when endpoints.json has a search section. |
CHAT_WEB_SEARCH_API_KEY | unset | Key for a hosted provider (tavily); setting it also enables the tool, since a hosted provider needs no URL. |
CHAT_WEB_SEARCH_TIMEOUT | 10s | Per-search timeout. |
CHAT_WEB_SEARCH_MAX_RESULTS | 5 | Result limit from 1 to 10. |
CHAT_TOOL_CALLING | true | Set false for an environment-backed endpoint without tool support. |
CHAT_WEB_FETCH | false | Set true to enable the fetch_url tool. |
CHAT_WEB_FETCH_PROVIDER | direct | direct (in-core, SSRF-guarded) or reader (external reader endpoint). |
CHAT_WEB_FETCH_URL | unset | Reader endpoint base URL; required when the provider is reader. |
CHAT_WEB_FETCH_TIMEOUT | 15s | Per-fetch deadline, 1s–60s. |
CHAT_WEB_FETCH_MAX_BYTES | 2097152 | Per-page download cap before truncation. |
CHAT_WEB_FETCH_USER_AGENT | browser UA | Overrides the User-Agent used by the direct provider. |
CHAT_WEB_FETCH_ALLOW_PRIVATE | false | Local dev only: relax the SSRF guard. |
CHAT_IMAGE_GEN_URL | unset | Image backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables the generate_image tool. Ignored when endpoints.json declares images[]. |
CHAT_IMAGE_GEN_PROVIDER | openai | Backend protocol: openai (POST /v1/images/generations) or responses (the Responses API's image_generation tool; needs carrier_model, so use endpoints.json images[] for it in practice). |
CHAT_IMAGE_GEN_MODEL | unset | Optional model id passed through to the backend (the bundled sidecar ignores it). |
CHAT_IMAGE_GEN_API_KEY | unset | Optional bearer token for the image backend; env-only, never stored in SQLite. |
CHAT_IMAGE_GEN_TIMEOUT | 120s | Idle deadline per image, 10s–10m (duration or bare seconds). Streamed progress events reset it; against a non-streaming backend it is the total deadline. |
CHAT_IMAGE_GEN_STREAM | true | Request streamed progress and partial previews (OpenAI Images stream/partial_images); JSON-only backends degrade gracefully. |
CHAT_TTS_URL | unset | Speech backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables read-aloud. Ignored when endpoints.json declares tts[]. |
CHAT_TTS_MODEL | tts-1 | Speech model id passed through to the backend (omnivoice for the bundled sidecar, gpt-4o-mini-tts on the OpenAI API). |
CHAT_TTS_API_KEY | unset | Optional bearer token for the speech backend. Env-only, never stored in SQLite. |
CHAT_TTS_VOICES | unset | Optional comma-separated voice list, pinning the voice picker instead of asking the backend. |
CHAT_TTS_DEFAULT_VOICE | unset | Voice used when an account has no saved preference. |
CHAT_TTS_TIMEOUT | 300s | Per-synthesis-call deadline (1s–10m). Long replies are chunked on sentence boundaries, so this bounds each chunk; the default is sized for CPU backends, where synthesis can take minutes. |
CHAT_STT_URL | unset | Transcription backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables voice input's server engine. Ignored when endpoints.json declares stt[]. |
CHAT_STT_MODEL | whisper-1 | Transcription model id passed through to the backend (large-v3-turbo for the bundled sidecar, whisper-1 on the OpenAI API). |
CHAT_STT_API_KEY | unset | Optional bearer token for the transcription backend. Env-only, never stored in SQLite. |
CHAT_STT_TIMEOUT | 120s | Per-transcription-call deadline (1s–10m). |
Search uses structured function calls and persists real results as cards. Snippets are untrusted summaries, not full-page content. The per-chat Search toggle gates them: off withholds web search, image search, and both fetch tools for that chat. Saved cards appear in the collapsed Sources section under each answer. Each assistant response records whether web search was unavailable, available but unused, successful, successful with no results, or failed; the Generation details panel reports that outcome.
With CHAT_WEB_FETCH=true a companion fetch_url tool lets the model
retrieve one specific URL — a web page, article, or (via a capable reader) a YouTube transcript —
extract its readable text, and summarize or quote it. The direct provider fetches
in-core behind an SSRF guard that refuses private, loopback, link-local, and cloud-metadata
addresses; the reader provider delegates to an external or self-hosted reader
endpoint for JS-heavy or anti-bot sites. Fetched content is treated as untrusted, persists as a
source card in the same Sources section as web-search cards, and its
outcome is folded into the single Web access line of the Generation details panel
alongside web search (that line reads Used whenever either tool ran).
For time-sensitive questions the model can limit a search to the past day, week, month or year;
with Tavily, a past-day or past-week search uses its news index, so asking for today's headlines
returns that day's articles rather than older roundups.
web_search and fetch_url share one bounded tool-call budget per turn. When
the fetch tool is enabled and your message contains a URL, the core fetches it automatically before
the model runs, so summarizing a pasted link works even if the model would otherwise try to search
for it.
The same setting lets the assistant show you a picture that already exists rather than drawing one — ask "show me a 3-month candlestick chart of TSLA" or "find me a picture of a cat and dog" and it looks one up, downloads it, and puts it in its reply. This works even on an instance with no image generation configured. (Looking pictures up by keyword needs a search provider that supports image search; without one the assistant can still show images whose address it already has, such as a chart endpoint or an image on a page it just read.) The image is saved with the message rather than hotlinked, so it stays put if the source disappears, comes along when you export the chat, and never tells the origin site who is looking at it; you can also ask for edits to it afterwards. A fetched image is always labelled Fetched with a link to where it came from, so it is never confused with one the assistant generated. If a URL turns out to be a web page rather than an image, the assistant says so instead of showing something broken. It is unavailable in off-the-record chats, since the image would have to be stored.
Every turn also carries a hidden, system-injected context block (separate from editable
Session Instructions): the current date and time in your browser's timezone, best-effort details
about the client device's OS/version/CPU architecture, and a short directive.
When the web_search tool is available the directive tells the model to search for
current or possibly-changed facts and to trust fresh results over its own memory; the
fetch_url tool adds guidance to read a specific URL for summarizing. When neither
web tool is available it instead tells the model it has no web access and should flag that its
knowledge may be out of date. The browser sends its IANA timezone name plus structured OS
metadata from Client Hints or coarse user-agent fallbacks; the raw user-agent string is never
sent. The server keeps the clock, treats device version/architecture as approximate and
distinct from a remote core host, and persists neither value. The prompt explicitly tells the
model to use the detected OS for commands, installation, and debugging unless you name a
different target. No configuration is required.
With an image backend configured, tool-capable models are offered
generate_image and edit_image tools (no per-chat toggle). The core
calls the backend server-side, stores the produced image with the conversation (bytes never
enter the model context), and streams it to the chat as it lands — including live progress and
blurred partial previews when the backend supports streamed partials (the bundled sidecar and
the Codex backend both do; those progress events also reset the
CHAT_IMAGE_GEN_TIMEOUT idle deadline, so long generations don't time out while
they're visibly working). The tool has its own budget of two
images per response, separate from the web tools; the assistant response's Generation details
panel records an Image generation outcome (used, failed, or not used), and each
image's caption tooltip names the backend that produced it.
Backends are named and plural, and reference a server. The simplest setup is
CHAT_IMAGE_GEN_URL — a single URL or a comma-separated list (one backend per URL,
named by host, first = default). For richer setups declare an images[] array +
default_image_endpoint in endpoints.json (see
docs/endpoints.example.jsonc): each entry names a connection — which
carries the base_url, a key (api_key/api_key_env) or an
auth OAuth block, and (via protocol) whether it speaks the OpenAI
Images API (POST /v1/images/generations with b64_json; the bundled
Z-Image-Turbo sidecar, compose.imagegen.yaml, GPU required — see
Running Omni Chat, combination g) or the OpenAI
Responses API's native image_generation tool (protocol: "responses")
— and keeps optional model, timeout, and stream fields.
The responses protocol is how a ChatGPT subscription generates images
(gpt-image-2): point the connection at https://chatgpt.com/backend-api/codex
with "auth": {"type":"oauth","provider":"openai-codex"} and
"protocol": "responses" — the same instance credential as a Codex text connection,
connected once by the owner, shared automatically when it's the same connection — plus
"model": "gpt-image-2" and a carrier_model (the text model slug that
carries the tool call, e.g. gpt-5.5). Each chat picks its backend in the settings
sidebar's Image generation select (empty = instance default; the choice is saved with
the chat), and the New chat defaults section sets the backend new chats start on. Each
backend appears separately in the service-availability panel as
Image generation: <name>.
Speech backends work the same way. The simplest setup is
CHAT_TTS_URL (single URL or comma-separated list); for richer setups declare a
tts[] array + default_tts_endpoint in endpoints.json
(see docs/endpoints.example.jsonc): each entry names a connection (any
server speaking the OpenAI speech API, POST /v1/audio/speech — the bundled
OmniVoice sidecar, compose.tts.yaml, or a hosted endpoint such as
https://api.openai.com) and keeps a required model, an optional
pinned voices list with a default_voice, and a timeout.
Users pick their voice — across every configured backend — under Voice Engine in the
settings sidebar's Voice & Dictation section (see
Read aloud). Each backend appears in the service-availability panel
under Speech.
Transcription backends work the same way. The simplest setup is
CHAT_STT_URL (single URL or comma-separated list); for richer setups declare an
stt[] array + default_stt_endpoint in endpoints.json
(see docs/endpoints.example.jsonc): each entry names a connection (any
server speaking the OpenAI transcription API, POST /v1/audio/transcriptions — the
bundled whisper sidecar, compose.stt.yaml, or a hosted endpoint such as
https://api.openai.com) and keeps a required model, an optional
default language, and a timeout. Users pick their engine and
dictation language under Dictation in the settings sidebar's Voice &
Dictation section (see Voice input).
Unlike read-aloud, nothing configured does not hide dictation — the browser fallback still
covers it where supported. Each backend appears in the service-availability panel as
Dictation: <name>.
Multiple backends — endpoints.json
Create <config-dir>/endpoints.json to combine backends. Servers are
described once in a top-level connections[] array (name,
label, base_url, a key inline as api_key or via
api_key_env, and optional auth/protocol); the
endpoints[] array then names one connection per chat backend and keeps
only the capability fields below. schema_version must be
2 — an older per-endpoint-URL file fails startup with a message pointing
at docs/endpoints.example.jsonc and this chapter; there is no automatic migration, so an existing
file from before this change must be recreated. Per-connection/per-endpoint capability fields
mirror the CHAT_* model settings and apply only to that backend, including
optional thinking adapter overrides and model-pattern overrides.
{
"schema_version": 2,
"connections": [
{ "name": "local-litellm", "label": "Local LiteLLM", "base_url": "http://localhost:4000" },
{
"name": "openai",
"label": "OpenAI",
"base_url": "https://api.openai.com",
"api_key_env": "OPENAI_API_KEY"
},
{
"name": "vllm-local",
"label": "vLLM (local)",
"base_url": "http://localhost:8000",
"api_key": "sk-local-only-not-a-real-secret"
}
],
"default_endpoint": "local-litellm",
"endpoints": [
{
"connection": "local-litellm",
"thinking": {
"on_body": { "reasoning_effort": "medium" },
"off_body": { "reasoning_effort": "none" },
"parse_think_tags": true
}
},
{
"connection": "openai",
"max_tokens_field": "max_completion_tokens",
"context_window": 128000,
"max_output_tokens": 16000,
"honors_sampling": true
},
{
"connection": "vllm-local",
"context_window": 32768,
"prompt_profile": "auto",
"model_overrides": [
{ "model_pattern": "qwen2.5-vl*", "vision": true, "tool_calling": true, "params_b": 7 },
{ "model_pattern": "lfm2-*", "params_b": 1.2 },
{ "model_pattern": "qwen3-30b-a3b*", "active_params_b": 3, "capability": "large" },
{ "model_pattern": "phi-4-mini*", "prompt_profile": "full" }
]
}
]
}
| Field | Meaning |
|---|---|
connection | Names a connections[] entry (server, credential, protocol). At most one endpoints[] entry per connection. |
connections[].name / .label | Stable id ([a-z0-9-]+, immutable once saved) other sections reference, and the display name / group label in the model picker — safe to rename freely. |
connections[].base_url | Server root supporting /v1/models, /v1/chat/completions, /v1/embeddings. |
connections[].api_key / .api_key_env | Inline secret, or the name of an env var to read it from. |
connections[].runtime | Optional server type: llamacpp, vllm, ollama, lmstudio, omlx, litellm, sglang, openai or other. Filled in by Scan now or Test connection and editable in Settings (Server type); shown as the server's tile. Informational today — it gives backend-specific handling a stable id to key on. |
max_tokens_field | Output-tokens field name (default max_tokens). |
context_window | Context size for budgeting (default 8192). |
max_output_tokens | Per-endpoint output cap; clamps the response-length setting. |
honors_sampling | Whether the backend respects temperature/top_p (default false). |
extra_body | JSON object of static top-level request params for this backend. |
sampling_defaults | Gates the built-in anti-loop sampling defaults (hybrid Qwen3.x, Gemma, gpt-oss, DeepSeek-R1, Nemotron, GLM-4.5+ model-card values): "auto" (default) applies them when the backend is recognized as self-hosted from the model's owned_by (llama.cpp, SGLang, oMLX, Ollama, LM Studio, vLLM), "on" forces them (for servers that omit owned_by, e.g. mlx_lm.server), "off" disables them. Qwen3.5/3.6 and Qwen3.8 have generation-specific thinking/non-thinking profiles. Endpoint extra_body and per-model model_overrides[].extra_body override individual keys. |
prompt_profile | How much tool surface models on this backend are given: "auto" (default) derives it from the resolved parameter size, "full" keeps the complete surface, "terse"/"minimal" force the small-model behavior. Also settable per model via model_overrides[].prompt_profile, which wins. See Small models on modest hardware. |
vision_budget | Automatic image-input ceiling: low (1024 px, 1 MP/image, 2 MP/turn), balanced (default: 1568 px, 2.5 MP/image, 5 MP/turn), or high (2560 px, 6 MP/image, 12 MP/turn). Also settable per model via model_overrides[].vision_budget. |
thinking | Optional adapter mapping for the per-chat Thinking switch: on_body, off_body, optional parse_think_tags, and optional budget_tokens — the reasoning-guard cap for this backend, added on top of the answer budget (default: the framing profile's, 1024/2048/16384). |
thinking_overrides | Model-pattern overrides for on_body, off_body, parse_think_tags, and budget_tokens. Bodies may contain mode-specific sampling parameters as well as custom template kwargs; partial nested bodies merge over the inferred adapter, and the first matching pattern wins. |
tool_calling | Optional boolean; set false when this endpoint does not support streamed tool calls (a hard cap on every model here). |
model_overrides | Per-model capability metadata, matched by a model_pattern glob on the model id: vision, tool_calling, params_b, vendor, family, prompt_profile. params_b is also read from the backend when it reports one (llama.cpp's meta.n_params) — the only size signal for an alias like lfm2-tool whose id carries no size — and an override here still wins. Wins over the built-in model catalog and id inference. An entry may also carry extra_body — request params for matching models only (e.g. a custom top_k for one family), overriding the endpoint-wide extra_body and the inferred sampling defaults per key. |
If endpoints.json exists, CHAT_ENDPOINT_URL is ignored. Connection
keys in this file are never written to SQLite, but this owner-managed file is the one place a
plaintext key may live — prefer api_key_env if you'd rather not store the secret
in the file. A ready-to-copy, annotated starter template is in
docs/endpoints.example.jsonc — the app never loads the .jsonc
itself (it's plain JSON with // comments); copy the parts you need into your
real endpoints.json and delete the comments, or build the file from the desktop
app's AI Servers page instead.
Assistant soul — SOUL.md
The built-in assistant soul makes Omni Chat warm, thoughtful, conversational, and lightly
playful by default. To customize it, create <config-dir>/SOUL.md or point
CHAT_SOUL_PATH at a Markdown file. A missing default file is fine and uses the
built-in soul; an explicit path must exist and must not be empty. SOUL is personality text,
not a safety boundary: the app safety invariants are still hard-coded and injected before the
single upstream system message is sent.
Custom profiles — profiles.json
Extend or override the built-in profiles by creating <config-dir>/profiles.json.
{
"schema_version": 1,
"profiles": [
{
"id": "review_code",
"label": "Review Code",
"description": "Find correctness, maintainability, security, and test gaps.",
"persona_prompt": "Review code for correctness, maintainability, edge cases, security and privacy risks, and missing tests. Prioritize concrete findings...",
"icon": "code",
"style": {
"tone": "direct",
"response_length": "detailed",
"creativity": "precise",
"format": "structured"
}
}
]
}
| Field | Meaning |
|---|---|
id | Stable identifier in lower snake_case. |
label / description | Card title and one-line summary. |
persona_prompt | The persona text added to editable session instructions. |
icon | Optional; must name a bundled icon in lower kebab-case. Unknown names fall back to a generic icon. |
skill | Optional skill id. Empty means “if a skill has this profile’s id, attach it.” An explicit id that is not in the skills catalog fails startup. |
style | A bundle of tone, response_length, creativity, format using valid option keys. |
Built-ins load first; a matching id overrides in place, and new IDs append.
The welcome screen shows at most 8 cards total, with Auto first when present; the dropdowns
still list every profile. A missing default profiles.json
is fine and uses built-ins; an explicit CHAT_PROFILES_PATH must exist and is
validated at startup. A starter file is in docs/profiles.example.json; try it
with CHAT_PROFILES_PATH=docs/profiles.example.json make dev.
Skills — skills/*.md
A profile is who the assistant is this session (persona + style). A skill
is a procedure it should follow. Skills are Markdown files, injected as a managed
Skill: block that is not shown in the session-instructions textarea. The resolved
skill for the selected profile is separately visible as read-only Markdown in the expanded
Profile controls. Keep persona_prompt short; put process in a skill.
A profile attaches a skill in two ways: an explicit "skill": "write_howto"
field, or the same id (a profile id: "review_code" loads a skill named
review_code if one exists). An explicit skill id that is missing from the catalog
fails startup before READY.
Built-in skills match the six welcome-card profiles:
brainstorm_ideas, draft_content, write_code,
learn_topic, write_howto, and write_minibook. Overlay
more as *.md under
<config-dir>/skills/, or point CHAT_SKILLS_PATH at a
directory. A missing default directory is fine; an explicit path must exist. Each file is at
most 8 KiB. The same id replaces a bundled skill. Do not put a README.md in
that directory — every .md file is loaded as a skill.
---
id: review_code
name: Review Code
---
Review a code change for correctness, tests, and safety. You are a reviewer, not the author.
A starter overlay is docs/skills.example/: same-id files for
docs/profiles.example.json (review_code, debug_issue,
plan_project, and the rest). Those example profiles leave skill
empty so profiles-only startup still succeeds. Copy both:
cp docs/profiles.example.json ~/.config/omni-chat/profiles.json
mkdir -p ~/.config/omni-chat/skills
cp docs/skills.example/*.md ~/.config/omni-chat/skills/
Or try both for one run:
CHAT_PROFILES_PATH=docs/profiles.example.json CHAT_SKILLS_PATH=docs/skills.example make dev.
Typst templates for how-tos and mini-books seed only when the resolved skill is
write_howto or write_minibook; a custom skill gets the prompt, not
those templates.
Custom router policy — router.json
Customize how the Auto (Smart Router) model tags and picks
models by creating <config-dir>/router.json.
{
"schema_version": 1,
"routes": [
{ "name": "code", "description": "writing, debugging, or reviewing code or technical artifacts" }
],
"tag_rules": [
{ "pattern": "*coder*", "tags": ["coder", "code"] },
{ "pattern": "*mini*", "endpoint": "openai", "tags": ["small", "fast"] }
],
"policies": [
{
"task": "code",
"prefer": [
{ "tags": ["coder"], "min_context": 32768 },
{ "tags": ["general"] }
]
}
]
}
| Field | Meaning |
|---|---|
schema_version | Must be 1. |
classifier | Optional structured section (url, model, format, timeout_ms) configuring the semantic classifier described below. classifier_model is a deprecated top-level alias for classifier.model. |
classifier.format | Adapter selector: auto (detect arch-router in the model id), arch (Arch-Router native routes), or schema (JSON-schema classification for general instruct models). Env equivalent: CHAT_ROUTER_CLASSIFIER_FORMAT. |
routes[] | Optionally overrides a built-in task's description shown to the classifier. Each entry is { name, description }, where name must be one of the task classes (code | reasoning | simple_fast | creative | general). Providing any routes[] replaces the full built-in list, not just the named entries. |
tag_rules[] | pattern (glob over the model id), optional endpoint (glob scope), and tags (lower snake_case). A model's tags are the union of every matching rule, built-ins included. |
policies[] | task (code | reasoning | simple_fast | creative | general) and an ordered prefer list of selectors (tags and/or min_context). A user policy replaces the built-in policy for the same task. |
A missing default router.json is fine and uses built-in tag rules and
policies — Auto still works. An explicit CHAT_ROUTER_PATH must exist and is
validated at startup (invalid JSON, unknown fields, a bad schema_version, empty
patterns, and invalid or duplicate policy tasks all fail before the READY
handshake). A starter file is in docs/router.example.json; try it with
CHAT_ROUTER_PATH=docs/router.example.json make dev.
How Auto tags and routes models
Every available model gets a set of lower-snake_case tags before Auto picks
one for a turn. Tags come from three layers, and later layers only add tags — none of them
remove a tag a candidate already has:
- Built-in glob rules match substrings of the model id:
*coder*/*codestral*/*devstral*/ … →coder;*deepseek-r1*/*qwq*/*reason*/*o1*/*o3*/*gpt-oss*→reasoning;*mini*/*small*/*phi*/*1b*…*8b*→fast;*70b*/*large*/*opus*/*gpt-4*/*gpt-5*→large. - Computed tags are attached automatically, not from glob rules:
general(every candidate),large_context(context window ≥ 64K), and capability tagsvision,reasoning,size_tiny/size_small/size_medium/size_large(size_tinyis additive — a sub-2B model carriessize_smalltoo — and the small and large ends also alias tofast/large), andvendor_<v>/family_<f>. These resolve strongest-first from a per-modelmodel_overridesentry inendpoints.json, a built-in catalog of common model ids, then inference from the id itself. - Your own
tag_rules[](above) add more tags the same way the built-ins do — a model's final tag set is the union of every matching rule.
For each turn, Auto classifies a task (code, reasoning,
simple_fast, creative, or general), then narrows the
candidates: models that can't fit the estimated input are dropped (unless that would empty
the list), and an image attachment keeps only vision-tagged models when one
exists. It keeps loaded models first (see Auto), then
matches capability to difficulty: every model has a capability band —
tiny, small, medium, large or frontier — from a capability override, its
parameter count (tiny < 2B ≤ small < 10B ≤ medium < 40B ≤ large), the built-in catalog
for hosted models, or a guess (large for a cloud API, medium for your own server), one step
higher for a true reasoning model (R1, QwQ, gpt-oss — not one that can merely switch
thinking on) and for a coder on a coding turn. A simple turn needs small+,
a moderate one medium+, a hard one large+ (Jev's five-level scale adds tiny and frontier);
Auto keeps the models that reach the band and prefers the smallest of them — the fastest
model that is good enough — or, if none does, the strongest. It then applies soft boosts
(large_context when the turn needs it, and tool-calling/thinking-capable models
per your toggles) before walking the task's policies[].prefer list top to
bottom and picking the first selector that still matches at least one candidate. If nothing
in prefer ever matches, the default endpoint's first remaining candidate wins.
A user policy for a task replaces the built-in policy for that task
outright rather than merging with it, so writing your own prefer list is the
reliable way to pin a task to a specific model — for example, to keep creative
turns on one backend model even though a built-in glob like *large* or
*opus* would otherwise tag it (and win) differently:
{
"schema_version": 1,
"tag_rules": [
{ "pattern": "*my-writer-model*", "tags": ["storyteller"] }
],
"policies": [
{
"task": "creative",
"prefer": [
{ "tags": ["storyteller"] },
{ "tags": ["general"] }
]
}
]
}
This tags one specific backend model storyteller and makes it the first
choice for creative turns, falling back to any available model
(general) if that one is offline. See docs/router.example.json for
a fuller starter covering the classifier, routes, and multiple policies
together, and docs/endpoints.example.jsonc for the model_overrides
syntax used to set capability tags explicitly per backend.
Semantic router classifier
By default Auto classifies each turn with fast, deterministic heuristics only. To classify
turns semantically instead, either configure the auxiliary model
(CHAT_AUX_*) — which automatically doubles as the classifier via the JSON-schema
adapter with no extra setup (disable with CHAT_AUX_ROUTER=0) — or point the router
at a dedicated classifier: set CHAT_ROUTER_CLASSIFIER_URL and
CHAT_ROUTER_CLASSIFIER_MODEL (and optionally
CHAT_ROUTER_CLASSIFIER_API_KEY / CHAT_ROUTER_CLASSIFIER_TIMEOUT),
or add an equivalent classifier section to router.json (see the
commented example in docs/router.example.json) — env vars override the file
per field. A dedicated classifier always takes precedence over the auxiliary model.
The default classifier model is Arch-Router-1.5B, a purpose-built router
model. The core auto-detects it and speaks its native protocol: it is handed the task
classes (with descriptions overridable via routes above) and returns the single
best-fitting one, which maps to a model through the usual tag rules and policies.
classifier.format / CHAT_ROUTER_CLASSIFIER_FORMAT forces the
adapter — set schema to use a general instruct model with the richer
JSON-schema classification instead. Because Arch-Router returns only a route, under it (or
with the classifier disabled) the Auto-profile persona is chosen by a
deterministic task-affinity heuristic over the profile catalog rather than by the classifier;
the schema classifier gives a richer, content-aware persona pick.
On every routed turn the classifier races the heuristics against a deadline (default
3000ms, sized for a CPU-only sidecar classifying a fresh, full-context turn).
Whichever finishes first within the deadline wins; a classifier error,
timeout, or unset config falls back to heuristics, so a turn never fails or blocks on it.
The persisted routing reason records which one won, prefixed semantic: or
heuristic: , and the per-message stats hover's decision time includes any time
spent waiting on the classifier.
To self-host a tiny classifier model with no extra setup, use the
compose.router.yaml (or compose.router-only.yaml for host-run
cores) Compose overlay described under
Docker Compose deployment; it runs
ghcr.io/ggml-org/llama.cpp with the Arch-Router-1.5B GGUF and constrained
JSON output.
Arch-Router-1.5B is a Katanemo model published by DigitalOcean, LLC under the Katanemo Community License, not Omni Chat's own Apache-2.0 license. Built with DigitalOcean. Commercial use of Arch-Router-1.5B requires a separate license from DigitalOcean — see the model's license terms on Hugging Face before using it commercially. Katanemo Models are licensed under the DigitalOcean Community License, Copyright 2026 DigitalOcean, LLC. All Rights Reserved.
Chapter 12
Administration & Deployment
Multi-user isolation
Every owned resource carries an owner, and the store layer filters all queries by the session's user. Identity comes from the session, never the client. Accessing another user's private resource returns 404 — the app reveals nothing about resources you don't own.
Projects have a visibility of private or shared. Marking a project
shared makes its chats, documents, and memory readable and writable by members
of the group; private projects stay yours alone.
Deployment modes
Shared web instance
Run with CHAT_BIND=0.0.0.0 so household or team members can reach it in a
browser and log in. Each gets an isolated workspace; shared projects are the common ground.
Single-machine desktop
The Tauri shell bundles the core as a sidecar on 127.0.0.1 and opens it in
a native window. Auto-login for solo use is not built yet, so the first launch shows the
setup screen and later launches ask you to log in.
Building the desktop app (macOS)
The shell picks a free loopback port, starts the same core binary as a sidecar with that
port, waits for the core's READY port=<n> handshake, then points the webview
at it. All logic stays in the core — the Rust layer only supervises the process.
The app is built for Apple Silicon only and declares a minimum system version of 26.0, so
the DMG will not install or launch below macOS 26. make build-desktop runs
make check-macos-prereqs first, which fails on a non-arm64 host or any
DESKTOP_TARGET other than aarch64-apple-darwin.
On top of the usual toolchain you need Rust 1.82+ and the Tauri CLI:
cargo install tauri-cli --version "^2" --locked
make dev-desktop # run it
make build-desktop # bundle it
For a locally signed build, create an Apple Development certificate in Xcode,
find its exact name with security find-identity -v -p codesigning, then run:
APPLE_SIGNING_IDENTITY='Apple Development: Your Name (IDENTIFIER)' make build-desktop
The macOS signing config supplies the executable-memory entitlement required by
the core's SQLite/WebAssembly runtime, while keeping Hardened Runtime enabled.
A signed build missing this entitlement can stop with signal 9 before printing
anything (CODESIGNING: Invalid Page in the macOS crash report).
Rebuild with desktop/Entitlements.macos.plist configured; restoring
endpoint settings does not fix this packaging error.
The bundle lands at
desktop/target/release/bundle/macos/Omni Chat.app.
make build-desktop also builds the user-facing
omni-chat CLI into Contents/Resources/omni-chat
(not a Tauri sidecar — those get a target-triple suffix). The one-line
installer then symlinks that binary onto PATH.
Both targets first run make llama-sidecar, which downloads the
llama.cpp release pinned in desktop/llama.json, verifies its SHA256 and
keeps llama-server plus the libraries it actually links in
desktop/vendor/llama/ (git-ignored, ~23 MB). It needs the network once, then
no-ops until the pinned tag changes — the same discipline web/runtimes.json uses for
the Pyodide assets. To bump the version, edit the tag, run the target, and paste the checksum it
reports.
| What | Path |
|---|---|
| Config | ~/.config/omni-chat/ — the same directory the core uses in web mode, so one machine has one config. |
| Database | ~/Library/Application Support/omni-chat/app.db — separate from the repo's data/app.db, so the desktop app has its own accounts and chats. |
Without a signing identity, builds use ad-hoc signing. An Apple Development
certificate supports local development; Developer ID signing and notarization
are needed for normal distribution outside the App Store. The build accepts
Tauri's APPLE_SIGNING_IDENTITY setting; no identity is stored in the repo.
If a macOS build finishes compilation but fails at bundle_dmg.sh,
the verbose build output identifies the failing disk-image command. Retry packaging
with cargo tauri bundle --bundles app,dmg --verbose --config from
desktop/, supplying the same version configuration and signing identity
as the original build (see the README example). For Finder/AppleScript errors only,
CI=true skips window layout while retaining the Applications shortcut.
Linux AppImage
Linux desktop builds use the same thin Tauri supervisor and Go core as macOS. Prebuilt
AppImages are published for x86_64 and aarch64. The one file carries
both the GUI and the omni-chat launch CLI; the installer gives that file two names
using AppImage's multi-call entrypoint.
On first use, the shell materializes its embedded Go core and CLI as private, content-addressed executables under the user's XDG cache. They still come from the one downloaded and verified AppImage; no system-wide runtime is installed.
curl -fsSL https://geekaholic.gitlab.io/omni-chat/install-linux.sh | bash
omni-chat-app
omni-chat launch pi
Immutable, versioned downloads are available from the project's
GitLab package registry.
Make the matching Omni-Chat-x86_64.AppImage or
Omni-Chat-aarch64.AppImage executable and run it. Endpoint keys use the desktop
Secret Service (GNOME Keyring, KWallet or compatible) and are never written to SQLite or
plaintext config.
Release builds use Ubuntu 24.04 target-architecture tools, once per architecture with the same
explicit BUILD_VERSION. Published compatibility starts at Ubuntu 24.04 and is
smoke-tested on Debian 13, current Fedora, and a current rolling distribution. Linux
deliberately does not bundle
llama-server; the core uses an installed binary or Docker so the runtime can match
the host's GPU drivers.
make check-linux-prereqs
make build-linux BUILD_VERSION=2026.9.1120000
GITLAB_DEPLOY_TOKEN=your-deploy-token make publish-linux \
APPIMAGE_X86_64='/path/Omni Chat_2026.9.1120000_amd64.AppImage' \
APPIMAGE_AARCH64='/path/Omni Chat_2026.9.1120000_aarch64.AppImage'
Create a project deploy token with the write_package_registry scope and provide
it as GITLAB_DEPLOY_TOKEN only for the publish command. A personal, project or
group access token with api also works as GITLAB_TOKEN; GitLab CI uses
its automatic CI_JOB_TOKEN. AppImages and their checksums go to GitLab's generic
package registry; GitLab Pages stores only the small downloads/latest.json
manifest used by the installer.
On other Linux distributions, including Omarchy, the bundled Ubuntu 24.04 Docker builder supplies Go, Node, Rust, Tauri and the native libraries. Only Make and a working local Docker daemon accessible by your user are required:
make build-linux-docker
# Optional: pin the release version.
make build-linux-docker BUILD_VERSION=2026.9.1120000
The target checks Docker access, builds/reuses docker/build-linux/Dockerfile,
and runs the build stages as your UID/GID without a privileged container or FUSE device.
The first run downloads and compiles the tools; later runs reuse layers and caches.
AppImages are written to
.cache/linux-docker/amd64/target/release/bundle/appimage/ (use
arm64 for an ARM64 target). Rust and npm caches are separate from your host's
desktop/target and web/node_modules; Go sidecars and web assets
still use their normal checkout paths, so run only one build at a time in a checkout.
The default target is the host's native architecture. If Docker access fails, start the daemon
and check your Docker context and socket permissions before retrying.
To build x86_64 on ARM64 or ARM64 on x86_64, set LINUX_ARCH=amd64 or
LINUX_ARCH=arm64 (aliases x86_64/aarch64 also work).
The frontend builds in native Node/Go Docker userspace, which also cross-compiles the
pure-Go sidecars for the selected architecture. Rust and linuxdeploy's ELF dependency step
run in target Ubuntu userspace under QEMU when needed. The native stage then creates the
target AppImage from the completed AppDir and target runtime, avoiding the target AppImage
plugin on large-page ARM64 hosts. npm caches follow the host architecture in
.cache/linux-docker/assets-<host-arch>/; Rust caches and AppImages follow the
target architecture. The wrapper checks target execution before compiling.
Ubuntu's versioned Rust 1.91 packages avoid the rustup compiler's allocator crash;
a Rust compile/link/run smoke test runs before installing Tauri.
No special QEMU or Node environment settings are required.
If QEMU/binfmt is missing, enable the host executable handlers once:
docker run --privileged --rm tonistiigi/binfmt --install amd64,arm64
This setup changes host executable handlers and is never run automatically by Make. See Docker's QEMU documentation. Emulated compilation and compression can be much slower. Build both targets sequentially with the same version, then publish the two files:
make build-linux-docker LINUX_ARCH=amd64 BUILD_VERSION=2026.9.7032027
make build-linux-docker LINUX_ARCH=arm64 BUILD_VERSION=2026.9.7032027
GITLAB_DEPLOY_TOKEN=your-deploy-token make publish-linux \
APPIMAGE_X86_64='.cache/linux-docker/amd64/target/release/bundle/appimage/Omni Chat_2026.9.7032027_amd64.AppImage' \
APPIMAGE_AARCH64='.cache/linux-docker/arm64/target/release/bundle/appimage/Omni Chat_2026.9.7032027_aarch64.AppImage'
Building and self-signing the iPad app
The native iPad target requires an M-series iPad with at least 8 GB RAM, iPadOS 26.0+, and full Xcode 26.6+ on a Mac. The stable reference test point is iPadOS 26.5. Install CMake, CocoaPods, the iOS Rust target, and Tauri CLI, then generate and build:
sudo xcode-select -s /Applications/Xcode.app/Contents/Developer
brew install go node make cmake rustup cocoapods
export PATH="$(brew --prefix rustup)/bin:$PATH"
rustup default stable
rustup target add --toolchain stable aarch64-apple-ios aarch64-apple-ios-sim
rustup component add --toolchain stable llvm-tools
cargo install tauri-cli --version "^2" --locked
cmake --version
pod --version
make check-ipad-prereqs
make build-ipad
The Tauri Xcode project under desktop/gen/apple is tracked and ready after
checkout so Xcode Cloud can select it. Run make init-ios only when Tauri asks you
to regenerate the mobile scaffolding, then review and commit the generated changes.
Xcode Cloud's post-clone hook installs llvm-tools and rust-src
before compilation. The iOS wrapper selects the real stable compiler and propagates that
toolchain to child commands, preventing competing rustup downloads during parallel builds.
The iPad build explicitly enables Go modules, even if your global Go configuration
has GO111MODULE=off.
The iPad app uses the desktop icon artwork from desktop/app-icon.png.
Project initialization and each iPad build regenerate the required icon sizes. To change
the icon, replace that artwork, rebuild, and install over the existing app to keep its data.
For a cable-free test, install the iOS 26.5 Simulator runtime under
Xcode → Settings → Components, then run
make check-ipad-simulator-prereqs and
make dev-ipad-simulator. Tauri prompts for a destination; an exact one can be
supplied as IPAD_SIMULATOR="iPad Air 11-inch (M4)". No Personal Team or
Developer Mode is needed. The Apple-Silicon simulator exercises the embedded Go core and
Metal-linked library, but its model speed and memory pressure are not physical-iPad results.
Xcode 26.6 supports iPadOS devices only through 26.5. An iPad running iPadOS 27 beta needs a current Xcode 27 beta supported by the Mac; Developer Mode does not add missing device support. Check Apple's Xcode SDK and device-support table, especially when the device beta is newer than Xcode.
Install CocoaPods with Homebrew. If Tauri reports that the package is missing and falls
back to a system-Ruby installation requiring sudo, stop and run
brew install cocoapods; do not use the Ruby-gem fallback.
The llvm-tools component supports the pinned upstream
swift-rs Xcode 27 compatibility fix. The generated Xcode build phase uses the
repository's Rust wrapper, which locates Homebrew rustup and the Cargo-installed
Tauri CLI even though GUI applications start with a minimal PATH.
iOS 27 also turns UIKit's missing-scene-lifecycle warning into a launch-time breakpoint at
UIApplicationMain. Omni Chat supplies a tracked scene manifest naming tao's
delegate, mirrors it into generated Xcode files, and pins tao's merged scene-configuration
lifetime fix. After updating the source, rerun make build-ipad before pressing Run
in Xcode; continuing past the breakpoint does not fix the lifecycle declaration.
The first llama.cpp compile can take several minutes. Missing Git-repository and OpenMP
messages are expected warnings for the pinned source archive: iPad inference uses Metal and
Accelerate. The iOS build excludes llama.cpp's standalone executable and web UI and links only
its static API server. Keep the make build-ipad terminal open after Xcode launches;
a missing-signing-certificate warning is expected until a Personal Team is selected.
In Xcode → Settings → Accounts, add your Apple ID. In the generated Omni
Chat target's Signing & Capabilities pane, enable automatic signing and select
the resulting Personal Team. Keep bundle id
org.geekaholic.omni-chat, select the connected iPad, and press Run. Trust the Mac
when prompted. If requested, enable Developer Mode under iPad
Settings → Privacy & Security, restart, and Run again. After the first
install, approve the Personal Team certificate under Settings → General → VPN &
Device Management → Developer App → Trust, then Run or open Omni Chat again. This
certificate approval is separate from trusting the Mac and enabling Developer Mode.
A free Personal Team profile normally expires after seven days; reconnect and Run from
Xcode to reprovision. This does not produce an App Store/TestFlight or redistributable build.
Re-signing over the same installed bundle preserves data; deleting the app removes its local
database and downloaded model. The full procedure and troubleshooting are in
docs/ipad-self-signing.md.
TestFlight builds come from Xcode Cloud, not from a local archive. Each release begins by
bumping CFBundleShortVersionString and CFBundleVersion across the
five tracked files that carry the app version, then pushing the branch the workflow watches;
a repeated build number is rejected at upload. The procedure is in
docs/ipad-testflight.md. desktop/Info.ios.plist declares
ITSAppUsesNonExemptEncryption = NO, so builds skip the export compliance
prompt.
First launch offers a resumable ~1.75 GB download of the pinned Qwen 3.5 2B MLX snapshot. The Go core checks every file's pinned length and SHA-256 before the MLX helper can load it. It supports private offline text chat and tools; vision is not included. Leave the app in the foreground for downloads and inference because iPadOS may suspend the complete process. Docker services, host coding agents, stdio MCP servers, and SFTP are unavailable in the iPad sandbox; HTTP MCP servers, remote OpenAI-compatible endpoints, and other network features remain optional.
Editing endpoints from inside the app: AI Servers and Models
Omni Chat edits endpoints.json itself — in the desktop app, and for admins in the
web app too — so getting running does not mean copying
docs/endpoints.example.jsonc and hand-editing JSON. Open the user menu →
Settings; every user can open it. Admins (the owner, plus
anyone given the Admin switch in Members) edit everything; everyone else sees
Integrations and a read-only view of the servers, model assignments and web
tools, never their URLs or keys. On the web, saving restarts the server in place and a typed
API key is stored in the server's endpoints.json (there is no Keychain); under
Docker the ./config mount must be writable by the container user, uid 10001
(the shipped compose.yaml mounts it read-write; on Linux run
sudo chown 10001:"$(id -g)" ./config && sudo chmod 775 ./config once;
add :ro to lock web saving off).
Settings → Integrations connects ChatGPT (sign in, then Use for Chat /
Image generation; its model list follows your Codex CLI when installed, and so does each
model's context window, so long chats keep their history) and your MCP servers.
Local services, Agent handoff and Desktop stay desktop-app-only. Settings splits into AI Servers
("Discover and manage servers") and Models ("Defaults for each job"), followed
by Web & tools, Local services, Agent
handoff, and Desktop. It opens on AI Servers when you have no saved
servers yet, else on Models. On a phone-width screen (820 px or narrower, e.g. an iPhone
in portrait) the page list is a slide-in menu: Settings opens with it showing, and the ☰
button beside the page title brings it back. Server rows there fold their status into a dot
on the server-type tile plus a word on the address line.
AI Servers holds the servers themselves. Connections lists what you've saved — each row shows a status (Connected / Needs key / Unreachable / Unsaved), the label, URL, its "Used for" chips, and a model count, with Manage expanding the full editor in place — plus an Add server button for typing one in by hand. Below that (above it, when you have no servers yet) is Network discovery: an auto toggle (servers on this machine are always used, whatever the toggle says) and a Scan now button for an on-demand full pass — this machine, named local hosts, and the attached network — regardless of the toggle. Scan now lists every open server it finds with its model count, and a key-locked one as Locked · needs API key. Each server shows a lettermark tile for its detected server type (llama.cpp, vLLM, Ollama, LM Studio, oMLX, LiteLLM, SGLang, OpenAI; a generic server glyph when unknown) next to its model count. A status dot on each result says where it stands: green In use (auto-discovery already offers its models in the picker) or Saved, orange Needs key, grey Available; each result gets its own Add connection button, which opens the editor with the URL and a suggested label already filled in — and, for a locked one, the key field focused so you can type it straight in.
The connection editor itself (Add server / Manage / add-from-scan) has a Label
(an internal id is derived once behind the scenes and never shown again — rename the label as
often as you like), Base URL (the server address only — no /v1 at
the end), Server type (set by detection, or pick it yourself), Authentication (None / API key — ChatGPT sign-in lives in Integrations), and Used
for switches — Chat, Embeddings, Image generation, Read aloud, Voice input, at least one
required. Turning a use on reveals that job's settings inline (Chat: context window, max output,
sampling, thinking, model overrides; Image: model, carrier model, stream, timeout; Read aloud:
model, voices, timeout; Voice input: model, language, timeout). Test connection
runs one check per enabled use and shows a row per use, then Save & restart
applies it. Saving works even when a test fails or was never run, with a warning.
Models holds the default for each job — each select lists only connections with that use turned on on AI Servers. Default chat server picks a server only (no model; the model you used last still wins on a new chat). Embeddings is a connection + model + dimension, with a warning that changing the model re-indexes memory — it's optional; memory works without it. Image generation, Read aloud (plus a default voice), and Voice input each pick a default connection the same way. A job with no tagged connection shows "No servers are set up for this — add one in AI Servers" instead of an empty select. Removing on AI Servers a connection that a Models default still points at is blocked at Save, and the problem message names — and clicking it opens — whichever page fixes it.
Auxiliary model
The Models page's Auxiliary card offers On-device (AppleFoundation, native
Apple shells only), Managed (the Local services llama.cpp below), a connection
+ model id, or Disabled. This small, non-thinking model handles chat titles, background memory
jobs and long-chat summaries without occupying the chat model. The
Use for Smart Router switch is on by default; turn it off when routing should
remain heuristic or use a separately configured classifier. Show in model picker
is also on by default. When enabled, the model can be selected for an ordinary chat turn but
remains excluded from Auto routing and omni-chat launch. Turn it off to hide new
selection without interrupting auxiliary jobs or chats that already use it. Web operators can
set aux.show_in_model_dropdown: false in endpoints.json. The app reads
that model's live endpoint metadata for prompt budgeting; this includes apfel's reported context
limit for apple-foundationmodel. If Apple's framework returns its opaque retryable
generation error before emitting any text or tool call, the turn is retried once safely. Apfel
can also fall back to writing a tool_calls JSON block as ordinary text; the app
recognizes a complete block only when every named tool was actually offered, then executes it
through the normal validation and budget loop instead of showing raw JSON in the conversation.
Test sends one real completion using the values currently in the form, including a connection and key that have not been saved yet. A cold local model can take several seconds to load. The test never starts a service or downloads weights.
Image generation, read aloud and dictation
These three jobs are chips you tick on a connection (AI Servers), with defaults you pick on the Models page — so a local backend and a hosted one can sit side by side, and a server serving several jobs (chat, images, TTS) is typed and keyed only once. They live in the same file as the chat connections and are saved by the same Save & restart button.
Image generation's chip settings take a model, and — via the connection's
Authentication and Advanced → protocol — either an
OpenAI-compatible backend or a Codex subscription through the Responses API. The two want
different values: a local image server is something like z-image-turbo on a port
on this machine, while Codex serves gpt-image-2 from chatgpt.com with
protocol: "responses". The Codex case also asks for a carrier model — the
text model that carries the image request.
A sign-in or Responses API server (Codex) can only be used for Chat and Image generation: embeddings, read aloud, voice input and the auxiliary model send just a URL and key, so the editor disables those switches, the Models page leaves such servers out of the auxiliary and embeddings choices, and the core refuses the file if one is named there anyway.
A connection picks its credential explicitly on Authentication: an API key,
or one of the sign-in providers the core offers. Choosing Codex preselects Sign in with
ChatGPT, because that is how a subscription is authenticated — leaving it on the key
field is what used to produce connections with no credential at all, which then failed with a
bare 401. You sign in from the editor itself, so image generation no longer
requires creating a separate chat connection first — tick both Chat and Image generation on
one Codex connection instead.
A provider sign-in is stored once per connection, and every job section that references that connection rides it. If a Codex connection is ticked for both Chat and Image generation, one sign-in covers both — and signing out drops it for every job using that connection.
Test connection's Image row checks credential and reach without generating, listing the backend's models when it has a listing and falling back to a health check otherwise (which is what the bundled image sidecar does).
The smallest image the API offers is 1024×1024 — there is no thumbnail mode — which is real money against a hosted backend and a queued GPU slot on a local one. A settings button that quietly spent one on every press would be a trap, so this one checks the connection and says so.
Read aloud's chip works with any OpenAI-compatible speech server, including the bundled voice sidecar. Test connection's Read aloud row synthesizes a short phrase and plays it, which is the point: no status line can tell you the voice is the one you wanted. It also reports the voices the backend lists, with one click to adopt them into the Voices field — leave that field empty and the backend is asked at runtime instead.
Voice input's chip configures dictation. Test connection's Voice input row sends a half-second clip generated on the spot, containing silence. It passes when the backend accepts it, which proves the route exists, the key is accepted and the model id resolves — but not that anything was heard, because there is nothing to hear. Whisper given silence returns either nothing or a small hallucination, so a test that demanded a transcript would fail a backend that works perfectly.
connections — like endpoints, search, and fetch — always moves
on save, because every client that speaks schema 2 renders it; the other job sections
(images/tts/stt/aux/embeddings)
move only when the client actually sent them, so an older client leaves a section it doesn't
know about alone. Removing your last image connection really does remove it. The Auxiliary
card can replace or disable the background-model section written by Local
services, while preserving its advanced dedicated-router model field.
Local services
The Settings → Local services page can start two optional sidecars for you, so a fresh install is not stuck with search and background intelligence switched off. Starting one also writes the config that points the core at it, so it asks for a restart afterwards.
| Service | Needs | Unlocks |
|---|---|---|
| Web search (SearXNG) | Docker | web_search, search_images |
| Background models (llama.cpp) | Nothing on macOS — llama-server is bundled | Chat titles, memory extraction, rolling summaries, semantic routing |
For the background models you pick what to host. Background model runs one
~700 MB model — not a reduced setup, since that model also serves as the router's semantic
classifier by default, so titles, memory and routing all work.
Background + router adds Arch-Router-1.5B as a purpose-built
classifier; both are hosted by one llama-server in router mode, on one
port, loaded on demand. Models swap by default, which is safe on any machine;
Keep both models loaded holds them resident and wants roughly 32 GB.
Test background models runs one real request per configured model and
reports each as ready in green or not ready in red, with a timing. It is the only honest
readiness signal, because Running does not mean loaded: router mode loads a
model only when a request names it, and the service probe uses /health, which
answers 200 with nothing resident. The first test is therefore slow — it includes reading the
weights off disk, a few seconds on a cold machine — while later tests return in milliseconds.
The test never downloads anything; that is what Start does. With a dedicated router classifier
configured there is a row per model, so a failure says which of the two is wrong.
The core restarts on every config save, so a loaded model deliberately survives those
restarts rather than reloading hundreds of megabytes each time you edit a setting. Quitting
the app is different: everything the app started — the model server and its
omni-managed-* containers — is stopped with it, so nothing keeps eating memory
after the app is gone. A service you started yourself by hand still shows here as running
(state is discovered by probing), but quitting leaves it alone: only what the app launched is
stopped. If the app is force-killed, the leftovers are picked up on the next launch and
stopped at that session's quit.
On macOS, an off-by-default Keep running in background setting (under
Settings → Desktop) changes what closing the window means: the
window hides, the app stays visibly in the Dock, and the services keep running — useful when a
model server should stay warm. A Dock click brings the window back; quitting from the Dock or
with ⌘Q still stops everything. Note this also applies to a dev run:
Ctrl-C on make dev-desktop now stops the services it started too.
The macOS app ships its own llama-server (23 MB, in
Contents/Resources/llama/), so background models work on a Mac with neither Docker
nor Homebrew installed. A native binary is always preferred over a container because containers
on macOS never see the GPU, and the bundled build is Metal-enabled. Resolution order is the
bundled binary, then llama-server on PATH, then Docker — so a Linux
build with no bundled binary still works through a container.
Managed services can create an optional-service-only endpoints.json; no remote
chat endpoint is required first. This is the same valid shape used on iPad, where
@on-device is injected at runtime instead of being written to the file.
Image generation, TTS and STT are deliberately not managed: a container cannot reach the GPU on macOS, so a one-click container there would be unusably slow while looking supported. Those keep their native and remote-host paths described below.
Agent handoff
The Agent handoff page sets up Pi, Hola, or both — letting a chat pass a
coding task to an enabled agent, which reads and edits files on this machine. It writes
<config-dir>/agents.json, a different file from the endpoints config. The owner
console stages both documents and applies them with one Save & restart action.
Pick where the agent works first. The default gives every chat its own fresh subfolder under one workspace folder you choose — created on the first handoff, reused by follow-ups in the same chat, and never your existing files or the folder itself. The other mode lets the agent work directly inside folders you list, for real repositories; there the model proposes an existing directory under one of them and you approve the exact path.
Turn Pi and Hola on independently, then pick where each one runs. Installed on this
machine runs pi or hola-coder directly — install it from
pi.dev or the
Hola usage guide and the form finds
it. Docker container runs each delegation in a throwaway container instead.
The repository supplies only the Pi image, built with make agent-image; for Hola,
give a custom image containing Hola 0.6+ and hola-coder. The form discovers and
tests each enabled agent separately.
Hola profile values override provider environment variables, so Omni Chat generates an
omni-chat profile for each delegation. It pins the chat model and scoped proxy,
enables the coding toolsets and OpenAI tool-call parser, and applies the documented balanced
convergence policy. The profile omits api_key, so Hola reads the short-lived
token from OPENAI_API_KEY in the child process environment; the token never enters
the profile or command line. Omni Chat launches each handoff with --new, so files
persist in the workspace but Hola never replays stale model history from
hola-coder.db. Personal ~/.hola settings are
excluded. A repository-level .hola/profiles.json would take precedence, so that
workspace is refused with an explanation rather than silently routing somewhere else. See
Hola's profile guide and
0.6 options.
bridge network on macOS
Docker Desktop's host networking does not give a container this machine's loopback, so an
agent on the host network cannot reach the model and the run ends without
output. On bridge the core rewrites localhost endpoint URLs to
host.docker.internal, which works. The form defaults to bridge.
Folders the agent may work in is the safety boundary, and nothing is filled in for you: the agent can only ever read and write inside a folder that resolves under one of these, so it is worth choosing deliberately rather than accepting a guess.
Each Test button reports a row per check — the command resolved to a real path,
the Docker daemon answered, the image exists, the folders resolve, and finally the agent itself
runs and reports its version. That last row is the point: a binary being present says nothing
about whether it works, and a half-installed npm package, a broken node or an
image built for the wrong architecture all pass a file-exists check and fail here. It
deliberately does not run a real delegation — that needs a model, a chat, your consent
and minutes of waiting, and it would write into your own repository.
Turn off deletes agents.json (the format has no off switch, and
an empty agent list is invalid) and keeps an agents.json.bak beside it, so setting
it up again is not starting over.
PATH
An app launched from Finder or the Dock gets launchd's minimal environment, which contains
neither /opt/homebrew/bin nor /usr/local/bin — so a plainly
installed Pi/Hola command, or Docker itself, can look missing. The core therefore searches the usual
install directories after PATH, never instead of it. If your copy lives
somewhere unusual (a version manager, for instance), give the full path in the Command
field.
Self-hosting SearXNG
The search section of endpoints.json names one provider at a time.
For the self-hosted option:
{ "search": { "provider": "searxng", "base_url": "http://localhost:8888", "max_results": 5 } }
Bring the sidecar up on its own — it is not part of the normal Compose stack:
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.search-only.yaml up -d
That publishes SearXNG on 127.0.0.1:8888. Point the app at it —
Settings → Web & tools → Web search → SearXNG in the desktop app, or
CHAT_WEB_SEARCH_URL=http://localhost:8888 for web mode — and press
Test search. Running the app and the sidecar in one Compose project instead?
Use -f compose.yaml -f compose.search.yaml, which sets the URL for you;
docs/endpoints.docker.example.jsonc shows that form.
deploy/searxng/settings.yml must list json under
search.formats, or SearXNG answers 403 and the core reports "SearXNG rejected
JSON output; enable the json search format." That is exactly what Test search
surfaces.
The same screen configures web search and web fetch, which
used to be environment-only and so unreachable in a bundled app. Search offers Tavily (hosted —
paste a free API key, nothing to run) or SearXNG (self-hosted — give it a URL); a
Test search button runs one real query, so a rejected key or a SearXNG without
JSON output enabled shows up immediately. SearXNG additionally provides image search; Tavily
does not. On iPad, Tavily can be saved without adding a remote chat endpoint and is available
to the injected @on-device model. Web fetch is a single switch: its default
provider runs inside the core behind an SSRF guard, so fetch_url and
fetch_image need nothing else running.
Saving does three things in order: writes any new API key to the OS secure store, writes the
config file, then restarts the core. The restart is the apply mechanism — the core reads its
configuration once at startup — and it ends any answer in progress. The previous version is kept
at endpoints.json.bak, and a config the core would refuse to boot on is rejected
before the file is touched.
endpoints.json stores only api_key_env — the name of an
environment variable. The shell reads the value from Apple Keychain or Linux Secret Service and injects it when it
spawns the core, so the secret never reaches disk. Two consequences: the first launch after
storing a key raises a macOS permission prompt (an ad-hoc build gets a new signature on every
rebuild, so during development this recurs — the launch is not blocked, the shell waits ten
seconds then starts without the stored keys), and a core started outside the shell (web mode,
Docker) reading the same file will not have those variables set. Keys already inline in the
file keep working everywhere; the UI shows them as stored in plain text and offers to move
them.
This is native-shell-only by design. GET/PUT /api/config return
501 not_implemented unless the core was started with CHAT_DESKTOP=1,
which only the Linux, macOS and iPad Tauri shells set — a shared web instance never exposes instance
config, endpoint keys, or the ability to repoint the app at another backend.
If the core fails to start — most often an invalid file in ~/.config/omni-chat/
— the window stays on the shell's own splash page and shows the core's output instead of a
blank screen, with Try again and Restore previous config
buttons.
Three desktop differences
Confirmations (Archive, Delete chat, New project, discarding a temporary chat) are drawn by
the app rather than by the operating system. The webview does not implement the browser's
alert/confirm/prompt, so a native dialog would never
appear — the same in-app dialog is now used in the browser too, so both behave identically.
Saving a file — exporting a chat, downloading a code block or an image — opens the macOS save panel, because the webview ignores the browser's download mechanism. Print / Save as PDF is unavailable in the desktop app for the same reason; use Download HTML and print that from a browser.
Links open in your default browser. The webview cannot open a tab, so
citations, source cards and links in a reply are handed to the operating system instead — and so
is OAuth sign-in, which is why signing in to a provider was impossible in the app before. Only
http and https are ever opened, and only to another site: the app's own
links, such as an attachment, stay inside, because a separate browser has no session for the
local core.
Production build & run
Build the web SPA and the core binary (the SPA is embedded into the binary):
make build
# binary: .cache/bin/omni-chat-core
Run it with a persistent database path and your endpoint:
CHAT_ENDPOINT_URL=http://your-endpoint:port \
CHAT_BIND=0.0.0.0 \
CHAT_PORT=8080 \
CHAT_DB_PATH=/persistent/path/app.db \
./.cache/bin/omni-chat-core
CHAT_BIND=0.0.0.0 exposes the instance to the LAN. Only set it when you
intend to share, and put it behind appropriate network controls.
Docker Compose deployment
To include the optional self-hosted SearXNG provider, use:
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.yaml -f compose.search.yaml up -d
The overlay enables JSON output, connects Omni Chat over the private Compose network, and
exposes SearXNG only at 127.0.0.1:8888 by default. The generated
SEARXNG_SECRET is required, belongs in the ignored .env file, and
must not be committed or copied into settings.yml.
For a host-run core (make dev) instead of the full Compose stack, start just
SearXNG and point the core at its published port:
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.search-only.yaml up -d
CHAT_WEB_SEARCH_URL=http://localhost:8888 make dev
To include the optional llama.cpp auxiliary-model sidecar (see Auxiliary model above), use:
docker compose -f compose.yaml -f compose.aux.yaml up -d
The overlay wires CHAT_AUX_URL / CHAT_AUX_MODEL to an internal
aux-model service running ghcr.io/ggml-org/llama.cpp, which downloads
the configured Hugging Face GGUF (default LiquidAI/LFM2.5-1.2B-Instruct-GGUF,
~731 MiB) into a named volume on first run. This one model handles chat titles, memory
jobs, and smart-router classification, so on its own it needs no separate router sidecar. It is
stackable with the other overlays, e.g. docker compose -f compose.yaml -f
compose.search.yaml -f compose.aux.yaml up -d. For a host-run core, start just the
sidecar and point the core at its published port:
docker compose -f compose.aux-only.yaml up -d
CHAT_AUX_URL=http://localhost:8091 \
CHAT_AUX_MODEL=lfm2.5-1.2b-instruct make dev
To also run the optional dedicated llama.cpp router-classifier sidecar (see Semantic router classifier below), use:
docker compose -f compose.yaml -f compose.aux.yaml -f compose.router.yaml up -d
The overlay wires CHAT_ROUTER_CLASSIFIER_URL /
CHAT_ROUTER_CLASSIFIER_MODEL to an internal router-classifier
service running ghcr.io/ggml-org/llama.cpp, which downloads the configured
Hugging Face GGUF (default katanemo/Arch-Router-1.5B.gguf, ~1 GiB) into a named
volume on first run — allow a few minutes before the sidecar reports healthy and Omni Chat
starts. When present, this dedicated classifier outranks the auxiliary model for routing (the
aux model then does titles and memory only); omit it and the aux sidecar classifies via the
fallback. For a host-run core, start just the sidecar and point the core at its published
port:
docker compose -f compose.router-only.yaml up -d
CHAT_ROUTER_CLASSIFIER_URL=http://localhost:8090 \
CHAT_ROUTER_CLASSIFIER_MODEL=arch-router-1.5b make dev
To run the fetch_url tool through a self-hosted reader sidecar instead of the
in-core direct provider, use:
docker compose -f compose.yaml -f compose.fetch.yaml up -d
The overlay enables the tool (CHAT_WEB_FETCH=true), switches it to the
reader provider, and points Omni Chat at an internal reader service on
http://reader:8081 (also published on 127.0.0.1:3001 for debugging).
The service is ghcr.io/jina-ai/reader:oss, Jina AI's official self-host image
(multi-platform: linux/amd64 and linux/arm64, so it runs unmodified on
Apple Silicon) that renders JavaScript in a bundled headless browser; override
READER_IMAGE/READER_VERSION to pin or replace it. It is stackable with
the other overlays. When to use it: the default direct provider
needs no extra container and handles most static or server-rendered pages, but it only does a
plain HTTP GET — it cannot run JavaScript, is blocked by anti-bot defenses such as Cloudflare,
and cannot read YouTube transcripts. The reader sidecar runs a real browser, so it handles all
three, at the cost of a much larger image, more memory, slower fetches (consider raising
CHAT_WEB_FETCH_TIMEOUT, up to 60s), and sending target URLs to that
container. Prefer direct and switch to the reader overlay only when you need one of
those capabilities. For a host-run core, start just the reader and point the core at its
published port:
docker compose -f compose.fetch-only.yaml up -d
CHAT_WEB_FETCH=true CHAT_WEB_FETCH_PROVIDER=reader \
CHAT_WEB_FETCH_URL=http://localhost:3001 make dev
To offer the generate_image tool, use the image-generation overlay (see
Running Omni Chat, combination g, for the requirements —
an NVIDIA GPU with the container toolkit, and a ~12 GB first-run weight download into the
imagegen-models volume):
docker compose -f compose.yaml -f compose.imagegen.yaml up --build -d
The overlay builds the bundled Z-Image-Turbo sidecar (docker/imagegen/ — FastAPI
+ diffusers exposing POST /v1/images/generations) and sets
CHAT_IMAGE_GEN_URL=http://imagegen:8000 on the app. Tune it with
IMAGEGEN_MODEL_ID (any diffusers text-to-image model id; default
Tongyi-MAI/Z-Image-Turbo), IMAGEGEN_STEPS, and — for precise
instruction edits via the edit_image tool —
IMAGEGEN_EDIT_MODEL_ID (recommended: Qwen/Qwen-Image-Edit-2509,
Apache-2.0 but ~20B; alternative: black-forest-labs/FLUX.1-Kontext-dev, 12B,
gated + non-commercial — accept its license on Hugging Face and set HF_TOKEN in
.env; unset falls back to img2img with the base model, tuned by
IMAGEGEN_EDIT_STRENGTH, default 0.7 — raise toward 0.8 for stronger restyles —
and IMAGEGEN_EDIT_GUIDANCE, default 1.0, which only helps CFG-capable models as
the distilled base can burn above 1) in .env. Tongyi's announced Z-Image-Edit
has no public weights as of July 2026. On non-x86 GPU hosts (Jetson/L4T or arm64 CUDA),
override IMAGEGEN_BASE_IMAGE with a matching CUDA-enabled PyTorch base and set
IMAGEGEN_RUNTIME=nvidia — see the GPU note in
Getting started. For a host-run core — or to run the sidecar on
a separate GPU machine — start just the sidecar and point the core at its published port
(127.0.0.1:8001 by default; set IMAGEGEN_PUBLISH_ADDRESS=0.0.0.0 and an
IMAGEGEN_API_KEY/CHAT_IMAGE_GEN_API_KEY pair when it must be reachable
over the network):
docker compose -f compose.imagegen-only.yaml up --build -d
CHAT_IMAGE_GEN_URL=http://localhost:8001 make dev
Any other OpenAI-Images-compatible server also works — set CHAT_IMAGE_GEN_URL
(plus CHAT_IMAGE_GEN_MODEL/CHAT_IMAGE_GEN_API_KEY as needed) and skip
the sidecar.
To use the default in-core direct provider under the base Compose stack (no
sidecar), just add CHAT_WEB_FETCH=true to .env — no overlay needed. If
you are testing a feature (such as fetch_url) that predates the published
latest image, docker compose up -d alone keeps running that stale image;
rebuild from your checked-out source with docker compose up --build -d to pick up
local changes.
The repository's Compose configuration pulls the non-root, multi-platform image from the
GitLab Container Registry, publishes the app only on 127.0.0.1:8080 by default,
checks /api/health, and restarts the service unless it is stopped explicitly:
# Optional environment overrides; .env is ignored by Git.
cp .env.example .env
docker compose up -d
docker compose ps
Compose passes variables from .env into the container. Use it for
CHAT_ENDPOINT_URL, CHAT_API_KEY, variables referenced by
api_key_env, and other supported CHAT_* overrides. Set
OMNI_CHAT_PUBLISH_ADDRESS=0.0.0.0 only when the app should be reachable from
other machines, and protect the published port appropriately.
The default image is
registry.gitlab.com/geekaholic/omni-chat:latest for both AMD64 and ARM64. Every
successful main-branch pipeline also publishes an immutable full-commit-SHA tag. Set
OMNI_CHAT_IMAGE=registry.gitlab.com/geekaholic/omni-chat:<full-commit-sha>
in .env to pin a deployment. Private projects require
docker login registry.gitlab.com before pulling. Use
docker compose up --build -d to build from the local checkout instead.
Connecting a container to AI on the host
localhost inside a container is the container itself. Use
host.docker.internal for an OpenAI-compatible endpoint running on the Docker
host. Compose adds the host-gateway mapping needed on native Linux; Docker Desktop provides
the same hostname. On native Linux, the AI process must also listen on an interface reachable
from Docker's bridge rather than only 127.0.0.1. Bind the AI server to an
appropriate host interface and restrict that port with host firewall rules.
Docker data and configuration
SQLite is stored in the named omni-chat-data volume. Operator-managed JSON
files are bind-mounted read-write from ./config to /config, so
admins can save Settings and the files survive image upgrades. The process runs as uid
10001; on Linux a directory you created is not writable by that user until you hand it
over, even with no :ro on the mount. Settings tells the two causes apart:
"cannot write to the config directory" means ownership, fixed once with the commands
below; "the config directory is read-only" means the mount carries :ro.
Append :ro yourself to edit files only on the host.
sudo chown 10001:"$(id -g)" ./config
sudo chmod 775 ./config
Files that Settings saves (endpoints.json and its .bak) are
written mode 0600 and owned by uid 10001, because they can hold API keys. After
the first save from Settings, edit them on the host with sudo.
cp docs/endpoints.docker.example.jsonc config/endpoints.json # then delete the // comments
cp docs/profiles.example.json config/profiles.json
cp docs/SOUL.example.md config/SOUL.md
mkdir -p config/skills && cp docs/skills.example/*.md config/skills/
# Edit the files, then reload startup configuration.
docker compose restart
docs/endpoints.docker.example.jsonc is an annotated example (schema 2,
connections[]) the app never loads — copy from it, then delete its comments
before the result is valid JSON. It points to host.docker.internal:8000. If
config/endpoints.json exists, it takes precedence over
CHAT_ENDPOINT_URL. Existing JSON configuration is validated at startup exactly
as it is for native runs.
Update and replace the container without deleting its data:
git pull
docker compose pull
docker compose up -d
docker compose down preserves omni-chat-data.
docker compose down -v permanently deletes it.
For a consistent backup, stop writes while archiving the volume:
docker compose stop
mkdir -p backups
docker run --rm \
-v omni-chat-data:/data:ro \
-v "$PWD/backups:/backup" \
alpine:3.22 tar -czf /backup/omni-chat-data.tar.gz -C /data .
docker compose start
Docker without Compose
docker volume create omni-chat-data
docker run -d --name omni-chat --restart unless-stopped \
--add-host host.docker.internal:host-gateway \
-p 127.0.0.1:8080:8080 \
-e CHAT_ENDPOINT_URL=http://host.docker.internal:8000 \
-v omni-chat-data:/data \
-v "$PWD/config:/config" \
registry.gitlab.com/geekaholic/omni-chat:latest
The image runs as uid 10001. On Linux, ./config must be writable by that user
(sudo chown 10001:"$(id -g)" ./config && sudo chmod 775 ./config). Add
:ro to the config mount to forbid saving from Settings. For a local source build,
run docker build -t omni-chat:local . and use omni-chat:local as the
final argument instead.
Importing from ChatGPT
Bring existing ChatGPT history into Omni Chat with the core import chatgpt CLI
subcommand. In ChatGPT, go to Settings → Data controls → Export data; a
download link for the export zip arrives by email.
./.cache/bin/omni-chat-core import chatgpt ~/Downloads/chatgpt-export.zip --user alice
<export> accepts the export zip as downloaded, an already-unzipped export
directory, or a bare conversations.json. A bare JSON file has no attachments to
pull from, so the importer prints a note to stderr and continues without them. Both the classic
single-conversations.json layout and the newer sharded layout
(conversations-000.json, conversations-001.json, … plus
.dat assets with a conversation_asset_file_names.json restoring their
original names) are supported.
| Flag | Effect |
|---|---|
--user <username> | target omni-chat username (required) |
--db <path> | SQLite database path (defaults to $CHAT_DB_PATH or data/app.db) |
--skip-attachments | opt out of importing images/files referenced by messages |
--skip-memories | opt out of importing ChatGPT memories |
--skip-archived | opt out of importing conversations archived in ChatGPT |
--dry-run | parse the export and print the report without writing anything |
For each conversation the importer takes the active branch only — exactly
what was last seen in ChatGPT, not every edited-away alternative. It imports the title,
timestamps, user/assistant messages, reasoning ("thoughts") folded into the following
assistant message, images/files as attachments, and archived state; every imported chat gets a
chatgpt tag. ChatGPT memories (bio-tool writes) are harvested into the user's
user-scope memories, deduplicated against what's already there. Tool-call plumbing (code
execution, browsing traces, canvas) is not imported and is counted in the summary.
Re-running the import against a newer export is incremental: conversations that haven't changed are skipped, conversations that grew in ChatGPT get only their new messages appended, and conversations whose active branch changed (an edit that switched branches) are skipped with a warning — existing messages are never rewritten or deleted. It is safe to import while the server is running (SQLite WAL). Imported chats have no model/endpoint set, so they pick the user's default model the next time they use them.
Importing under Docker Compose
When Omni Chat runs via docker compose, the same binary is the container's
entrypoint, so the import runs as a one-off container that shares the service's database
volume and environment:
docker compose run --rm \
-v "$PWD/chatgpt-export.zip:/import/export.zip:ro" \
omni-chat import chatgpt /import/export.zip --user alice
docker compose run reuses the service's CHAT_DB_PATH=/data/app.db
and the omni-chat-data volume, so no --db flag is needed and the
import lands in the same database the server uses — safe to run while the service is up.
Bind-mount the export read-only at a fresh path such as /import/… (avoid
/tmp, which the service config covers with a small tmpfs). All the flags above
work the same way; add --dry-run first for a report before writing anything. If
the image predates the import feature, rebuild it with docker compose build.
Quality gate
For contributors: make check runs fmt + lint + test + build and must pass before
a change is considered done.
Chapter 13
Troubleshooting & FAQ
No models appear in the picker
The model list comes from the backend's /v1/models. Confirm
CHAT_ENDPOINT_URL is set and reachable, or that your endpoints.json is
valid. Remember that endpoints.json, if present, makes
CHAT_ENDPOINT_URL ignored. A backend that fails at startup is shown as offline and
skipped.
No models appear when Omni Chat runs in Docker
Do not use localhost for an AI server running on the host. Set the endpoint to
http://host.docker.internal:<port>. If you mounted
config/endpoints.json, update its base_url because that file
overrides CHAT_ENDPOINT_URL, then run docker compose restart. On
native Linux, also confirm the AI server listens on a bridge-reachable interface.
SearXNG exits because server.secret_key is unchanged
SearXNG rejects its bundled ultrasecretkey. Run
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env, then run
docker compose -f compose.yaml -f compose.search.yaml up -d again. Compose
reports a clear setup error before startup when the variable is missing.
Empty model picker with a LAN endpoint on macOS
If CHAT_ENDPOINT_URL points to a machine on your local network (e.g.
http://192.168.x.x:8000 or http://<hostname>.local:port) and the
picker stays empty even though curling the same /v1/models URL works,
the cause is usually macOS Local Network privacy. Recent macOS releases block a
process from reaching local-network addresses until you grant it access, which surfaces here as a
"no route to host" error and an offline endpoint; public/internet endpoints are unaffected.
Fix: open System Settings → Privacy & Security → Local Network, enable the
toggle for the terminal app you launch make dev from (Terminal, iTerm2, Ghostty,
VS Code, …) and for Omni Chat when using the desktop build, then
quit and reopen that app. The entry only appears after a LAN connection is attempted. On
Sequoia the grant is per-binary, so an ad-hoc go run or unsigned sidecar can stay
blocked even after Terminal or Omni Chat is allowed. If you launch through tmux/screen/ssh,
start from a plain terminal window so the permission is attributed correctly.
I get a context_overflow error
The input didn't fit the context window. Lower Response Length, shorten the conversation, or
raise CHAT_CONTEXT_WINDOW to match your model. The current user message is never
trimmed — if it alone exceeds the budget you'll see this error.
Creativity changes nothing in the model's behavior
Sampling parameters are only sent when the backend honors them. Set
CHAT_HONORS_SAMPLING=true (or honors_sampling on the endpoint).
Otherwise Creativity only changes the style prompt text.
The model wants a different max-tokens field
Some endpoints require max_completion_tokens. Set
CHAT_MAX_TOKENS_FIELD=max_completion_tokens globally, or
max_tokens_field on the specific endpoint.
No sources are ever returned
Check that an embeddings endpoint is configured and that documents have actually been uploaded and indexed. For web results, check that Search is on for the chat. An empty corpus never produces citations.
Reasoning parameters for thinking models
Use the UI Thinking switch first, and Thinking effort in Settings for models that support
it (Qwen 3.8). Omni Chat detects common adapters such as Qwen on llama.cpp and
capability-advertising models on OMLX.
Configure thinking.on_body and thinking.off_body in
endpoints.json only for custom endpoints or models with different request knobs.
To opt a non-inferred model into thinking effort, set
thinking_overrides[].effort to qwen38.
Chapter 14
Acknowledgements
Omni Chat is released under its own Apache-2.0 license, but it stands on — and interoperates with — a large ecosystem of open-source projects. This chapter credits the projects it bundles, ships as optional sidecars, or connects to. Each project remains under its own license; follow the links for their terms. Thank you to their authors and maintainers.
Models
The default smart-router classifier, Arch-Router-1.5B, is a Katanemo model published by DigitalOcean, LLC under the Katanemo Community License, not Omni Chat's own Apache-2.0 license. Built with DigitalOcean. Commercial use of Arch-Router-1.5B requires a separate license from DigitalOcean — see the model's license terms on Hugging Face before using it commercially. Katanemo Models are licensed under the DigitalOcean Community License, Copyright 2026 DigitalOcean, LLC. All Rights Reserved.
- LFM2.5-1.2B-Instruct (Liquid AI) — the recommended default auxiliary "background intelligence" model (chat titles, memory jobs, router-classifier fallback).
- OmniVoice (k2-fsa, Apache-2.0) — the default zero-shot text-to-speech model, with an Apple-Silicon port from mlx-audio.
- Whisper (OpenAI) — the speech-to-text model, run via faster-whisper in containers and mlx-whisper on Apple Silicon.
- Z-Image-Turbo (Tongyi-MAI) — the default image-generation model, served through Hugging Face diffusers. Optional image-editing models include Qwen-Image-Edit-2509 (Apache-2.0) and FLUX.1-Kontext-dev (Black Forest Labs — a gated, non-commercial license; review its terms before commercial use).
Inference & serving backends
Omni Chat bundles llama.cpp (ggml) as the
Docker sidecar image (ghcr.io/ggml-org/llama.cpp) that serves the router and
auxiliary models. As an OpenAI-compatible client it also interoperates with
Ollama, vLLM,
LM Studio,
LiteLLM, and
Apple MLX (mlx_lm).
In-browser code sandbox
- Pyodide / CPython (MPL-2.0) — Python execution (bundles NumPy and pandas).
- QuickJS (MIT) — JavaScript/TypeScript execution.
- Yaegi — the embedded Go interpreter that runs Go snippets.
- Wasmer (MIT) with WASIX
bashandcoreutils(GPL) — shell command execution. - xterm.js (MIT) — the interactive terminal UI for the sandbox.
Agents
- Pi — the external coding agent behind the
delegate_taskhandoff (npm package@earendil-works/pi-coding-agent). - Codex (OpenAI) — supported as a hosted backend over the OpenAI Responses API ("Sign in with ChatGPT").
Model Context Protocol
Tool integrations use the Model Context Protocol and its official Go SDK. Under Docker, stdio MCP plugins run as sidecars bridged to HTTP by supergateway. Example servers referenced in the docs include @modelcontextprotocol/server-filesystem, figma-developer-mcp, the GitHub MCP server, and Robinhood's MCP server.
The server and web app also build on many open-source libraries — among them Svelte, Vite, marked, DOMPurify, SQLite, and pdfcpu — each under its own license.