Manual

oMNI Chat

A local-first, multi-user AI chat application that runs against any OpenAI-compatible backend. This manual covers everyday use and the configuration that operators need to run it.

Chapter 1

Introduction

Omni Chat is a lightweight chat application built around a single Go core. There is no database server to install and no external service to manage — conversations, projects, documents, and memory all live in a local SQLite file. The core talks to any endpoint that speaks the OpenAI API (/v1/chat/completions, /v1/embeddings, /v1/models), so it works with local engines like llama.cpp and vLLM, a LiteLLM proxy, or hosted providers.

Local-first

One binary, local SQLite storage with vector search. Your data stays on the machine running the core.

Multi-user

OS-login style accounts with private workspaces. Projects can be marked shared for a household or team.

Any backend

Point it at one endpoint, or combine several backends so all their models appear in one picker.

Memory & sources

Optional durable memory of facts and decisions, plus retrieval over documents you upload.

Native Apple apps automatically select AppleFoundation as the auxiliary model on iOS 26 / macOS 26 and later, including OS 27. Apple Intelligence must be enabled and its model downloaded in system Settings. This works independently of the Qwen chat-model download. Under Settings → Models → Auxiliary, choose AppleFoundation, Managed (the Local services llama.cpp), a connection + model, or Disabled; use Test to check availability. Existing custom connections and explicit disabled settings are preserved. A dedicated router classifier takes precedence. On-device Qwen 3.5 2B displays reasoning in the Thinking block and allows opt-in Python/JavaScript code execution through the client sandbox.

Two shells, one core

The same Go core powers two front ends:

  • Web app — runs in any browser over HTTP. This is the supported way to run Omni Chat today.
  • Desktop app — a native Tauri shell that bundles the core as a sidecar and opens it in its own window. Apple Silicon Macs running macOS 26 or later; see Chapter 12.
  • iPad app — the Tauri shell links the same Go core into the application (iPadOS cannot spawn a sidecar), serves the same API on loopback, and can run a verified Qwen 3.5 2B model through MLX on iPadOS 26 or later on M-series iPads with at least 8 GB RAM. Its context is 16,384 tokens, room for web-search results and a longer conversation. Simulator is remote-only. How much the model can write in one reply scales with the device — roughly 4,000 tokens on an 8 GB machine, 6,000 on a large Mac — and is shared between its thinking and its answer, so a model that thinks for too long is stopped and asked again without thinking rather than running out of room mid-reply. Answers stream as they are written, including on turns that can use tools, and the answer after a web search starts quickly because the model does not re-read the whole conversation. The model loads from the downloaded snapshot; it does not compile on the device. The 2B's thinking is capped so a turn where it keeps second-guessing itself is answered within seconds. Apple Silicon Macs on macOS 26 or later use the same model controls and run Qwen 3.5 4B (16 GB of memory or more) or Qwen 3.5 2B (8 GB) with a 32,768-token context. Weights from earlier versions — the GGUF download or the Mac's retired Core AI bundle — are listed as old weights; use Remove old weights to reclaim their space. Close any main-window on-device model notice with the × in its top-right corner; this is remembered for your account on this device. The notice returns if what it says changes — a different reason, download state or error — and model controls appear when a verified bundle becomes available. Open Settings → Models → On-device model at any time to download, resume, pause, or remove the model, even after dismissing the notice. These actions apply immediately; Save & Restart is not needed.
Where the specs live

This manual describes how to use and configure the app. Build and process notes live in AGENTS.md.


Chapter 2

Getting Started

The fastest way to run Omni Chat is Docker. This chapter walks through everything end-to-end: getting an AI backend running if you don't have one yet, picking how much of the optional stack (web search, URL fetch, the auxiliary model, smart routing) fits your hardware, and starting the app. Building from source instead? Skip to Local Development Setup.

What you need

Docker (Docker Desktop, or Docker Engine + the Compose plugin on Linux) and access to an OpenAI-compatible AI backend — something that answers /v1/chat/completions. That can be:

  • Something you already have running (llama.cpp, vLLM, a LiteLLM proxy, an existing hosted provider) — skip ahead to Running Omni Chat.
  • A model you set up locally right now with Ollama or LM Studio — see below.
  • A hosted/cloud endpoint with no local model at all — see Or use a hosted endpoint instead.

Setting up a local AI backend

Ollama

Install: curl -fsSL https://ollama.com/install.sh | sh (Linux), brew install ollama or the installer from ollama.com/download (macOS), or the Windows installer from the same page.

ollama pull llama3.2

ollama run <model> also starts the server if it isn't already running. Ollama serves an OpenAI-compatible API on port 11434 with no /v1 in the base URL:

# Docker
CHAT_ENDPOINT_URL=http://host.docker.internal:11434

# Host-run (make dev)
CHAT_ENDPOINT_URL=http://localhost:11434 make dev

LM Studio

Download from lmstudio.ai, search for and download a model, load it, then open the Developer (Local Server) tab and click Start Server. LM Studio's default port is 1234:

# Docker
CHAT_ENDPOINT_URL=http://host.docker.internal:1234

# Host-run (make dev)
CHAT_ENDPOINT_URL=http://localhost:1234 make dev

Either way, the model you pulled/loaded shows up in Omni Chat's Model dropdown automatically — it's fetched live from the backend's /v1/models.

Or use a hosted endpoint instead

No GPU, or you'd rather not run a local model at all: point Omni Chat at a cloud endpoint. This is also the recommended path at 8GB of memory or less — see How much memory do you have? below.

  • Ollama Cloud — an OpenAI-compatible hosted endpoint with no local install. Create a key at ollama.com/settings/keys (a free tier exists), then:
    CHAT_ENDPOINT_URL=https://ollama.com
    CHAT_API_KEY=<your ollama.com key>
    Pick from cloud-hosted models like gpt-oss:120b, qwen3-coder:480b, or deepseek-v3.2 in the Model dropdown — no local GPU or RAM spent on inference.
  • RouteLLM — an open-source router (lm-sys/RouteLLM) that fronts a cheap/fast model and a strong/expensive one and exposes its own OpenAI-compatible endpoint, routing each request to control cost. Self-host it and point CHAT_ENDPOINT_URL at wherever it listens (hosted RouteLLM-compatible options also exist). This is a different routing axis than Omni Chat's own Auto (Smart Router) — RouteLLM chooses between two backends for cost; Auto chooses among your configured backends for task fit — and the two can be combined.
  • Any other hosted OpenAI-compatible provider you already use — just set CHAT_ENDPOINT_URL and CHAT_API_KEY.

How much memory do you have?

Search and Fetch are lightweight: SearXNG is CPU/RAM only, and the Fetch reader sidecar is a headless-browser process — neither touches VRAM or unified memory, so they're safe to add regardless of hardware. The chat model itself, plus the optional Auxiliary model (recommended: LiquidAI/LFM2.5-1.2B-Instruct, ~730MB GGUF), the Router classifier (recommended: katanemo/Arch-Router-1.5B, ~1GB GGUF), and the Image generation sidecar (see the note below the table) all compete for the same VRAM/unified-memory pool — size your setup by how many of those you plan to run at once.

Available VRAM / unified memoryWhat fitsRecommended path
≤ 8 GB One small chat model (1–3B, Q4) or no local model at all Barebones Compose + a small local model (e.g. Llama 3.2 3B, Qwen2.5 3B), or skip local inference and use Ollama Cloud / RouteLLM / a hosted provider above. Search and Fetch overlays are still fine to add.
8–16 GB A 7–8B chat model (Q4) comfortably Barebones, or + Search/Fetch freely. Add the Auxiliary-model sidecar once you're past ~12GB.
16–32 GB A 7–14B chat model and the Auxiliary-model sidecar together + Search + Fetch + Aux.
32 GB+ 14B+ chat model plus both the Auxiliary-model and Router-classifier sidecars The full stack: + Search + Fetch + Aux + Router.
Rule of thumb, not a guarantee

These are starting points. Actual usage depends on quantization and context length — watch your system monitor the first time you load a new model.

Image generation is the heavyweight

The image sidecar's recommended model, Tongyi-MAI/Z-Image-Turbo (a fast 8-step 6B model, ~12 GB download), wants roughly 16 GB of VRAM on its own in bf16 — and setting IMAGEGEN_EDIT_MODEL_ID for precise instruction edits keeps a second, typically larger model resident (Qwen/Qwen-Image-Edit-2509, the reference open editor, is ~20B — plan for 40–60 GB on top). Plan for it the way you'd plan for a second large chat model: run it on a dedicated GPU box with compose.imagegen-only.yaml when your main machine is already busy with chat inference, or point CHAT_IMAGE_GEN_URL at a hosted OpenAI-compatible images endpoint (e.g. gpt-image-1) and spend no local memory at all. All of these model choices ship as ready-to-uncomment examples in .env.example.

Running Omni Chat: pick your combination

Each step below just adds one more -f flag to the previous command — start barebones and layer on whichever of these your memory tier and needs call for. What each overlay does in depth (healthchecks, image pinning, reader vs. direct fetch, running any sidecar without Compose) is covered in Administration & Deployment; this is just the fast path.

# Once, before any of the combinations below
cp .env.example .env

a. Barebones — chat only, no web search:

docker compose up -d

b. + Search — adds self-hosted web search (SearXNG):

printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.yaml -f compose.search.yaml up -d

c. + Fetch — adds JS-page rendering and YouTube-transcript support for the fetch_url tool:

docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml up -d

d. + Auxiliary model — adds a small dedicated model for chat titles, memory jobs, and (with no dedicated classifier configured) smart-router classification:

docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml \
  -f compose.aux.yaml up -d

e. + Auxiliary model + Router classifier — the full stack; adds a dedicated routing model that outranks the auxiliary model for smart-router classification:

docker compose -f compose.yaml -f compose.search.yaml -f compose.fetch.yaml \
  -f compose.aux.yaml -f compose.router.yaml up -d

f. + Agent handoff & SFTP — lets the assistant delegate real coding tasks to a containerized agent (each conversation gets a private workspace) and gives users SFTP access to their files. The core runs each approved delegation as a sibling container through the host's Docker socket, so this overlay needs a one-time setup before the first run.

1. Build the agent image. This is the container each delegation runs in (Pi plus git):

docker build -t omni-agent-pi docker/agent-pi

2. Create the shared workspace directory. Managed workspaces live here, and the same absolute path must exist on the host, in .env, and in agents.json — because the host Docker daemon resolves the sibling container's mounts against host paths. On Linux:

sudo mkdir -p /srv/omni-agent-work
sudo chown 10001:10001 /srv/omni-agent-work   # 10001 is the core container's user

On macOS (Docker Desktop) the /srv path is on the sealed read-only system volume, so put the directory under your home instead — and skip the chown, since Docker Desktop maps ownership automatically:

mkdir -p ~/omni-agent-work

3. Add the agent config. Copy the example and set managed_root to the directory from step 2 (a literal absolute path — no ~, which the container cannot expand):

cp docs/agents.compose-example.json config/agents.json

The example is managed-mode with a Docker-runtime Pi agent and SFTP enabled. On macOS, edit managed_root to e.g. /Users/you/omni-agent-work.

4. Pre-generate the SFTP host key. The core runs as uid 10001, so a key you create yourself must be owned by that user (no chown needed on macOS). When ./config is writable by uid 10001 the server can also create sftp_host_key on first start:

ssh-keygen -t ed25519 -f config/sftp_host_key -N ""
sudo chown 10001 config/sftp_host_key          # Linux only

5. Point .env at the workspace — this value must match managed_root exactly:

echo 'OMNI_AGENT_WORK_DIR=/srv/omni-agent-work' >> .env    # macOS: /Users/you/omni-agent-work

6. Bring the stack up with the overlay (stackable with the others). On Linux, pass the group that owns the Docker socket so the non-root core may use it; on macOS Docker Desktop, omit DOCKER_GID entirely:

DOCKER_GID=$(getent group docker | cut -d: -f3) \
  docker compose -f compose.yaml -f compose.agents.yaml up --build -d
docker compose -f compose.yaml -f compose.agents.yaml up --build -d   # macOS Docker Desktop
Getting at the files

SFTP publishes on 127.0.0.1:2222 by default; set OMNI_SFTP_PUBLISH_ADDRESS=0.0.0.0 in .env to reach it from the LAN. Users sign in with their Omni Chat username and password and see only their own chats' folders — sftp -P 2222 you@host, or the VS Code SSH FS extension pointed at sftp://you@host:2222. See Getting at your files (SFTP) above for client details.

g. + Image generation — lets the assistant create images ("generate an image of a cyberpunk cat riding a skateboard") with a bundled Z-Image-Turbo sidecar. It requires an NVIDIA GPU (with the NVIDIA Container Toolkit installed; ~16 GB of VRAM) and downloads ~12 GB of model weights on first start:

docker compose -f compose.yaml -f compose.imagegen.yaml up --build -d

Watch docker compose logs -f imagegen on the first run — the app waits for the sidecar's healthcheck, which stays starting until the weights are downloaded and loaded. The default model is Tongyi-MAI/Z-Image-Turbo; image edits fall back to img2img re-imagining with those same weights unless you set IMAGEGEN_EDIT_MODEL_ID in .env to a dedicated instruction-edit model — see the verified choices and their (large) sizes in .env.example and the memory note above. No NVIDIA GPU on this machine? Run compose.imagegen-only.yaml on a GPU box and set CHAT_IMAGE_GEN_URL=http://gpu-host:8001 in .env instead — or point it at any other server that speaks the OpenAI Images API. See Generating images for how it behaves in chat.

Jetson/L4T or arm64 CUDA hosts: the default base image is x86-64 CUDA, so override IMAGEGEN_BASE_IMAGE in .env with a CUDA-enabled PyTorch image matching your host's driver stack — torch is not reinstalled by the build, so it must come from the base — and set IMAGEGEN_RUNTIME=nvidia to expose the GPU. For example IMAGEGEN_BASE_IMAGE=dustynv/pytorch:2.7-r36.4.0 on a Jetson (match your JetPack), or nvcr.io/nvidia/pytorch:25.09-py3 for arm64 CUDA (DGX Spark / Grace-Hopper) — the GB10 (sm_121) needs a CUDA ≥ 12.9 base, or edit-model JIT kernels fail with nvrtc: invalid value for --gpu-architecture.

Apple Silicon Macs: run the sidecar natively, not in Docker

Containers on macOS never see the GPU (the overlay's NVIDIA reservation fails, and CPU diffusion is unusably slow), but the sidecar runs fine directly on the Mac's GPU via PyTorch MPS. From the repo: cd docker/imagegen && python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt torch && uvicorn app:app --port 8001 — it auto-detects MPS and downloads the weights on first start. Then set CHAT_IMAGE_GEN_URL=http://localhost:8001 for a host-run core, or http://host.docker.internal:8001 in .env when the core itself runs under Compose. All the IMAGEGEN_* knobs from .env.example apply as plain process environment variables when running natively — .env only feeds Docker Compose — e.g. IMAGEGEN_EDIT_MODEL_ID=Qwen/Qwen-Image-Edit-2509 uvicorn app:app --port 8001.

h. + Read aloud (text-to-speech) — lets the assistant's responses be read out loud with a bundled OmniVoice sidecar (docker/tts/, Omni Chat's own OpenAI-compatible wrapper around the k2-fsa OmniVoice model). No GPU required — the sidecar auto-detects its device: on CPU it drops to 8 inference steps (usable, though slower than real-time), and a CUDA GPU runs full 32-step quality much faster than real-time. On a Mac, run the sidecar natively instead of in Docker — containers cannot use the Mac's GPU, but the native run picks Apple-Silicon MPS and delivers full quality at nearly real-time speed (from the repo: cd docker/tts && python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt torch && uvicorn app:app --port 8880, then point CHAT_TTS_URL at it):

docker compose -f compose.yaml -f compose.tts.yaml up --build -d

The first start downloads the OmniVoice model weights into the tts-models volume — watch docker compose logs -f tts. For CUDA set TTS_BASE_IMAGE to a torch-CUDA image that also ships torchaudio (omnivoice needs it) and TTS_RUNTIME=nvidia in .env — the same switch on every CUDA host: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime on x86-64, dustynv/pytorch:2.7-r36.4.0 on Jetson/L4T (match your JetPack), or nvcr.io/nvidia/pytorch:25.09-py3 on arm64 NGC (DGX Spark / Grace-Hopper; GB10 needs CUDA ≥ 12.9). The overlay passes NVIDIA_VISIBLE_DEVICES through so the GPU is exposed on all three with no per-host deploy.devices block (TTS_DEVICE is optional — auto-detect picks cuda). To run the sidecar on another machine (or beside a host-run core) use compose.tts-only.yaml and set CHAT_TTS_URL=http://host:8880 plus CHAT_TTS_MODEL=omnivoice; any other OpenAI-compatible speech server works too — e.g. the hosted OpenAI API with CHAT_TTS_URL=https://api.openai.com, CHAT_TTS_MODEL=gpt-4o-mini-tts, and CHAT_TTS_API_KEY. See Read aloud for how it behaves in chat.

i. + Voice input (dictation) — lets you dictate a message instead of typing it, with a bundled whisper sidecar (docker/stt/, Omni Chat's own OpenAI-compatible wrapper around OpenAI's Whisper model). No GPU required — the sidecar runs the faster-whisper engine and auto-detects CPU int8 compute; a CUDA GPU is much faster than real-time. On a Mac, run the sidecar natively instead of in Docker — containers cannot use the Mac's GPU, but the native run picks the Apple-Silicon MLX engine (from the repo: cd docker/stt && python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt mlx-whisper && uvicorn app:app --port 8881, then point CHAT_STT_URL at it):

docker compose -f compose.yaml -f compose.stt.yaml up --build -d

The first start downloads the whisper model weights into the stt-models volume — watch docker compose logs -f stt. For CUDA set STT_BASE_IMAGE to a CUDA-capable image, STT_DEVICE=cuda, and on Jetson/L4T STT_RUNTIME=nvidia in .env. To run the sidecar on another machine (or beside a host-run core) use compose.stt-only.yaml and set CHAT_STT_URL=http://host:8881 plus CHAT_STT_MODEL=large-v3-turbo; any other OpenAI-compatible transcription server works too — e.g. the hosted OpenAI API with CHAT_STT_URL=https://api.openai.com, CHAT_STT_MODEL=whisper-1, and CHAT_STT_API_KEY. With no server backend configured at all, dictation still works via the browser's own on-device speech recognition where supported (Chrome, Safari). See Voice input for how it behaves in chat.

Whichever combination you run, open http://127.0.0.1:8080. The default AI endpoint is http://host.docker.internal:8000; edit CHAT_ENDPOINT_URL in .env to point at the backend you set up above. See Chapter 12 for running without Compose, pinning an image tag, backups, and upgrades.

First run: create the owner

The first time you open a fresh instance, no users exist yet. You'll be taken to a setup screen to create the owner account (username, display name, password). After that, the instance shows a login screen on every visit.

Sessions, not clients

Logging in sets an opaque session token in an HttpOnly, SameSite cookie. Your identity always comes from that session — never from anything the browser sends — which is how the app keeps each user's data isolated.

Add to iPhone Home Screen

When the web app is reachable from your phone, open it in Safari, tap Share, then tap Add to Home Screen. iOS uses the bundled Omni app icon and launches Omni Chat in a standalone browser window from that Home Screen icon.

Building from source instead of using Docker? See Local Development Setup. Ready to combine multiple backends or try the Auto Smart Router? See Models & Backends.


Chapter 3

Local Development Setup

For contributors, or anyone who'd rather build from source than use Docker. If you just want to run Omni Chat, see Getting Started instead.

Prerequisites

Building Omni Chat from source needs the following toolchain. Everything runs through the Makefile, so make is required too.

ToolMinimumWhy
Go1.25+Builds the core (matches core/go.mod).
Node.js20 or 22 LTSBuilds the Svelte/Vite web app.
npmbundled with NodeInstalls and runs the web build.
makeany recentRuns every documented command.
gitany recentClones the repository.
No C compiler needed

Storage uses the pure-Go github.com/ncruces/go-sqlite3 (no CGO), so there is nothing else to install. Rust and the Tauri CLI are needed only to build the desktop app — the core and the web app build without them.

To actually run the app you also need an OpenAI-compatible endpoint already listening — see Getting Started for setting one up with Ollama, LM Studio, or a hosted provider if you don't have one. That's a runtime dependency, not a build dependency.

The README has copy-paste install commands for macOS (Homebrew) and Ubuntu/Debian.

Quick launch

Point the app at a running endpoint and start the development build. This compiles the web app and runs the core with the SPA embedded:

# Ollama defaults to port 11434, LM Studio to 1234, llama.cpp/vLLM commonly to 8000/8080
CHAT_ENDPOINT_URL=http://localhost:11434 make dev

make dev serves the app on port 8080, so open http://127.0.0.1:8080 in your browser. The core also prints one handshake line to stdout when it's ready:

READY port=8080

Add CHAT_PORT to use a different port, CHAT_API_KEY if your endpoint needs a token, and CHAT_DB_PATH to choose where the SQLite file lives — see Chapter 11 for the full list. Running the compiled binary directly (Chapter 12) defaults CHAT_PORT to 0, which picks an ephemeral port and reports it in the READY line. First run creates the owner account the same way as the Docker path — see Getting Started.

Running modes during development

CommandWhat it does
make devBuild the web SPA and run the Go core with it embedded. Simplest single-process workflow.
make dev-coreRun only the Go core. Pair with dev-web for hot reload.
make dev-webRun the Vite dev server for the Svelte app (auto-reloads on changes), proxying API calls to the core.
make build-cliBuild the standalone omni-chat wrapper at .cache/bin/omni-chat.
make build-cli-sidecarBuild the same CLI into desktop/binaries/omni-chat so the desktop bundle can ship it.
make dev-desktopBuild the SPA and core sidecar, then run the Tauri desktop shell. Needs Rust and the Tauri CLI.
make check-ipad-prereqsVerify full Xcode, the Rust iOS device target, CMake, CocoaPods, and Tauri CLI before generating the mobile project.
make check-ipad-simulator-prereqsAdditionally verify the ARM64 iOS Simulator SDK, the Rust target, and that Simulator.app is present in the active Xcode (Tauri boots the device by launching it).
make init-iosIntentionally regenerate the tracked Tauri Xcode project when its mobile scaffolding changes.
make dev-ipadBuild the on-device model framework and embedded Go core, then install/run on an attached iPad.
make build-ipadCreate an iPadOS device build and open it in Xcode for signing.
make dev-ipad-simulatorBuild and run on a selected ARM64 iPad Simulator without device signing.
make build-ipad-simulatorCreate an unsigned ARM64 iPad Simulator bundle without launching it.

See Chapter 12 for production builds and running the compiled binary.


Chapter 4

The Interface

Omni Chat uses a three-column, adaptive layout. The side columns collapse to give the conversation more room; the right Settings panel starts collapsed by default, and each panel remembers its expand/collapse state per user.

Left sidebar — Library

Your projects and chats in a nested tree. Search by title or tag from the Search chats or #tags box — press ⌘⇧P on macOS or Ctrl+Shift+P on Windows/Linux to jump there (it expands the sidebar if it was collapsed). Filter by hashtag from the tag chips. Collapse the sidebar to widen the conversation.

Center — Conversation

Message history plus the composer. The composer grows as you type and shrinks after you send, so answers get the space. It also tucks into a slim peek bar when an answer starts streaming (or when you hide it with the chevron); press ⌘⇧L on macOS or Ctrl+Shift+L on Windows/Linux to expand it and put the cursor in the input for a follow-up. Press ⌘F on macOS or Ctrl+F on Windows/Linux to find text in the open conversation — a small bar highlights matches in visible message text; Enter goes to the next hit and Escape closes it. Those shortcuts do not re-collapse the composer, and they are ignored while a dialog is open. Send becomes Stop while a response is streaming, and a thin orange glow sweeps around the composer border (with the collapse chevron gently breathing) for as long as the model is generating — the hint shows on the collapsed peek bar too, so a live turn stays visible even with the composer tucked away. With reduced motion enabled in your OS, the sweep becomes a static accent ring instead.

Right sidebar — Settings

Per-session response controls: profile, tone, length, creativity, format, sources, and the editable session instructions. It starts collapsed by default to give the conversation more room — open it from the gear button on its rail, and your choice is remembered per user on that device. Expand Profile to reach its style controls, nested Session Instructions, and the selected profile's Skill when one is attached. Both text viewers start collapsed and have their own eye icons.

The top-right user button opens the account menu. From there you can edit your username, display name, and local avatar image, toggle Memory for the current chat, manage saved memories, or sign off.

The composer row also carries lean Model and Profile dropdowns so you can switch either without opening the Settings panel. The rest of this manual walks through each control.

Light & dark theme

The top bar carries a sun/moon button beside the switch-user button that toggles between the light and dark themes. The theme defaults to your operating system's preference and is remembered per user in the browser, so each account keeps its own choice on that device.

Make the interface larger or smaller

In the desktop app, press Cmd++ or Cmd+- on macOS (Ctrl instead of Cmd on other platforms) to zoom the whole interface. Cmd/Ctrl+0 returns to 100%. Zoom steps through 80%, 90%, 100%, 110%, 125%, 150%, 175%, and 200%, and the selected level is remembered for each user on that device. The web app uses the browser's native zoom shortcuts and limits; touch and trackpad pinch zoom remains available.


Chapter 5

Chatting

Starting and organizing chats

  • New chat — start a conversation inside a project or leave it unassigned. After the first answer completes, Omni Chat may replace the initial first-message title with a short generated summary of that first exchange.
  • Projects — group related chats. A project can hold shared instructions that are added to every chat inside it (see Chapter 8).
  • Hashtags — type # to create or attach a tag (for example #research or #implementation). Tags appear as chips and filter the left sidebar.

Chat Details

The Chat Details control in the top bar lets you rename a chat, move it to another project, add or remove hashtags, archive it, or delete it.

Archiving hides a chat from the list without deleting it. Turn on Show archived in the sidebar filters to bring archived chats back into view; Chat Details then offers Unarchive. Deleting is permanent, and both actions ask first.

Sharing & export

The share icon in the top bar exports the current conversation as a standalone document. Both options render what's already on screen into one self-contained HTML page — no server round-trip.

  • Download HTML saves a single .html file with the chat title, every message, rendered markdown, code-run outputs, and web-source citations. Image attachments are inlined as data URLs so the page opens fully offline; other attachments are listed by filename. A chat containing math also carries the KaTeX stylesheet and fonts inside the file. Assistant thinking is not included.
  • Save as PDF opens that same document in a print view — choose your browser's Save as PDF destination to keep crisp, selectable text without any extra tooling. (Allow pop-ups for the app if the print view doesn't open.)
  • Download Markdown saves a .md file with each assistant answer kept as Markdown, code-run outputs as fenced blocks, and web sources as links. Image attachments are referenced by filename rather than embedded, so the file stays small and portable.

Working with messages

  • Copy — copy a message to the clipboard. It is on your messages as well as the assistant's, and copies the text as you typed it (including any code blocks).
  • Edit & resend — change one of your earlier messages and re-run from that point.
  • Edit response — revise an assistant response from a short edit instruction, such as adding sound to a generated game or rephrasing one paragraph. The original stays visible while the edit runs; revision thinking can be expanded, and the answer is replaced only after the edit succeeds. The instruction is not saved as a chat message.
  • Regenerate — re-run the assistant's answer from any prior turn (handy after switching model or profile).
  • Continue — when a response stops short (it reached the length limit, or the connection dropped mid-stream), the newest message shows a "cut off / interrupted" note with a Continue button. It resumes the same message in place — the model picks up where it left off instead of starting over — so long code or documents can be finished across a few continuations. Raise the Response length style setting to reduce how often this happens.
  • Background generation — a turn keeps generating even if you click into another chat, reload the page, or briefly close the lid: generation runs server-side, detached from the browser connection. The generating chat pulses gently in the sidebar so you can find your way back; opening it shows everything streamed so far — thinking trace included — and continues live. On reload the app re-attaches automatically. Only pressing Stop (or leaving no browser attached for the detach grace, CHAT_STREAM_DETACH_GRACE, default 5 minutes) ends a turn early. While a turn runs, other chats are read-only — their Send button explains where the model is busy.
  • Stuck-thinking recovery — a reasoning model that loops in its thinking or thinks past its reasoning budget is stopped automatically instead of burning the whole output window. Running out of thinking budget no longer costs you the turn: the model is asked once more with thinking off, handed the notes it already wrote, and finishes the answer — the message then carries a short "its thinking ran long, so it wrote this answer from its own notes" line. A repetition loop, or an overrun on a model whose thinking cannot be switched off, still shows "The model got stuck thinking, so it was stopped" with two one-click actions: Retry without thinking (re-runs the turn with thinking off, for that turn only) and Regenerate. The partial reasoning stays viewable in the collapsed Thinking block. While a model is thinking, the Thinking label shows a live elapsed timer so long reasoning reads as progress, not a hang. Known local model families (Qwen3, Gemma, gpt-oss, DeepSeek-R1, Nemotron, GLM-4.5+) also automatically get their publishers' recommended anti-repetition sampling settings on self-hosted backends — see sampling_defaults in the endpoints reference.
  • Small models on modest hardware — tool schemas are not free. Shown a wide tool surface, a very small model stops answering and starts routing: asked to pull some fields out of a message as JSON, a 1.2B model replies "I don't have a tool that can do that — would you like me to search?", even though it answers the same question correctly with no tools attached. Testing showed the cause is the presence of tool schemas, not prompt length or tool count — a shorter prompt, an explicit "just answer directly" instruction, and cutting ten tools to two all failed to help. So Omni Chat sizes the tool surface to the model. Below 2B parameters a tool is offered only on a turn that actually calls for it (you pasted a link, asked for a picture, told it to remember something, or asked about something current); below 10B only the external-agent handoff is withheld. Once the model uses a tool the full set returns, so multi-step work like search-then-read still chains. Models whose size is unknown are treated as large and behave exactly as before. Tools a model is too small for are dimmed in the Tools menu with the reason. Override any of this with prompt_profile per endpoint or per model in endpoints.json (auto / full / terse / minimal), or CHAT_PROMPT_PROFILE without a config file. To see how a model behaves under this framing, run make probe ENDPOINT=http://localhost:8000 MODEL=<id>: it replays a fixed set of prompts at the resolved profile and again forced to full, and prints a pass/fail table. It needs a live backend and is not part of make check.
  • Delete with undo — remove a message with a short undo window.
  • Stop — cancel a streaming response; the partial answer is kept.
  • Generation details — hover or focus the info icon for a quick view, or click it to pin the panel until you close it, click elsewhere, or press Escape.

Writing code in the composer

Type three backticks on an otherwise empty composer line to open a monospace code block. Paste or type code there without syntax highlighting; Enter and Shift+Enter both add another line and never submit. Press Arrow Down at the end of the block, or click the blank line below it, to continue with normal text. A populated code block disappears only after all of its contents are deleted or cut; Backspace dismisses a newly opened empty block. After sending, the user message keeps the visual code block while ordinary user text remains literal.

The Generation details panel shows the profile used (or No profile), the exact model and endpoint that handled the response, the SOUL source/hash, whether private identity context was included, memory retrieval mode/fallback and the selected memories with reasons, thinking mode, and the tool outcomes for the response — web access, and, when those features were in play, code execution and agent delegation — plus labeled input/output counts, prefill and generation speed, time to first token, thinking time when available, and total request-to-completion time. Omni Chat saves these details with each new assistant response, so changing a chat's settings later does not rewrite its history. Older responses show only values that can be proved from their saved data and label the rest Not recorded.

Attaching files (images & documents)

Use the paperclip button in the composer — or drop files anywhere on the conversation (the transcript, welcome screen, or composer; Finder drops work in the desktop app too), or paste an image from the clipboard — to attach files to your next message. A drop expands a hidden composer so you can see the chips. Off-the-record chats and an in-flight response refuse the drop, same as the paperclip. Supported types: images (PNG, JPEG, WebP, GIF) and documents (PDF, plain text, Markdown, HTML), up to 20 MB per file and 8 files per message. Attached files appear as chips; remove one with its ✕ before sending. Sent images show as thumbnails in the chat history and documents as clickable chips.

  • Images are sent to the model as image content on that turn, so you can ask "what's in this photo?". This needs a vision-capable model (for example a local Qwen2.5-VL served by llama.cpp or Ollama); the composer warns you when the selected model cannot view images, and Auto (Smart Router) picks a vision model automatically. Original source pixels are preserved and each current turn is resized once for the resolved endpoint's Low VRAM, Balanced, or High detail budget. Use the bounded High image detail option when small text or diagram detail matters; it never sends unrestricted original dimensions. To keep follow-up turns cheap on local hardware, images are only sent in full on the turn you attach them — later turns reference them by name.
    Vision support is detected from, strongest first: model_overrides in endpoints.json; the endpoint's own reported modalities (the llama.cpp router exposes input_modalities per model — a llama-server started without --mmproj reports text-only and is treated as such, so load the mmproj file to enable images); and a built-in catalog of known multimodal families (Gemma 3/4, Qwen 3.5/3.6, Qwen-VL, LLaVA, Pixtral, GLM-4.5V/4.1V, GPT-4o/5, Claude, Gemini). If a vision model is still misdetected, add {"model_pattern": "…", "vision": true} to that endpoint's model_overrides. That override is also the answer for a model whose vision support depends on the build rather than the name — GLM-5.2 ships as both a text-only and an image-capable deployment under the same model id, so the catalog claims no vision for it and you opt in on the endpoint serving the multimodal build.
  • Documents have their text extracted and given to the model as source context, on the turn you attach them and on follow-up questions in the same chat (within the model's context budget).
  • Scanned PDFs with no text layer are handled by extracting their page images and sending those through the vision path — chat about a scanned document exactly like a photo. If nothing can be extracted, the upload is rejected with a clear message.

Reasoning / thinking traces

When a model emits a separate thinking or reasoning block, Omni Chat keeps it apart from the answer and shows it in a collapsible region. Reasoning is stored separately and is never fed back into later prompts — only the answer content is. The compact brain switch beside the composer Profile selector and the labeled switch in Settings are synchronized and update the same per-chat state. When the leading reasoning phase ends, its folded title shows the elapsed time from the first reasoning token to the first visible answer token; this duration remains available after reloading the chat. If the selected model has no known reasoning channel or adapter, the Thinking switch is shown as unavailable. OMLX models are detected from their chat-template capability metadata rather than a model-name allowlist.

Math rendering

Assistant answers render LaTeX with KaTeX, bundled into the app rather than loaded from a CDN, so formulas work offline and look the same in the browser and the desktop app. $$…$$ and \[…\] become a centred display equation — a long one scrolls inside the message instead of stretching it — while $…$ and \(…\) render inline on the text baseline.

Money in ordinary prose is left alone: "it costs $5, so $10 total" stays plain text. A bare $…$ counts as math only under the strict rules TeX itself uses — no space after the opening $, no space before the closing one, no digit immediately after it, and no line break in between — and anything money-shaped fails at least one of them.

Because models routinely emit slightly invalid LaTeX, the renderer is deliberately forgiving. A literal % (as in \mathbf{77.18%}, which TeX would read as a comment and use to swallow the rest of the line) and a literal $ (as in \frac{$140}{$860}, which would otherwise be a parse error) both render as written; the trade-off is that TeX comments cannot be used inside math. Genuinely broken LaTeX shows its source in red rather than failing the whole message, and math inside code spans or fenced code blocks is never rendered. Exported chats carry the KaTeX stylesheet and its font faces inside the file, so math survives offline in a downloaded HTML page.

Code preview

Assistant code blocks include toolbar controls to copy or save the displayed code. The save action opens a Save As picker when the browser supports it, otherwise it falls back to a normal local download. Named blocks keep their filename and unnamed blocks use a language-based snippet.* name; both include an ordering suffix such as main-01.py or snippet-02.js. HTML code blocks in an assistant message can be rendered in a sandboxed preview that runs inline CSS and JavaScript, including dynamically evaluated code used by apps such as calculators, and may load resources over HTTPS from a CDN. Adjacent css and js blocks in the same message are bundled into the same preview. Use the popout control on an HTML preview to open the same sandbox on its own — a new browser tab in the web app, a separate app window in the desktop app; the inline message returns to code view while keyboard-driven demos such as games keep focus in the popout.

Sandboxed by design

The preview runs without allow-same-origin, so previewed code cannot read your app cookies or call the authenticated /api/* endpoints. A popped-out preview in the desktop app gets the same sandbox on an app-internal address of its own, and reaches none of the app's own commands.

Running code blocks

JavaScript, TypeScript, Python, Go, and bash code blocks show a Run button that executes the snippet locally, in a sandboxed WebAssembly runtime (QuickJS for JS/TS, Pyodide/CPython for Python, a Yaegi interpreter for Go, WASIX GNU bash for bash/sh) inside a dedicated Web Worker — nothing is sent to a server. Output (stdout and stderr, the exit code, and how long it took) streams into a panel below the code, with Stop, Re-run, and Close controls. The runtime downloads once on first use and is then cached (the Python runtime is about 6 MB, Go about 8 MB, bash about 4 MB).

Interactive programs: a JavaScript or TypeScript program that reads input with Node readline or process.stdin, a Python program that calls input(), or a bash block runs in a small terminal below the block — type into it and the run keeps going. For JS/TS the run pauses while it waits for a line (idle time doesn't count against the 30-second limit; a runaway loop after you answer still does) and works in any browser with no special setup. Python input() and bash need a cross-origin-isolated context (a SharedArrayBuffer): they work in Chrome and Firefox and show a "not supported" notice in Safari. bash runs your block as a real script (bash -c) with coreutils (ls, cat, and friends); if the script calls read, the terminal stays live for input. It runs without the 30-second wall-clock limit (Stop and the 1 MiB output cap are the guards). Because bash runs in WebAssembly without full POSIX signals (no SIGPIPE), a filter reading an unbounded stream piped into another command — e.g. tr < /dev/urandom | head -c 16 — produces no output; bound the source first (head -c 64 /dev/urandom | …).

Multi-file programs: when code blocks in one answer carry filenames — on the fence (```python main.py) or as a bold/heading line right above the block (**utils.py**) — they form a workspace. Running a named block mounts its named siblings into the sandbox's in-memory filesystem, so import utils, a TypeScript import './greet', or a multi-file Go package works. The files exist only for that run and are discarded with it; blocks without filenames run alone exactly as before. If the same filename appears twice in an answer, the later block wins.

Untrusted by default

Because the code is often AI-generated, the sandbox has no network, no DOM, and no filesystem: fetch, XMLHttpRequest, WebSockets, and browser storage are all removed before the code runs, so it cannot reach your session or the network. Each run is capped at 30 seconds of wall-clock time and 1 MiB of output, memory is capped at 128 MiB, and nothing runs until you click Run. CPU use is bounded only by the timeout — a busy loop simply runs until it is killed.

Python: standard library + bundled offline packages

A curated set of packages is bundled offline and loads automatically when your code imports it: numpy and pandas (plus their dependencies — a one-time ~8 MB download, fetched from the app itself, never the internet). Anything else fails with a ModuleNotFoundError followed by an honest note: package installation (pip/micropip) is not possible because the sandbox has no network access — that is the security boundary, not a missing feature. Graphical packages (pygame, tkinter) wouldn't work in the sandbox anyway: there is no display. If a runtime takes longer than two minutes to download, the run is stopped with “runtime failed to load in time”.

bash has no pip, python, or package managers

The bash sandbox ships GNU bash + coreutils only. When a script calls a missing command (pip, python, node, apt…), the normal command not found error is followed once by a note that only bash + coreutils are available. If an AI answer pairs a Python program with pip install … / python file.py shell instructions, skip the shell block and click Run on the Python code block itself.

Go runs interpreted, standard-library subset

Go support runs a Go interpreter (Yaegi) compiled to WebAssembly: the standard library is available, but networking (net, net/http), os/exec, syscall/syscall/js, and unsafe are withheld from the interpreter as a security boundary, and third-party modules and go get are not supported. Interpreted Go is slower than native and has no goroutine preemption — a runaway program is stopped by the 30-second limit rather than interrupted mid-computation.

Letting the model run code itself (Code Execution)

With a tool-capable model, the assistant can execute code as part of answering — for example "sort these numbers and tell me the median, use a tool". Turn on the Code Execution toggle in the composer — the </> tray button or its Tools-menu switch, which outlines orange when active (off by default). The model is then offered a run_code tool for Python and JavaScript: it writes a program, your browser executes it in the same sandbox as the Run button — no network, no filesystem, no input, 30-second cap — and the printed output flows back into the model's answer. The code never leaves your machine for execution; the server only relays it to your own browser tab.

While it runs, the response shows a Running Python… chip that turns into a collapsed card such as Ran Python · ok · 0.4 s. Expand the card to see exactly what code executed and its stdout/stderr; the card is saved with the message, so you can audit it after a reload too. The model gets at most three runs per response. If the run fails or times out, the model is told so and answers as best it can — check the card when a computed answer looks off. Go and bash are not offered to the model (they stay Run-button-only), and with the toggle off nothing ever executes without your explicit click.

The tool tray and the Tools menu

Per-chat tools live in the composer as a tray of up to four icon toggles plus a Tools menu (the grid button) that lists every tool with an on/off switch; hover a row for its description. Tools whose backing service was never configured don't appear anywhere — a bare install shows only Code Execution, Thinking, Memory, and (when the browser supports speech) Read aloud. A tool whose service is configured but currently down stays in the menu dimmed with the reason and re-enables by itself when the service returns. Enabled tools claim the tray slots first, so whatever you switched on stays one click away; the rest waits in the menu. All of these apply to the chat you're in right now. To set what a fresh chat starts with, use the New chat defaults switches in the right settings sidebar. Those defaults are saved to your account and sync across devices; changing one only affects chats you create afterward, never a conversation already open. The Thinking default is the exception for omni-chat launch: a launch has no chat, so it uses that account switch (on/off) plus Thinking effort. (Agent handoff and MCP Tools are not new-chat defaults — they're situational and start off in every chat until you turn them on.)

Off the record (temporary chats)

The privacy-shield toggle next to the theme switch in the top bar starts a temporary chat: while it is on, nothing about the conversation is saved anywhere. No chat appears in History, no messages are written to the database, and no memory, rolling summaries, or chat title are generated. The transcript lives only in your browser for the session and disappears when you reload the page or open another chat — you'll be asked to confirm first, because it can't be recovered.

Because nothing is recorded, some features are turned off in this mode: Memory is forced off and its switch is locked, and file attachments, image generation, and the Agent handoff are unavailable (they would each leave something on disk). Web search, URL fetch, Code Execution, and MCP tools still work — none of them persist once the turn isn't saved. You can only switch a chat to off-the-record while it is brand new and empty; once you've sent the first message it stays on until you start or open a different chat. A banner across the top of the conversation reminds you the chat isn't being saved.

Opening interactive Pi with Omni Chat models

If Pi is installed on your computer, the standalone wrapper can launch your normal Pi interface with every text model that the Omni Chat model picker can currently see. The macOS one-line installer puts omni-chat on your PATH. From source, build it first:

make build-cli
.cache/bin/omni-chat launch pi

A packaged desktop app already contains the same binary at Omni Chat.app/Contents/Resources/omni-chat. Launch uses your New chat defaults → Thinking switch (and Thinking effort), not the composer toggle of the chat you have open. Restart the launch after changing that default.

The command checks that the backend and its database are healthy, asks for your Omni Chat username and password, then opens Pi. Use Pi’s /model picker at any time; model calls travel through Omni Chat, so OAuth subscription endpoints and Responses-protocol endpoints work without copying their credentials into Pi. Your password and endpoint keys are never written to disk. A short-lived launch token lives only in Pi’s environment, is renewed while Pi runs, and is revoked when it exits. If an endpoint connection drops before producing any model output, the core retries it twice with a short bounded backoff; after any model event arrives, the request is never replayed. When a model does not report an output cap, the temporary Pi provider uses the proxy’s 32,768-token file-writing fallback; an explicit endpoint/model cap still wins, and the proxy clamps the request to the context room left after Pi’s transcript and tool schemas. Reasoning may use that entire resolved output allowance: the proxy does not apply the smaller chat-response thinking budget and does not need to recognize the model family or inference backend. Repetition-loop protection remains active; if a backend streams past the full allowance, the proxy returns a normal finish_reason: "length" and completed stream rather than disconnecting Pi.

The wrapper looks for the server in this order: --server, OMNI_CHAT_URL, the desktop app’s remembered port, then http://127.0.0.1:8080. Remote servers should use HTTPS; non-loopback HTTP requires an explicit --allow-insecure-http. You can choose the initial model and pass normal Pi arguments after --:

omni-chat launch pi --server https://chat.example.com --username alice
omni-chat launch pi --model local/qwen3-coder -- --continue

Your regular Pi sessions, skills, themes, and extensions remain available. The wrapper owns the provider/model routing flags, and its v1 catalog is text-only: it excludes the virtual Auto Router and image-generation backends. Trusted Pi extensions can still make their own unrelated network calls. This command runs Pi on your computer and does not need agents.json; it is separate from Agent handoff.

Handing a task to an external agent (Agent)

Code Execution runs in a sealed sandbox — no files, no network. When a task needs the real thing (edit a project, run its tests, make a git commit), the assistant can hand it to an external coding agent running on the machine that hosts Omni Chat — either installed there directly or inside a Docker container the operator configures (see Running Omni Chat, combination f, for the Compose setup). Built-in adapters support Pi and Hola's hola-coder tail (0.6 or newer). This only works if whoever runs your instance has enabled it (by creating an agents.json that lists which folders agents may touch — or, in the desktop app, filling in Settings → Agent handoff); if you don't see the Agent toggle do anything, it isn't configured on your instance.

Turn on the Agent icon (the robot button) and ask for something that needs a real workspace — "in ~/Work/myproj, add a failing test for the parser and make it pass — delegate it." The assistant writes a self-contained brief and a confirmation box appears showing the folder and the brief. Nothing runs until you approve, and you can edit the brief first — the agent sees exactly what you approve. The agent uses the same model your chat is using — its model calls are relayed through Omni Chat itself, so it works with every configured backend (including subscription ones like the ChatGPT Codex endpoint) and the agent never sees your endpoint's API key. It also inherits the effective Thinking choice for that turn: switching Thinking off applies the endpoint/model’s off-body and non-thinking sampling policy to every Pi or Hola model call, leaving the agent’s own harness to manage planning and convergence.

While it works you'll see a live activity card with an animated working indicator, a live elapsed timer, and a Stop task button — so a quiet agent still looks active. When it finishes, a collapsed card records the brief, what it did, a short summary, and a list of the files it changed — saved with the message so it survives a reload. That file list is a git diff when the workspace is a repository and a plain created/modified/deleted list when it is not, so a fresh per-chat folder reports its work like anything else. If you deny it, stop it, or it times out, the assistant answers with what it has — including the files the agent had already finished, so asking again continues that work instead of starting it over.

When the task succeeds the assistant tells you where the files are rather than printing them into the chat. That is deliberate: the assistant cannot read the agent's workspace, so anything it wrote out would be its own second version of work that is already done, not the file on your disk. Open the file from the workspace folder to see the real result.

If the code is already in this conversation, the assistant can pass saved code blocks directly to the agent. The approval dialog lists the selected blocks. After approval, Omni copies their exact bytes into separate input files in a .omni-handoff-* folder; the transfer does not overwrite project files. Up to eight blocks (1 MiB total) can be passed, separately from the task brief. Their references remain in the saved run card. This includes code in your current or edited message, and the original answer during an assistant revision.

Each turn permits an initial handoff and one corrective handoff, with a new approval for each. A workspace picked or reset during approval also applies to the corrective handoff. Denying or stopping a task ends further handoffs for that turn. Run finished means the agent process finished; check its summary and validation results to see whether the requested task succeeded. After three consecutive failed tool calls, Omni stops tool attempts and asks the assistant to explain the remaining blocker.

Image tools are for pictures relevant to your request, not for retrieving code from chat history. Omni rejects invented image links before downloading: links must come from your messages, image-search results, a fetched page, or a supported live-chart endpoint. Revising an answer does not turn assistant-invented links into user-supplied links; links in your actual edit request are still eligible.

A finished Pi task leaves its session behind, so you can open a terminal, cd into the workspace folder and run pi --continue to pick the work up by hand. Give it a provider and model when you do — pi --continue --provider <yours> --model <yours> — because the session records the temporary routing the app set up for that one run, which no longer exists once the run ends. A task you stop mid-run leaves the session folder but no transcript: the agent is killed rather than asked to exit, so there is nothing for it to flush. The agent can only ever work inside the folders your instance allows, one task at a time.

On a shared (multi-user) instance the operator usually configures managed workspaces: each conversation gets its own private folder, created the first time you delegate and reused for follow-ups in the same chat — so "now add sound effects" continues where the last task left off. Other users can never see your folders. In this mode you don't pick a directory; the confirmation box shows the one assigned to the chat.

In the desktop app the confirmation box also lets you change that folder: press Change… to browse, or create a new subfolder from the file dialog, and the task runs there instead. The choice sticks for the rest of that conversation — follow-ups reuse it without asking again — while a new chat starts fresh with its own folder. Use this chat’s own folder undoes it. You can only pick somewhere inside your own workspace area, so a chosen folder is still private to you and still reachable over SFTP; anywhere else is refused with the allowed folder named, and the box stays open so you can pick again. If a folder you chose is later renamed or on a drive that is not mounted, the task falls back to the chat's own folder and says so, and your choice is remembered for when it comes back.

Getting at your files (SFTP)

If your instance has SFTP enabled (managed workspaces only), you can browse and copy everything the agent built — and drop files in for it to work on next — using any SFTP client. Sign in with your Omni Chat username and password; you only ever see your own folders, one per conversation. Each run card shows that conversation's folder name (cht_…) next to two copy buttons: one copies the folder's path on the server, the other (the server icon, shown only when SFTP is on) copies a ready-to-paste sftp://you@your-server:2222/cht_… address — drop it into any of the clients below and you land straight in that conversation's folder. Paste it into an SFTP client or the Terminal, not a web browser — browsers can't open sftp:// links, so Safari or Chrome will show an "invalid address" error.

  • VS Code: install the SSH FS extension and add a connection to sftp://you@your-server:2222 — your workspaces mount like a local folder. Note it must be SSH FS, not Microsoft's Remote - SSH: Remote - SSH needs to run a server program on the remote machine, and this port deliberately allows file access only — no commands.
  • FileZilla / Cyberduck: protocol SFTP, host your-server, port 2222, password login.
  • Terminal: sftp -P 2222 you@your-server.

The port may differ on your instance — ask whoever runs it. There is no shell login on this port, just file access.

Calling tools on MCP servers (MCP Tools)

Beyond the built-in tools, the assistant can call tools from MCP servers (Model Context Protocol) that the operator has configured — remote services or local tool programs. Like Agent handoff, this only exists if your instance enables it (by creating an mcp.json); when configured, an MCP Tools icon (the server button) appears in the composer and the right settings sidebar lists the configured servers. It is off in every chat until you turn it on, and it is not a new-chat default.

With the toggle on, the model may call the servers' tools while answering. Every call asks you first: a confirmation box names the tool and server in plain language, explains what the tool does in the server's own words, and lists every argument on its own labelled row — so you can read what will happen instead of decoding JSON. Hover the ⓘ beside an argument for the server's description of it. A long value (a file's contents, say) is clamped to a readable preview with Show all N lines to see the rest, and View raw JSON shows the exact payload whenever you want to check it character for character. Nothing runs until you approve. For a tool you trust (say, a read-only search), tick "always allow this tool in this chat" and it stops asking — the grant applies only to that tool in that chat, never globally. Denying a call is fine: the assistant is told and answers with what it has. The response's info popover shows whether MCP tools were used, and each call (server, tool, outcome, duration) is stored with the message.

Some servers never let you skip that prompt. If the operator marked a server always ask — because its tools do something in the real world that shouldn't repeat unattended, like placing a trade — you'll see a shield note where the always allow checkbox usually sits, and a shield beside the server in the settings sidebar. Approving one call on such a server never carries over to the next one.

Each configured server has an on/off switch in the settings sidebar's MCP servers list. Turn one off and it offers no tools in any of your chats until you turn it back on — a quick way to silence a single server (say, one you're not using right now) without affecting the others. This is your own setting and applies across all your chats; it's separate from the per-chat MCP Tools switch in the composer, which turns every MCP server on or off for just that one chat.

Some servers require signing in with your own account — they show Needs connection in the settings sidebar's MCP servers list. Click Connect: the provider's sign-in page opens — in the desktop app, in your default browser, since many providers refuse to sign in inside an embedded window. Approve it there and close the tab; the list checks for a couple of minutes and updates itself, or use Refresh status. Your sign-in is yours alone — other users on the instance connect their own accounts — and Disconnect removes it at any time. Until you connect, such a server simply offers no tools. Setup for operators is in Chapter 11 and the README's "MCP tools" section.

Trading on Robinhood. Robinhood publishes an MCP server, so if your operator has added it you can ask about your portfolio and place orders from a chat. Connect it the same way — Connect next to robinhood in the settings sidebar, then sign in on Robinhood's page (this needs a desktop browser; Robinhood sets up your Agentic account during that first sign-in). Two things to know: Robinhood only lets an agent place trades in that separate, separately funded Agentic account — reads cover your other accounts, orders don't — and every single call asks you first, showing the symbol, side, and quantity as labelled rows. Approving one order never pre-approves another. Read each prompt: a model can misread a request or act on a stale quote, and denying is always the safe answer.

Operators offering several servers: the model is offered at most max_offered_tools MCP tools per turn (a top-level field in mcp.json, default 128, max 512), filled in server order. If you connect several large servers whose combined tools exceed that, the later servers' tools are silently dropped (the core logs mcp: offered tool cap reached) — raise max_offered_tools, reorder the servers, or narrow a big one with tool_allowlist.

Operators running under Docker: the omni-chat container ships without Node, so MCP servers launched with npx cannot run inside it. Use the compose.mcp.yaml overlay instead — it runs each npx-installed plugin (a Figma server and a filesystem server come ready-made) in its own Node sidecar bridged to Streamable HTTP, and config/mcp.json points at them with ordinary "transport": "http" entries (e.g. http://mcp-figma:8000/mcp). Secrets like FIGMA_API_KEY go in .env and reach only the sidecar. Details and the add-another-plugin recipe are in the README's "Docker: running npx MCP plugins" section.

Generating images

If your instance has an image backend configured, just ask: "Generate an image of a cyberpunk cat riding a skateboard." A tool-capable model writes a detailed prompt (you can ask for square, portrait, or landscape) and an animated placeholder appears in the response the moment generation starts, already shaped like the image to come. With the bundled backend you then watch the image form: a blurred preview of the actual picture appears part-way through and sharpens as the percentage climbs, until the finished image replaces it in place, with the prompt as its caption. Generation happens on the server's configured backend and can take from seconds to a couple of minutes depending on the hardware — the response simply continues once it lands.

Each image has three buttons: Download saves it (named after the prompt), Variation drops a ready-made variation request into the composer — edit it if you like, then send, and the model generates a fresh take — and Edit starts an image-to-image edit of that exact picture (see below). Images are stored with the conversation (they survive reload and appear in chat exports) and are private to your account. The assistant can produce at most two images per response; if the backend is offline the assistant says so and answers without one. The image server draws one picture at a time — if someone else's generation is running, the placeholder reads "Waiting for the image server… (N ahead)" until it's your turn (waiting never counts against the timeout), and when the wait queue is full the assistant simply reports that the server is busy and suggests trying again shortly. If nothing happens when you ask for images, the instance has no image backend configured — see Running Omni Chat, combination g.

Editing images (image-to-image). The assistant can also transform an existing picture: attach a photo and say "make this look like a watercolor", say "now make it night time" about an image it just generated, or click Edit on any generated image — it prefills the message with that image's reference so you just type the change. While an edit runs, the placeholder starts from a blurred copy of the source image, and the result carries an Edited caption. This works with any tool-capable chat model — the chat model does not need vision; it hands your instruction and an image reference to the image backend, which does the actual transformation. (If your instance's image backend runs without a dedicated edit model, edits are re-imaginings guided by your instruction rather than surgical changes — style changes land better than element edits, and describing the full desired result works better than a terse instruction. Ask whoever runs it about IMAGEGEN_EDIT_STRENGTH for a stronger restyle, or IMAGEGEN_EDIT_MODEL_ID for precise instruction edits.)

Read aloud (text-to-speech)

If your instance has a speech backend configured, every assistant response grows a speaker button in its action row. Click it to hear the response read aloud in your chosen voice — the icon becomes a stop button while the audio is prepared and while it plays; click again to stop. Only one message speaks at a time, and replaying the same message in the same session starts instantly (the audio is kept in memory, not re-synthesized). Code blocks are skipped with a short spoken note, links read their text, and tables are read row by row.

For hands-free use, the composer has a Read aloud toggle (in the tray or the Tools menu): while it's on for a chat, each response is read aloud automatically the moment it finishes streaming. Responses that error out or that you stop are never spoken. A matching switch under New chat defaults in the settings sidebar makes auto-speak the default for your new chats.

Expand the default-collapsed Voice & Dictation section in Settings and use its Voice Engine group: it contains the engine, a voice list (from the instance's speech backend, or several grouped lists when more than one is configured), a speed control (0.5×–2×), and a Preview button that speaks a sample sentence. With the bundled OmniVoice backend, a voice that isn't one of the listed presets is used verbatim as a voice description — save your own, e.g. male, elderly, low pitch, british accent (attributes: gender, age, pitch, accent, whisper style). The preference is saved to your account, so it follows you across devices. Nothing about read-aloud is stored on the server — audio is synthesized on demand and streamed to your browser, and it works in off-the-record chats too.

No speech backend? Your device speaks. When the instance has no speech backend configured — or every configured one is down — read-aloud automatically falls back to your device's own OS voice through the browser's built-in speech synthesis: instant, free, and offline. The Engine select in Voice Engine (shown when both options exist) lets you force it: Auto prefers the server backend when it's healthy, This device always uses the OS voice. The device engine speaks with the voice configured at the OS level (macOS: System Settings → Spoken Content) — there is deliberately no in-app list of the hundred-plus system voices — and it honors your speed preference. To get the richer OmniVoice/hosted voices instead, see Running Omni Chat, combination h.

Voice input (dictation)

The composer has a mic button before Send. Click it (or press Ctrl+Shift+D) to start recording; click it again (or press the shortcut again) to stop — the recording is transcribed and the text lands at your cursor. Press Escape while recording to cancel and discard the audio instead. A recording auto-stops after 3 minutes.

Dictation uses one of two engines: the instance's server transcription backend when one is configured, or your browser's own on-device speech recognition (Chrome, Safari) when it isn't — or when you'd rather keep the audio on-device. Pick which under Dictation in the default-collapsed settings sidebar Voice & Dictation section: an engine select (Auto / Server / This device — Auto prefers the server backend when it's configured and healthy) and a language select (auto-detect, or pin a language for more accurate transcription). The preference is saved to your account, so it follows you across devices.

Privacy: audio is transcribed and discarded, never stored in the server's database. The core holds it in memory only; the bundled whisper sidecar writes a transient temp file it needs for transcription and deletes it immediately afterward, so nothing is retained on either engine. It works in off-the-record chats too. If the mic button doesn't appear at all, the instance has no server transcription backend configured and your browser doesn't support on-device speech recognition either — see Running Omni Chat, combination i.


Chapter 6

Models & Backends

Choosing a model

Pick a model from the Model dropdown in the composer row. The list is fetched from your backend's /v1/models endpoint. Each chat remembers both its model and its backend, so reopening a chat restores the right one. For a fresh chat, the picker defaults to the model you used last — preferring it only when it is online (loaded, loading, or from a backend that doesn't report residency). If it isn't, the next online model on the same backend is chosen, else an online model from another backend; and if a backend disappears while you're on the welcome screen, the selection heals itself the same way on the next refresh — returning to your remembered model once its backend is back.

You can't switch models mid-stream — the picker is disabled while a response is in flight. Use Stop (or regenerate afterwards) to change models.

A small status LED sits beside each real model inside the composer Model picker, and the closed picker repeats the selected model's LED next to its name. Green means the selected backend confirms the model is loaded and ready; amber means it is loading; neutral means it is unloaded, sleeping, or the backend does not expose a trustworthy signal. The Auto (Smart Router) entry and the backend divider rows carry no LED. The sidebar's Image generation dropdown uses the same dots for availability instead: green = connected and healthy, amber = needs connecting (e.g. a Codex sign-in), red = configured but failing, grey = off — matching the Services panel. Omni Chat refreshes the state while the app is visible, and polls more often during the first local-network scan so a server found after the window opens still appears in the picker. It auto-detects Ollama, LM Studio, llama.cpp (single-server and router modes), vLLM, and oMLX status APIs. A failed optional status check never removes an otherwise available model.

Combining multiple backends

Omni Chat can present models from several OpenAI-compatible backends in one picker — for example a local vLLM/LiteLLM proxy alongside a hosted provider. The dropdown groups models under a divider per backend; choosing a divider selects that backend's first model. A backend that fails to load at startup is marked offline and skipped — the others keep working.

This is configured with an endpoints.json file, covered in Chapter 11. With no such file, the app uses CHAT_ENDPOINT_URL, which may itself be a single URL or a comma-separated list.

Subscription backends (Sign in with ChatGPT)

Some backends authenticate with a subscription sign-in instead of an API key. The first supported one is Codex ("Sign in with ChatGPT"): declare it in endpoints.json with "auth": {"type": "oauth", "provider": "openai-codex"} and "protocol": "responses" (see the README's "Multiple backends" section for the full snippet). The instance owner then connects it once from the settings sidebar's Services section — like an API key, the connection powers everyone's chats on that backend. On a desktop or homelab install the sign-in completes automatically — and under Docker Compose too, once you set OMNI_OAUTH_PUBLISH_PORT=1455 in .env, which forwards the provider's fixed localhost:1455 callback into the container when your browser runs on the Docker host. That mapping is off by default: claiming a fixed host port for an optional feature would stop the whole stack from starting on a machine where another app already listens on 1455. Without it — or from any other machine — the sign-in tab ends on an unreachable or unrelated localhost:1455 page; that's expected, not an error: copy that tab's full address into the paste the callback URL field the panel offers. Until it's connected, the backend's row shows an amber dot with a Connect button (image backends sharing the same sign-in show the same state), and its models simply don't appear in the picker. The sign-in key is created in the config directory, or beside the database when the config directory is read-only (e.g. an :ro Docker bind mount) — no setup needed either way.

Auto (Smart Router)

Whenever at least one real model is available, the Model dropdown also offers a distinct ✨ Auto (Smart Router) entry pinned above the per-backend groups. Select it and every message you send is automatically classified and routed to the best available model for that turn — weighing the task type (coding, reasoning, a quick question), the capabilities the turn needs (tool calling for web search, a thinking channel), and the request's complexity and size, so heavier or larger-context turns favor bigger models and trivial ones favor faster models. No manual model-switching needed. Auto is opt-in: it is never the default for a new chat.

While a chat is on Auto, the composer shows Auto → <model> once a response has routed, and the assistant message's Generation details panel shows which model and backend handled it, the routing reason, and how long the routing decision took. Because the concrete model — and therefore its context window — is chosen per message, the context meter is hidden while Auto is selected.

Routing also weighs each model's resolved capabilities — vision, tool calling, parameter size, and vendor/family — taken from a built-in catalog of common model ids, from id inference, or from a per-endpoint model_overrides array in endpoints.json. A message with an image attachment is routed to a vision-capable model whenever one is available.

Loaded models first. On servers that report which models are in memory (Ollama, LM Studio, the llama.cpp router and oMLX — the same state the picker's status LEDs show), Auto picks only among models that are already loaded, so a quick question never makes the server unload one model and load another. It loads a different model only when no loaded one can handle the message: an image with no loaded vision model, a prompt longer than their context, or web search with no loaded tool-calling model. The routing reason then reads no loaded model fits; loads <model>. Models on your own servers whose load state can't be read count as not loaded; hosted cloud APIs always count as ready.

Jev by TypeSafe AI. For better judgement of each message, open Settings → Models → Smart Router, turn on Use Jev and paste a TypeSafe API key. Jev is a decision model rather than a chat model: one call per message (about 0.2 s) rates how demanding your latest message is on five levels — trivial, simple, moderate, hard, expert — whether it needs specialist knowledge, and whether you asked for a stronger model. Auto then picks the model, so no local model is tied up classifying. Press Test Jev to check the key and connection before saving; it makes one tiny call and reports the Jev version that answered and how long it took. Your message and a little recent context are sent to TypeSafe. If Jev errors or takes longer than its timeout (default 1500 ms), the built-in router decides instead — a message is never held up. Routing reasons judged by Jev start with jev:, and Smart Router (Jev) appears in the service status panel. On the desktop the key lives in the Keychain as OMNI_KEY_ROUTING; a web install saves it in endpoints.json's routing section.

Cloud models in Auto. If you have both your own models and a cloud subscription (such as ChatGPT), Settings → Models → Cloud models in Auto decides when Auto may use the cloud. Only for expert requests (the default) sends expert-level messages, explicit requests like "use your strongest model", and follow-ups in a thread a cloud model already answered; everything else stays on your own models. Always lets cloud models compete on every message; Never keeps Auto local whenever a local model can answer.

Follow-ups. A substantive follow-up stays within one band of the model that answered the thread, and prefers that same model, so a derivation that started on a strong model is not continued by a weak one. A simple follow-up — "thanks", "sum that up in one line" — drops to the fastest able model. A reply that continues the previous request — answering the assistant's question, pointing it to something earlier in the chat ("the code is above"), or saying "yes, do that" — stays with the model that was handling it, however short it is.

Auto works out of the box with a built-in policy. Power users can customize model tagging and per-task preferences with a router.json file — see Chapter 11 for the full schema and how tagging and routing decisions work.

Service availability & self-healing

Optional integrations — the agent runner, the image-generation sidecar, web search, web fetch, embeddings, and MCP servers — are tracked live via GET /api/status. While the app is visible the core re-checks each configured service with a cheap probe (the image sidecar's /health, a docker version for containerized agents, the search/reader base URL, the embeddings server's /v1/models), and every real tool call also feeds the same status. Nothing is probed while no browser tab is open.

The same holds at startup. If an optional service is configured but its environment is not ready — the agent workspace folder cannot be created, Docker is not running, an agent's command or image is missing — Omni Chat prints a warning saying exactly which service and why, disables just that feature, and carries on serving chat. Only a genuinely broken config file stops it from starting. The agent workspace is re-checked on every status poll, so once the folder exists (or its permissions are fixed) agents come back on their own, without a restart.

When a configured service is unreachable the UI degrades instead of pretending: the composer's Agent toggle dims with the reason in its tooltip, and a small alert triangle appears in the composer row. Clicking the triangle lists what is degraded right now (failed backends included) plus the recent offline/online history. The Settings panel's Services section shows everything the instance can talk to in one list, grouped as Chat backends, Image generation, Speech, Tools, and MCP servers, with one status language: green = working, amber = needs your action (e.g. a subscription sign-in to connect — the button is right on the row), red = failing (hover the dot for the reason), grey = not set up. A subscription credential is never shown as a separate line item; its state appears on every backend that uses it. Everything recovers automatically: within one status poll (~15 s) of a service coming back, toggles re-enable, dots turn green, and the alert clears. Each transition is also logged by the core (service down / service recovered) for debugging.


Chapter 7

Profiles

A profile is a one-click starting point. Selecting one applies a persona prompt plus a bundle of style settings (tone, length, creativity, and format) for the current session.

Built-in profiles

ProfileFor
Brainstorm IdeasExplore options and tradeoffs. Attaches a diverge-then-converge skill.
Draft ContentCreate reusable content. Attaches an audience-first drafting skill.
Write CodeBuild, debug, and refactor. Attaches a run_code vs delegate_task skill.
Learn TopicLearn a subject in the chat; offers a how-to or mini-book if you want a PDF.
Write a How-toA short how-to (one skill, one sitting) compiled to a PDF.
Write a Mini-bookA focused mini-book on one subject, chapter by chapter.

Selecting a profile

The welcome screen shows up to eight profile cards total; every configured profile is also in the Profile dropdowns. The synthetic ✨ Auto entry described below is shown first when present and counts as one of the eight cards. Pick a card to apply its persona and style. In Settings, the Profile dropdown remains visible while its accordion is collapsed; selecting a profile expands it. The expanded area contains the style controls, the nested Session Instructions editor, and a read-only Skill viewer when the selected concrete profile has an attached procedure.

Profile identity is saved; its settings remain session-only

Each assistant response stores the canonical profile ID and a label snapshot for that turn. Reopening a chat selects the profile used by its newest visible assistant response. A newest No profile response restores No profile. If a custom profile was removed, its old responses keep the saved label but the selector falls back to No profile. If that response was Auto-routed (labeled Auto → <Profile Label>, or a bare Auto), reopening restores the Auto profile selection instead of the concrete persona it happened to route to that turn. The persona, style bundle, generated system prompt, and prompt edits are not persisted as profile settings.

Every built-in welcome card attaches a matching skill (a procedure, not extra persona text). Write a How-to and Write a Mini-book are the writing skills: the assistant authors the teaching, an agent only typesets, and generated cover images from the chat are copied into the workspace so they can actually land in the PDF. After a compile, the PDF and page previews are attached to the assistant message so you can open them in the chat. Pick one of those cards before asking for a document.

You can define your own profiles to extend or override the built-ins, and attach a skill (a Markdown procedure) with an optional skill field or by using the same id — see Chapter 11 (Custom profiles and Skills). Starter overlay skills for the example catalog live in docs/skills.example/.

Auto profile

The Profile dropdown also offers a synthetic ✨ Auto entry. auto is a reserved id — you cannot define your own profile with that id. Picking Auto is an autopilot: it also switches the chat to the Auto (Smart Router) model (when that model is available), so one choice hands both the backend model and the per-turn persona to the router. Auto only picks a persona when the chat's model is Auto (Smart Router). On such a turn the profile catalog is offered to the router: the semantic router classifier, when configured in schema mode and answering in time, picks the best-fitting profile; otherwise (Arch-Router, or the classifier disabled or too slow) a deterministic task-affinity heuristic matches the turn's task to a profile whose description reads like that kind of work. Either way that profile's persona and style are used and the response is labeled Auto → <Profile Label>, just like a manually-picked profile. When no profile fits (e.g. a trivial greeting) or the chat uses a manually-selected model, Auto behaves exactly like No profile and the response is simply labeled Auto.


Chapter 8

Response Controls

The Settings panel shapes each response. These are session-level controls; a profile simply sets several of them at once.

Tone

Single-select: Default, Professional, Friendly, Direct, Executive. Controls tone only. Default adds no tone instruction.

Response Length

Single-select target for the output budget:

SettingTarget output tokens
Short2000
Medium8000
Detailed12000

Creativity

Single-select: Precise, Balanced, Exploratory. This always changes the style prompt. It sends temperature and top_p sampling parameters only when the backend is configured to honor sampling (CHAT_HONORS_SAMPLING=true, or honors_sampling on the endpoint). Otherwise those parameters are intentionally omitted.

Format

Single-select shape for the output: Natural, Structured, Bullets, Step-by-step, Table-first.

Session instructions

Session Instructions is nested inside the expanded Profile controls. It keeps an independent eye button for showing or hiding the editor and a reset button for restoring the generated instructions. Omni Chat assembles a single upstream system message from, in order: app safety invariants, hidden real-time context, assistant soul, editable session instructions, retrieved memory, retrieved sources, and summaries. The panel shows only the editable session instructions generated from project/profile/style settings; managed safety and SOUL are hidden.

Edits are session-only

Session-instruction edits apply to the current session and are not written to SQLite. Reloading discards those edits and regenerates the defaults for the restored current profile. App safety and assistant soul are enforced separately by the core and cannot be removed from the Settings textarea.

Skill

When the selected concrete profile has an attached skill, a Skill row appears after Session Instructions. Its independent eye button reveals the exact managed Markdown in a read-only panel. It starts hidden after reload and keeps its disclosure state while you switch profiles during that page session. Default and Auto show no Skill row because neither has one stable attached procedure; Auto's concrete profile is chosen per turn.


Chapter 9

Memory

Memory lets Omni Chat remember durable facts, preferences, and decisions across turns and chats — separate from the raw message history.

Your private assistant context

On first sign-in, an optional, skippable form lets you tell Omni Chat the name it should use, pronouns, location/timezone, role, organization, interests, goals, and a short “about” note. This small private profile is included on every normal turn independently of Memory, so basic identity does not depend on search ranking. Edit it later from Edit Profile. It is visible only to your own session and never in the instance member list.

The Memory toggle

  • On (default) — the assistant can save durable memories and recalls the relevant ones on later turns, injected into its prompt.
  • Off — chat history is still stored, but nothing is written to or read from durable memory for that chat.

The toggle lives in the composer's Tools menu and is remembered per chat.

Asking the assistant to remember

With Memory on, just tell the assistant — for example, "Remember that I'm vegetarian" — and it saves a durable memory; a "Memory · Saved" notification in the top-right corner confirms it for about 15 seconds, with Undo in case it saved something you didn't want kept. In a later chat, ask "What kind of food do I eat?" and it answers from that memory without being told again. Say "Forget that I'm vegetarian" to remove it.

Memory kinds

preference, fact, decision, constraint.

Scopes

ScopeVisible to
userYou, across all your projects (private).
projectEvery chat within the project.
workspaceMembers of a shared group (arrives with sharing).

Managing memories

Open Manage Memories from the top-right user menu to see everything the assistant remembers about you. From there you can filter by scope, edit a memory's text or kind, pin the ones that should always be eligible for context, add a memory manually, delete any of them, and see each one's provenance — the chat and exact user-message evidence it came from. Use the status filter to audit quarantined automatic memories and restore only the ones you trust. Your memories are private to your account.

Smarter recall (optional)

If the operator has configured an embeddings endpoint (CHAT_EMBEDDINGS_URL and friends), memory recall becomes semantic: the assistant surfaces the memories most relevant to what you're currently asking, and saving something you already mentioned — even in different words — updates the existing memory instead of creating a near-duplicate. Without an embeddings endpoint, recall still uses lexical, entity, and dotted-acronym matching; it never fills the context with unrelated recent rows. Either way, memory works.

Automatic memories

If the operator has configured the auxiliary model (CHAT_AUX_URL and CHAT_AUX_MODEL), the assistant quietly captures durable facts you mention — lasting preferences, your ongoing projects, decisions, and constraints — without you having to say "remember this". This runs in the background after a reply, so it never slows a conversation, and it uses a separate small model rather than your main chat model. It is conservative by design: it saves only clearly-stated, lasting facts, skips sensitive data (passwords, card or account numbers, secrets), and requires an exact quote from one of your messages as evidence. The assistant's own answer can never become a fact about you. Every automatic memory is announced with a brief "Memory · Remembered" notification in the top-right corner a few seconds after the reply — click Undo to drop it on the spot, or View to open the Memories panel (hovering the chip keeps it open). Every automatic memory shows up in the Memories panel with its evidence, so you can review, edit, quarantine, restore, or delete it. Existing unsupported automatic memories are quarantined on upgrade. Auto-capture is off unless the operator turns it on, and you can disable it per chat with the memory save policy.

Staying coherent in long chats (automatic context compaction)

When a conversation gets very long, it eventually exceeds the model's context window and the oldest messages would normally fall out of view. Instead, the assistant keeps a short rolling recap of the earlier part of the chat, refreshed quietly in the background as the conversation grows — each refresh folds the new turns into the previous recap, so even the very beginning of a long chat is never forgotten. When old messages no longer fit, the assistant sees the "Earlier in this conversation…" recap in their place, so it stays on topic instead of losing the thread. This works even without the optional auxiliary model: the recap is then generated by the chat's own model between turns, and it never slows down a reply.

Your transcript is never changed — you can always scroll back through the full conversation. When compaction was active on the latest response, a subtle dashed divider marks the boundary: messages above it reached the model only as the recap. Click View summary on the divider to read exactly what the model is told, and the response's info popup shows a "Context: Compacted" row. You can still edit, regenerate, or delete any message, including ones above the boundary — doing so simply discards the now-stale recap and rebuilds it in the background a turn later.

Tidying up over time

As your memories accumulate, the assistant occasionally tidies them in the background (with the same optional memory model, rate-limited so it runs at most once in a while). It merges duplicates, retires memories that a newer one has contradicted or made obsolete, and distills recurring themes into higher-level “insight” candidates. Because those candidates are model inferences rather than direct user quotes, they stay quarantined until you restore them. Your pinned memories are never touched. Nothing is ever permanently erased: a retired memory keeps a link to the one that replaced it, so the whole history stays auditable and reversible. The result is that a long-lived memory set gets sharper and less cluttered over time instead of filling up with near-duplicates.

Remembering across conversations

Sometimes the thing you need was discussed in a different chat entirely. When cross-chat recall is available (it needs the embeddings model configured), the assistant can pull in a short recap of your other conversations that are relevant to what you're asking now — shown to it as "From earlier conversations…". So if you start a fresh chat and ask about a project, decision, or topic you worked through days ago, it can pick up the thread instead of starting from nothing. It only ever draws on your own chats, surfaces just the few most relevant ones, and follows the per-chat Memory switch — turn Memory off for a chat and no past-conversation context is pulled in.


Chapter 10

Sources & Retrieval

Retrieval grounds answers in documents you've added, attaching source cards that show which documents and chunks were used.

The Sources section

Any answer that used sources — indexed documents, web search results, or fetched pages — shows a collapsed Sources · N row directly below it, with up to three of the source domains in the header. Click it to expand the source cards; it starts collapsed every time and there is nothing to switch on.

Search

The composer's Search toggle is on by default. Turning it off keeps that chat away from the web entirely: the model is not offered web search, image search, URL fetching, or image fetching, a link you paste is not opened, and the model is told that you turned web access off. It is a lighter privacy option than going off the record — the chat is still saved. Delegated agents and MCP tools have their own toggles and are not affected. Set whether new chats start with Search on under New chat defaults in the right settings panel. The toggle only appears when web search or web fetch is configured.

Document retrieval runs when the question looks source-worthy and the corpus has indexed documents (the configuration-level sources.mode, default auto); it has no switch in the UI.

Adding documents

Attach documents via upload or drag-and-drop. Once indexed, they become available to retrieval. Retrieval requires an embeddings endpoint to be configured (see Chapter 11).

No citations without context

If the corpus is empty, Omni Chat will not retrieve or fabricate citations. Source cards appear only when retrieval actually ran and returned context.


Chapter 11

Configuration Reference

Omni Chat is configured through environment variables set at startup, plus two optional JSON files for backends and profiles.

Environment variables — network & storage

VariableDefaultPurpose
CHAT_BIND127.0.0.1Bind address. Use 0.0.0.0 only when intentionally sharing the instance on a LAN.
CHAT_PORT00 chooses an ephemeral port and prints it in the READY line.
CHAT_DB_PATHcore/data/app.db via make devSQLite database file path.

Environment variables — backend & model

Configurable local network discovery. Native apps default to lan; standalone servers default to local. Optional "discovery": {"mode": "lan"} in endpoints.json accepts only lan, local, or off. Explicit CHAT_ENDPOINT_DISCOVERY wins over the file, then the platform default applies. Server operators edit this file or environment and restart. The admin-only Settings → AI Servers → Network discovery auto toggle maps on to lan, off to local, and applies with Save & Restart. A saved off survives until the toggle is explicitly changed. Environment overrides disable the toggle with an explanation. Saving without manual connections is valid; older clients omitting discovery preserve it. Manual connections remain usable in every mode and discovered servers remain ephemeral. The same AI Servers page also has a Scan now button: an on-demand full pass (loopback, named local hosts, and the attached network) regardless of the auto toggle, which also surfaces servers locked behind an API key as Locked · needs API key instead of hiding them. Both the auto toggle and Scan now sweep only this machine's own private-network interfaces and mDNS-named .local hosts — a server reachable only over an overlay network such as Tailscale (its 100.64.0.0/10 range isn't "private" here, and it isn't mDNS-advertised) will not appear in either; add it by its reachable address with Add server instead.

First-run network ask. A fresh desktop or iPad install does not search your network while you create the owner account, so the system's Local Network prompt doesn't appear out of nowhere. Right after setup, a card explains what the search is for and that your device will ask next. Look for servers turns on network discovery at once (the prompt follows); Not now keeps Omni Chat to this device — turn Network discovery on later in Settings → AI Servers, or press Scan now.

On iPad, Apple prompts for Local Network access on first network use. After denial, enable Omni Chat in Settings → Privacy & Security → Local Network and return to the app. Taking a photo prompts for Camera access before the camera opens. Denying it leaves the app open, with a note and a way to attach an existing file. Dictation prompts for Microphone and Speech Recognition. Those descriptions have to be in the app: iOS quits Omni Chat instead of asking when they are missing. Explicit native denial is surfaced beside the toggle; a routing error alone is not evidence of denial. Discovery starts on first active launch, cancels LAN work while inactive, and refreshes interface targets on resume. Scans retry every five seconds during the first active minute, then every 60 seconds, coalescing without overlap. Permission dialogs never block startup or the local model. Apple DNS-SD browses _ssh._tcp, _sftp-ssh._tcp, _http._tcp, and _workstation._tcp; verified address results supply names. iOS does no raw multicast or reverse-PTR and needs no restricted multicast entitlement. Non-Bonjour servers remain discoverable by IP via the bounded IPv4 subnet scan.

VariableDefaultPurpose
CHAT_ENDPOINT_URLunsetOpenAI-compatible base URL. Accepts a comma-separated list to combine backends (each shares CHAT_API_KEY). Ignored when endpoints.json is present. When unset, Omni Chat also scans for keyless local OpenAI-compatible servers and offers their models in the picker.
CHAT_ENDPOINT_DISCOVERYnative: lan; server: locallan (this machine, mDNS .local names, and each attached IPv4 /24), local (loopback and host.docker.internal), or off. Verified aliases of the same host and port are grouped when their complete model catalogs match, keeping all models together. DNS aliases such as host.docker.internal are recognized; the Docker host and container remain separate. Configured chat backends take precedence and saved interface selections continue to resolve. Matching model IDs on different machines do not cause merging; addresses without a discoverable host relationship remain separate. An auxiliary model on the same URL (for example oMLX on port 8000) does not hide the rest of that server's models. Hostname labels use Apple DNS-SD on iOS and Go mDNS reverse-PTR elsewhere. Native apps default to lan; standalone servers default to local. The environment overrides endpoints.json → discovery.mode, which overrides the platform default.
CHAT_ENDPOINT_DISCOVERY_PORTSbuilt-in popular listOptional comma-separated port list (max 32) replacing the defaults (Ollama 11434, LM Studio 1234, vLLM 8000, llama.cpp 8080, and other well-known OpenAI-compatible ports).
CHAT_API_KEYunsetOptional bearer token. Kept out of SQLite — keep secrets in env/keychain, not plaintext storage.
CHAT_MAX_TOKENS_FIELDmax_tokensSet to max_completion_tokens for endpoints/models that require it.
CHAT_CONTEXT_WINDOW8192Approximate input+output context window used for preflight budgeting.
CHAT_VISION_BUDGETbalancedAutomatic image-input budget for an environment-backed endpoint: low, balanced, or high.
CHAT_MAX_OUTPUT_TOKENSunsetOptional hard cap on generated tokens; response-length presets and model metadata are clamped to it.
CHAT_CONTEXT_SAFETY_MARGIN512Tokens reserved (subtracted) from the input budget.
CHAT_RESPONSE_HEADER_TIMEOUT2mHow long to wait for the endpoint's first response headers. On streaming backends these arrive with the first token, so this is effectively a time-to-first-token limit — raise it (e.g. 5m) for slow self-hosted routers that cold-load large models, or a turn errors with a timeout while the model is still loading. Accepts a Go duration (5m, 90s) or a bare number of seconds (300).
CHAT_STREAM_IDLE_TIMEOUT60sMaximum silence between streamed tokens before a turn is treated as a stalled upstream and ended (finish_reason=error, retryable). The streaming body has no read deadline otherwise, so a backend that wedges mid-response without closing the socket would hang the turn. A healthy stream never trips it; raise it only for extremely slow token generators. Accepts a Go duration or a bare number of seconds. An interrupted response can be resumed in place with Continue.
CHAT_STREAM_DETACH_GRACE5mHow long a generation keeps running with no browser attached before it is cancelled. Generation is detached from the connection: navigating around the app, reloading the page, or briefly closing the laptop never kills a turn — the client re-attaches and replays what it missed, and the generating chat pulses in the sidebar. A genuinely closed browser keeps the backend busy for up to this grace, so tune it to taste. Accepts a Go duration or a bare number of seconds; 0 restores the old cancel-on-disconnect behaviour. The Stop button always cancels immediately.
CHAT_HONORS_SAMPLINGfalseWhen true, Creativity sends temperature and top_p.
CHAT_EXTRA_BODYunsetJSON object merged into the top-level chat request body for static endpoint parameters.
CHAT_THINKING_ON_BODYunsetAdvanced adapter override: JSON object merged into the top-level chat request body when the per-chat Thinking switch is on/auto.
CHAT_THINKING_OFF_BODYunsetAdvanced adapter override: JSON object merged into the top-level chat request body when the per-chat Thinking switch is off.
CHAT_THINKING_BUDGET_TOKENSunsetReasoning-guard token cap per response. A reasoning model that thinks past this budget — or starts repeating itself — is stopped gracefully with a "got stuck thinking" note offering Retry without thinking and Regenerate, instead of burning the whole output window. The budget is added on top of the response-length preset rather than taken out of it, so thinking never starves the reply, and an overrun is retried once with thinking off (seeded with the model's own notes) instead of losing the turn. Unset applies the framing profile's default — 1024 for sub-2B models, 2048 up to 10B, 16384 above that and for models of unknown size. Also settable per endpoint (thinking.budget_tokens) or per model (thinking_overrides[].budget_tokens) in endpoints.json.
CHAT_PARSE_THINK_TAGSfalseAdvanced adapter override: parse streamed <think>...</think> answer text into the reasoning channel and strip it from the answer.

Reasoning models often have endpoint-specific parameters. Omni Chat infers common adapters such as Qwen served by llama.cpp and DeepSeek V4 served by ds4.c. For OMLX endpoints, Omni Chat reads the optional /v1/models/status metadata and enables the switch when thinking_default is either true or false; both values mean the model template supports enable_thinking. Custom endpoints can still define thinking request bodies in endpoints.json; the switch sends on_body for auto and off_body for off. Models with no inferred or configured adapter keep the switch disabled.

Thinking effort (Settings → Thinking effort) is an account-wide Low / Balanced / High control, default Balanced. It changes how thoroughly supporting models reason — it is not a token cap, and it is not per-chat. Qwen 3.8 thinks at its maximum (xhigh) unless this is set; Omni Chat maps Balanced to medium and High to xhigh (the template rejects the name high). The same preference is applied to agent handoff and to omni-chat launch. Models without a native effort knob are left unchanged. Operators can opt a model in with thinking_overrides[].effort: "qwen38".

ChatGPT (Codex) models reason through the Responses API's reasoning block. Omni Chat sends it for every Responses connection that has no thinking settings of its own: Thinking on asks for a reasoning summary (so the thinking streams) at the effort above — Low → low, Balanced → medium, High → high. These models have no "none" effort, so Thinking off drops to low and hides the summary.

GLM-4.5 and later (GLM-4.5, GLM-4.6, GLM-5.x) are hybrid reasoners toggled through the chat template, so a self-hosted GLM gets a fully controllable switch backed by chat_template_kwargs.enable_thinking. Served with --reasoning-parser glm45 the trace arrives on the native reasoning channel; served without it, Omni Chat sniffs the leading <think> tag and routes the trace anyway. This inference is limited to self-hosted backends on purpose — Z.ai's hosted API toggles the same models with a different shape (thinking: {"type": "enabled"}), so an endpoint pointed at it should set thinking explicitly. GLM-4 and older have no thinking mode and correctly show no switch.

Environment variables — embeddings

VariablePurpose
CHAT_EMBEDDINGS_URLEmbeddings endpoint URL.
CHAT_EMBEDDINGS_MODELEmbedding model name.
CHAT_EMBEDDINGS_DIMEmbedding dimension; must match the model.
Don't change the embedding dimension after indexing

CHAT_EMBEDDINGS_DIM must match your embedding model. Changing it after documents are indexed invalidates the stored vectors and forces a reindex.

Environment variables — auxiliary model

One optional small, fast, non-thinking model powers the background intelligence: chat-title generation, memory extraction, rolling summaries, consolidation, and — when no dedicated classifier is set — smart-router classification. It runs off the response path and never uses the main chat model. (Rolling chat summaries are the one exception when no auxiliary model is configured: they then fall back to the chat's own model between turns, so long-chat compaction always works.) A 1–2B non-reasoning instruct model is ideal — the recommended default is LFM2.5-1.2B-Instruct (Liquid AI, ~731 MB at Q4_K_M), served with llama-server -hf LiquidAI/LFM2.5-1.2B-Instruct-GGUF:Q4_K_M; see .env.example for smaller and larger alternatives.

VariableDefaultPurpose
CHAT_AUX_URLunsetOpenAI-compatible chat endpoint for the auxiliary model. Unset disables every aux-powered feature.
CHAT_AUX_MODELunsetModel name (small, fast, non-thinking). Required when the URL is set.
CHAT_AUX_API_KEYunsetOptional bearer token for the endpoint.
CHAT_AUX_ROUTERtrueSet false/0 to stop the aux model from serving as the smart-router classifier.
CHAT_AUX_SHOW_IN_MODEL_DROPDOWNtrueSet false/0 to hide the aux model from manual chat selection while keeping its background jobs enabled.
Migration from CHAT_MEMORY_*

The auxiliary model replaces the former CHAT_MEMORY_URL/CHAT_MEMORY_MODEL/CHAT_MEMORY_API_KEY variables, which are no longer read. Point CHAT_AUX_URL/CHAT_AUX_MODEL at the same small model to restore the memory jobs; setting CHAT_MEMORY_* now only logs a startup warning.

Environment variables — config locations

VariableDefaultPurpose
CHAT_CONFIG_DIROS config dirApp config directory. Defaults to ~/.config/omni-chat on macOS/Linux (unless Linux XDG_CONFIG_HOME is set); Windows uses the user config directory.
CHAT_PROFILES_PATH<config-dir>/profiles.jsonOverride only the profiles file.
CHAT_SOUL_PATH<config-dir>/SOUL.mdOverride only the assistant personality Markdown file.
CHAT_SKILLS_PATH<config-dir>/skillsOverride the overlay directory of *.md skills. Missing default dir is fine. Starter: docs/skills.example/.
CHAT_ENDPOINTS_PATH<config-dir>/endpoints.jsonOverride only the multi-backend file.
CHAT_ROUTER_PATH<config-dir>/router.jsonOverride only the Auto (Smart Router) policy file.
CHAT_AGENTS_PATH<config-dir>/agents.jsonOverride only the external-agent file that enables the Agent handoff (delegate_task) tool.
CHAT_MCP_PATH<config-dir>/mcp.jsonOverride only the MCP-server file that enables MCP tools (see the README's "MCP tools" section and docs/mcp.example.json).
CHAT_OAUTH_CALLBACK_BIND127.0.0.1Bind host for the temporary OAuth callback listener used by subscription sign-ins (port 1455 for Codex). Containers set 0.0.0.0 so the published port reaches it — the bundled compose.yaml does this. Publishing the host side is opt-in: set OMNI_OAUTH_PUBLISH_PORT=1455 in .env (default: an ephemeral port, so the automatic callback is off and sign-in uses the paste fallback).

Environment variables — semantic router classifier

VariableDefaultPurpose
CHAT_ROUTER_CLASSIFIER_URLunsetOpenAI-compatible endpoint of a dedicated router-classifier model (e.g. Arch-Router). Unset falls back to the auxiliary model (CHAT_AUX_*) when configured, else heuristics only.
CHAT_ROUTER_CLASSIFIER_MODELunsetRequired when the URL is set; the one explicit classifier model id. The router overlay defaults it to arch-router-1.5b.
CHAT_ROUTER_CLASSIFIER_FORMATautoClassifier protocol adapter: auto (detect arch-router in the model id), arch (Arch-Router native routes), or schema (JSON-schema classification for general instruct models).
CHAT_ROUTER_CLASSIFIER_API_KEYunsetOptional bearer token for the classifier endpoint.
CHAT_ROUTER_CLASSIFIER_TIMEOUT3000msRace deadline; accepts 100ms–10s. Requires a Go duration unit, e.g. 3000ms — a bare number like 3000 fails to parse and is fatal at startup.

Environment variables — web search

VariableDefaultPurpose
CHAT_WEB_SEARCH_PROVIDERsearxngProvider implementation.
CHAT_WEB_SEARCH_URLunsetProvider base URL (SearXNG); setting it enables the model tool. Ignored when endpoints.json has a search section.
CHAT_WEB_SEARCH_API_KEYunsetKey for a hosted provider (tavily); setting it also enables the tool, since a hosted provider needs no URL.
CHAT_WEB_SEARCH_TIMEOUT10sPer-search timeout.
CHAT_WEB_SEARCH_MAX_RESULTS5Result limit from 1 to 10.
CHAT_TOOL_CALLINGtrueSet false for an environment-backed endpoint without tool support.
CHAT_WEB_FETCHfalseSet true to enable the fetch_url tool.
CHAT_WEB_FETCH_PROVIDERdirectdirect (in-core, SSRF-guarded) or reader (external reader endpoint).
CHAT_WEB_FETCH_URLunsetReader endpoint base URL; required when the provider is reader.
CHAT_WEB_FETCH_TIMEOUT15sPer-fetch deadline, 1s–60s.
CHAT_WEB_FETCH_MAX_BYTES2097152Per-page download cap before truncation.
CHAT_WEB_FETCH_USER_AGENTbrowser UAOverrides the User-Agent used by the direct provider.
CHAT_WEB_FETCH_ALLOW_PRIVATEfalseLocal dev only: relax the SSRF guard.
CHAT_IMAGE_GEN_URLunsetImage backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables the generate_image tool. Ignored when endpoints.json declares images[].
CHAT_IMAGE_GEN_PROVIDERopenaiBackend protocol: openai (POST /v1/images/generations) or responses (the Responses API's image_generation tool; needs carrier_model, so use endpoints.json images[] for it in practice).
CHAT_IMAGE_GEN_MODELunsetOptional model id passed through to the backend (the bundled sidecar ignores it).
CHAT_IMAGE_GEN_API_KEYunsetOptional bearer token for the image backend; env-only, never stored in SQLite.
CHAT_IMAGE_GEN_TIMEOUT120sIdle deadline per image, 10s–10m (duration or bare seconds). Streamed progress events reset it; against a non-streaming backend it is the total deadline.
CHAT_IMAGE_GEN_STREAMtrueRequest streamed progress and partial previews (OpenAI Images stream/partial_images); JSON-only backends degrade gracefully.
CHAT_TTS_URLunsetSpeech backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables read-aloud. Ignored when endpoints.json declares tts[].
CHAT_TTS_MODELtts-1Speech model id passed through to the backend (omnivoice for the bundled sidecar, gpt-4o-mini-tts on the OpenAI API).
CHAT_TTS_API_KEYunsetOptional bearer token for the speech backend. Env-only, never stored in SQLite.
CHAT_TTS_VOICESunsetOptional comma-separated voice list, pinning the voice picker instead of asking the backend.
CHAT_TTS_DEFAULT_VOICEunsetVoice used when an account has no saved preference.
CHAT_TTS_TIMEOUT300sPer-synthesis-call deadline (1s–10m). Long replies are chunked on sentence boundaries, so this bounds each chunk; the default is sized for CPU backends, where synthesis can take minutes.
CHAT_STT_URLunsetTranscription backend base URL — a single URL or a comma-separated list (one backend per URL, first = default). Setting it enables voice input's server engine. Ignored when endpoints.json declares stt[].
CHAT_STT_MODELwhisper-1Transcription model id passed through to the backend (large-v3-turbo for the bundled sidecar, whisper-1 on the OpenAI API).
CHAT_STT_API_KEYunsetOptional bearer token for the transcription backend. Env-only, never stored in SQLite.
CHAT_STT_TIMEOUT120sPer-transcription-call deadline (1s–10m).

Search uses structured function calls and persists real results as cards. Snippets are untrusted summaries, not full-page content. The per-chat Search toggle gates them: off withholds web search, image search, and both fetch tools for that chat. Saved cards appear in the collapsed Sources section under each answer. Each assistant response records whether web search was unavailable, available but unused, successful, successful with no results, or failed; the Generation details panel reports that outcome.

With CHAT_WEB_FETCH=true a companion fetch_url tool lets the model retrieve one specific URL — a web page, article, or (via a capable reader) a YouTube transcript — extract its readable text, and summarize or quote it. The direct provider fetches in-core behind an SSRF guard that refuses private, loopback, link-local, and cloud-metadata addresses; the reader provider delegates to an external or self-hosted reader endpoint for JS-heavy or anti-bot sites. Fetched content is treated as untrusted, persists as a source card in the same Sources section as web-search cards, and its outcome is folded into the single Web access line of the Generation details panel alongside web search (that line reads Used whenever either tool ran). For time-sensitive questions the model can limit a search to the past day, week, month or year; with Tavily, a past-day or past-week search uses its news index, so asking for today's headlines returns that day's articles rather than older roundups. web_search and fetch_url share one bounded tool-call budget per turn. When the fetch tool is enabled and your message contains a URL, the core fetches it automatically before the model runs, so summarizing a pasted link works even if the model would otherwise try to search for it.

The same setting lets the assistant show you a picture that already exists rather than drawing one — ask "show me a 3-month candlestick chart of TSLA" or "find me a picture of a cat and dog" and it looks one up, downloads it, and puts it in its reply. This works even on an instance with no image generation configured. (Looking pictures up by keyword needs a search provider that supports image search; without one the assistant can still show images whose address it already has, such as a chart endpoint or an image on a page it just read.) The image is saved with the message rather than hotlinked, so it stays put if the source disappears, comes along when you export the chat, and never tells the origin site who is looking at it; you can also ask for edits to it afterwards. A fetched image is always labelled Fetched with a link to where it came from, so it is never confused with one the assistant generated. If a URL turns out to be a web page rather than an image, the assistant says so instead of showing something broken. It is unavailable in off-the-record chats, since the image would have to be stored.

Every turn also carries a hidden, system-injected context block (separate from editable Session Instructions): the current date and time in your browser's timezone, best-effort details about the client device's OS/version/CPU architecture, and a short directive. When the web_search tool is available the directive tells the model to search for current or possibly-changed facts and to trust fresh results over its own memory; the fetch_url tool adds guidance to read a specific URL for summarizing. When neither web tool is available it instead tells the model it has no web access and should flag that its knowledge may be out of date. The browser sends its IANA timezone name plus structured OS metadata from Client Hints or coarse user-agent fallbacks; the raw user-agent string is never sent. The server keeps the clock, treats device version/architecture as approximate and distinct from a remote core host, and persists neither value. The prompt explicitly tells the model to use the detected OS for commands, installation, and debugging unless you name a different target. No configuration is required.

With an image backend configured, tool-capable models are offered generate_image and edit_image tools (no per-chat toggle). The core calls the backend server-side, stores the produced image with the conversation (bytes never enter the model context), and streams it to the chat as it lands — including live progress and blurred partial previews when the backend supports streamed partials (the bundled sidecar and the Codex backend both do; those progress events also reset the CHAT_IMAGE_GEN_TIMEOUT idle deadline, so long generations don't time out while they're visibly working). The tool has its own budget of two images per response, separate from the web tools; the assistant response's Generation details panel records an Image generation outcome (used, failed, or not used), and each image's caption tooltip names the backend that produced it.

Backends are named and plural, and reference a server. The simplest setup is CHAT_IMAGE_GEN_URL — a single URL or a comma-separated list (one backend per URL, named by host, first = default). For richer setups declare an images[] array + default_image_endpoint in endpoints.json (see docs/endpoints.example.jsonc): each entry names a connection — which carries the base_url, a key (api_key/api_key_env) or an auth OAuth block, and (via protocol) whether it speaks the OpenAI Images API (POST /v1/images/generations with b64_json; the bundled Z-Image-Turbo sidecar, compose.imagegen.yaml, GPU required — see Running Omni Chat, combination g) or the OpenAI Responses API's native image_generation tool (protocol: "responses") — and keeps optional model, timeout, and stream fields. The responses protocol is how a ChatGPT subscription generates images (gpt-image-2): point the connection at https://chatgpt.com/backend-api/codex with "auth": {"type":"oauth","provider":"openai-codex"} and "protocol": "responses" — the same instance credential as a Codex text connection, connected once by the owner, shared automatically when it's the same connection — plus "model": "gpt-image-2" and a carrier_model (the text model slug that carries the tool call, e.g. gpt-5.5). Each chat picks its backend in the settings sidebar's Image generation select (empty = instance default; the choice is saved with the chat), and the New chat defaults section sets the backend new chats start on. Each backend appears separately in the service-availability panel as Image generation: <name>.

Speech backends work the same way. The simplest setup is CHAT_TTS_URL (single URL or comma-separated list); for richer setups declare a tts[] array + default_tts_endpoint in endpoints.json (see docs/endpoints.example.jsonc): each entry names a connection (any server speaking the OpenAI speech API, POST /v1/audio/speech — the bundled OmniVoice sidecar, compose.tts.yaml, or a hosted endpoint such as https://api.openai.com) and keeps a required model, an optional pinned voices list with a default_voice, and a timeout. Users pick their voice — across every configured backend — under Voice Engine in the settings sidebar's Voice & Dictation section (see Read aloud). Each backend appears in the service-availability panel under Speech.

Transcription backends work the same way. The simplest setup is CHAT_STT_URL (single URL or comma-separated list); for richer setups declare an stt[] array + default_stt_endpoint in endpoints.json (see docs/endpoints.example.jsonc): each entry names a connection (any server speaking the OpenAI transcription API, POST /v1/audio/transcriptions — the bundled whisper sidecar, compose.stt.yaml, or a hosted endpoint such as https://api.openai.com) and keeps a required model, an optional default language, and a timeout. Users pick their engine and dictation language under Dictation in the settings sidebar's Voice & Dictation section (see Voice input). Unlike read-aloud, nothing configured does not hide dictation — the browser fallback still covers it where supported. Each backend appears in the service-availability panel as Dictation: <name>.

Multiple backends — endpoints.json

Create <config-dir>/endpoints.json to combine backends. Servers are described once in a top-level connections[] array (name, label, base_url, a key inline as api_key or via api_key_env, and optional auth/protocol); the endpoints[] array then names one connection per chat backend and keeps only the capability fields below. schema_version must be 2 — an older per-endpoint-URL file fails startup with a message pointing at docs/endpoints.example.jsonc and this chapter; there is no automatic migration, so an existing file from before this change must be recreated. Per-connection/per-endpoint capability fields mirror the CHAT_* model settings and apply only to that backend, including optional thinking adapter overrides and model-pattern overrides.

{
  "schema_version": 2,
  "connections": [
    { "name": "local-litellm", "label": "Local LiteLLM", "base_url": "http://localhost:4000" },
    {
      "name": "openai",
      "label": "OpenAI",
      "base_url": "https://api.openai.com",
      "api_key_env": "OPENAI_API_KEY"
    },
    {
      "name": "vllm-local",
      "label": "vLLM (local)",
      "base_url": "http://localhost:8000",
      "api_key": "sk-local-only-not-a-real-secret"
    }
  ],
  "default_endpoint": "local-litellm",
  "endpoints": [
    {
      "connection": "local-litellm",
      "thinking": {
        "on_body": { "reasoning_effort": "medium" },
        "off_body": { "reasoning_effort": "none" },
        "parse_think_tags": true
      }
    },
    {
      "connection": "openai",
      "max_tokens_field": "max_completion_tokens",
      "context_window": 128000,
      "max_output_tokens": 16000,
      "honors_sampling": true
    },
    {
      "connection": "vllm-local",
      "context_window": 32768,
      "prompt_profile": "auto",
      "model_overrides": [
        { "model_pattern": "qwen2.5-vl*", "vision": true, "tool_calling": true, "params_b": 7 },
        { "model_pattern": "lfm2-*", "params_b": 1.2 },
        { "model_pattern": "qwen3-30b-a3b*", "active_params_b": 3, "capability": "large" },
        { "model_pattern": "phi-4-mini*", "prompt_profile": "full" }
      ]
    }
  ]
}
FieldMeaning
connectionNames a connections[] entry (server, credential, protocol). At most one endpoints[] entry per connection.
connections[].name / .labelStable id ([a-z0-9-]+, immutable once saved) other sections reference, and the display name / group label in the model picker — safe to rename freely.
connections[].base_urlServer root supporting /v1/models, /v1/chat/completions, /v1/embeddings.
connections[].api_key / .api_key_envInline secret, or the name of an env var to read it from.
connections[].runtimeOptional server type: llamacpp, vllm, ollama, lmstudio, omlx, litellm, sglang, openai or other. Filled in by Scan now or Test connection and editable in Settings (Server type); shown as the server's tile. Informational today — it gives backend-specific handling a stable id to key on.
max_tokens_fieldOutput-tokens field name (default max_tokens).
context_windowContext size for budgeting (default 8192).
max_output_tokensPer-endpoint output cap; clamps the response-length setting.
honors_samplingWhether the backend respects temperature/top_p (default false).
extra_bodyJSON object of static top-level request params for this backend.
sampling_defaultsGates the built-in anti-loop sampling defaults (hybrid Qwen3.x, Gemma, gpt-oss, DeepSeek-R1, Nemotron, GLM-4.5+ model-card values): "auto" (default) applies them when the backend is recognized as self-hosted from the model's owned_by (llama.cpp, SGLang, oMLX, Ollama, LM Studio, vLLM), "on" forces them (for servers that omit owned_by, e.g. mlx_lm.server), "off" disables them. Qwen3.5/3.6 and Qwen3.8 have generation-specific thinking/non-thinking profiles. Endpoint extra_body and per-model model_overrides[].extra_body override individual keys.
prompt_profileHow much tool surface models on this backend are given: "auto" (default) derives it from the resolved parameter size, "full" keeps the complete surface, "terse"/"minimal" force the small-model behavior. Also settable per model via model_overrides[].prompt_profile, which wins. See Small models on modest hardware.
vision_budgetAutomatic image-input ceiling: low (1024 px, 1 MP/image, 2 MP/turn), balanced (default: 1568 px, 2.5 MP/image, 5 MP/turn), or high (2560 px, 6 MP/image, 12 MP/turn). Also settable per model via model_overrides[].vision_budget.
thinkingOptional adapter mapping for the per-chat Thinking switch: on_body, off_body, optional parse_think_tags, and optional budget_tokens — the reasoning-guard cap for this backend, added on top of the answer budget (default: the framing profile's, 1024/2048/16384).
thinking_overridesModel-pattern overrides for on_body, off_body, parse_think_tags, and budget_tokens. Bodies may contain mode-specific sampling parameters as well as custom template kwargs; partial nested bodies merge over the inferred adapter, and the first matching pattern wins.
tool_callingOptional boolean; set false when this endpoint does not support streamed tool calls (a hard cap on every model here).
model_overridesPer-model capability metadata, matched by a model_pattern glob on the model id: vision, tool_calling, params_b, vendor, family, prompt_profile. params_b is also read from the backend when it reports one (llama.cpp's meta.n_params) — the only size signal for an alias like lfm2-tool whose id carries no size — and an override here still wins. Wins over the built-in model catalog and id inference. An entry may also carry extra_body — request params for matching models only (e.g. a custom top_k for one family), overriding the endpoint-wide extra_body and the inferred sampling defaults per key.
Load order & secrets

If endpoints.json exists, CHAT_ENDPOINT_URL is ignored. Connection keys in this file are never written to SQLite, but this owner-managed file is the one place a plaintext key may live — prefer api_key_env if you'd rather not store the secret in the file. A ready-to-copy, annotated starter template is in docs/endpoints.example.jsonc — the app never loads the .jsonc itself (it's plain JSON with // comments); copy the parts you need into your real endpoints.json and delete the comments, or build the file from the desktop app's AI Servers page instead.

Assistant soul — SOUL.md

The built-in assistant soul makes Omni Chat warm, thoughtful, conversational, and lightly playful by default. To customize it, create <config-dir>/SOUL.md or point CHAT_SOUL_PATH at a Markdown file. A missing default file is fine and uses the built-in soul; an explicit path must exist and must not be empty. SOUL is personality text, not a safety boundary: the app safety invariants are still hard-coded and injected before the single upstream system message is sent.

Custom profiles — profiles.json

Extend or override the built-in profiles by creating <config-dir>/profiles.json.

{
  "schema_version": 1,
  "profiles": [
    {
      "id": "review_code",
      "label": "Review Code",
      "description": "Find correctness, maintainability, security, and test gaps.",
      "persona_prompt": "Review code for correctness, maintainability, edge cases, security and privacy risks, and missing tests. Prioritize concrete findings...",
      "icon": "code",
      "style": {
        "tone": "direct",
        "response_length": "detailed",
        "creativity": "precise",
        "format": "structured"
      }
    }
  ]
}
FieldMeaning
idStable identifier in lower snake_case.
label / descriptionCard title and one-line summary.
persona_promptThe persona text added to editable session instructions.
iconOptional; must name a bundled icon in lower kebab-case. Unknown names fall back to a generic icon.
skillOptional skill id. Empty means “if a skill has this profile’s id, attach it.” An explicit id that is not in the skills catalog fails startup.
styleA bundle of tone, response_length, creativity, format using valid option keys.
Override & append rules

Built-ins load first; a matching id overrides in place, and new IDs append. The welcome screen shows at most 8 cards total, with Auto first when present; the dropdowns still list every profile. A missing default profiles.json is fine and uses built-ins; an explicit CHAT_PROFILES_PATH must exist and is validated at startup. A starter file is in docs/profiles.example.json; try it with CHAT_PROFILES_PATH=docs/profiles.example.json make dev.

Skills — skills/*.md

A profile is who the assistant is this session (persona + style). A skill is a procedure it should follow. Skills are Markdown files, injected as a managed Skill: block that is not shown in the session-instructions textarea. The resolved skill for the selected profile is separately visible as read-only Markdown in the expanded Profile controls. Keep persona_prompt short; put process in a skill.

A profile attaches a skill in two ways: an explicit "skill": "write_howto" field, or the same id (a profile id: "review_code" loads a skill named review_code if one exists). An explicit skill id that is missing from the catalog fails startup before READY.

Built-in skills match the six welcome-card profiles: brainstorm_ideas, draft_content, write_code, learn_topic, write_howto, and write_minibook. Overlay more as *.md under <config-dir>/skills/, or point CHAT_SKILLS_PATH at a directory. A missing default directory is fine; an explicit path must exist. Each file is at most 8 KiB. The same id replaces a bundled skill. Do not put a README.md in that directory — every .md file is loaded as a skill.

---
id: review_code
name: Review Code
---

Review a code change for correctness, tests, and safety. You are a reviewer, not the author.

A starter overlay is docs/skills.example/: same-id files for docs/profiles.example.json (review_code, debug_issue, plan_project, and the rest). Those example profiles leave skill empty so profiles-only startup still succeeds. Copy both:

cp docs/profiles.example.json ~/.config/omni-chat/profiles.json
mkdir -p ~/.config/omni-chat/skills
cp docs/skills.example/*.md ~/.config/omni-chat/skills/

Or try both for one run: CHAT_PROFILES_PATH=docs/profiles.example.json CHAT_SKILLS_PATH=docs/skills.example make dev. Typst templates for how-tos and mini-books seed only when the resolved skill is write_howto or write_minibook; a custom skill gets the prompt, not those templates.

Custom router policy — router.json

Customize how the Auto (Smart Router) model tags and picks models by creating <config-dir>/router.json.

{
  "schema_version": 1,
  "routes": [
    { "name": "code", "description": "writing, debugging, or reviewing code or technical artifacts" }
  ],
  "tag_rules": [
    { "pattern": "*coder*", "tags": ["coder", "code"] },
    { "pattern": "*mini*", "endpoint": "openai", "tags": ["small", "fast"] }
  ],
  "policies": [
    {
      "task": "code",
      "prefer": [
        { "tags": ["coder"], "min_context": 32768 },
        { "tags": ["general"] }
      ]
    }
  ]
}
FieldMeaning
schema_versionMust be 1.
classifierOptional structured section (url, model, format, timeout_ms) configuring the semantic classifier described below. classifier_model is a deprecated top-level alias for classifier.model.
classifier.formatAdapter selector: auto (detect arch-router in the model id), arch (Arch-Router native routes), or schema (JSON-schema classification for general instruct models). Env equivalent: CHAT_ROUTER_CLASSIFIER_FORMAT.
routes[]Optionally overrides a built-in task's description shown to the classifier. Each entry is { name, description }, where name must be one of the task classes (code | reasoning | simple_fast | creative | general). Providing any routes[] replaces the full built-in list, not just the named entries.
tag_rules[]pattern (glob over the model id), optional endpoint (glob scope), and tags (lower snake_case). A model's tags are the union of every matching rule, built-ins included.
policies[]task (code | reasoning | simple_fast | creative | general) and an ordered prefer list of selectors (tags and/or min_context). A user policy replaces the built-in policy for the same task.
Load order & validation

A missing default router.json is fine and uses built-in tag rules and policies — Auto still works. An explicit CHAT_ROUTER_PATH must exist and is validated at startup (invalid JSON, unknown fields, a bad schema_version, empty patterns, and invalid or duplicate policy tasks all fail before the READY handshake). A starter file is in docs/router.example.json; try it with CHAT_ROUTER_PATH=docs/router.example.json make dev.

How Auto tags and routes models

Every available model gets a set of lower-snake_case tags before Auto picks one for a turn. Tags come from three layers, and later layers only add tags — none of them remove a tag a candidate already has:

  1. Built-in glob rules match substrings of the model id: *coder* / *codestral* / *devstral* / … → coder; *deepseek-r1* / *qwq* / *reason* / *o1* / *o3* / *gpt-oss* → reasoning; *mini* / *small* / *phi* / *1b*…*8b* → fast; *70b* / *large* / *opus* / *gpt-4* / *gpt-5* → large.
  2. Computed tags are attached automatically, not from glob rules: general (every candidate), large_context (context window ≥ 64K), and capability tags vision, reasoning, size_tiny / size_small / size_medium / size_large (size_tiny is additive — a sub-2B model carries size_small too — and the small and large ends also alias to fast / large), and vendor_<v> / family_<f>. These resolve strongest-first from a per-model model_overrides entry in endpoints.json, a built-in catalog of common model ids, then inference from the id itself.
  3. Your own tag_rules[] (above) add more tags the same way the built-ins do — a model's final tag set is the union of every matching rule.

For each turn, Auto classifies a task (code, reasoning, simple_fast, creative, or general), then narrows the candidates: models that can't fit the estimated input are dropped (unless that would empty the list), and an image attachment keeps only vision-tagged models when one exists. It keeps loaded models first (see Auto), then matches capability to difficulty: every model has a capability band — tiny, small, medium, large or frontier — from a capability override, its parameter count (tiny < 2B ≤ small < 10B ≤ medium < 40B ≤ large), the built-in catalog for hosted models, or a guess (large for a cloud API, medium for your own server), one step higher for a true reasoning model (R1, QwQ, gpt-oss — not one that can merely switch thinking on) and for a coder on a coding turn. A simple turn needs small+, a moderate one medium+, a hard one large+ (Jev's five-level scale adds tiny and frontier); Auto keeps the models that reach the band and prefers the smallest of them — the fastest model that is good enough — or, if none does, the strongest. It then applies soft boosts (large_context when the turn needs it, and tool-calling/thinking-capable models per your toggles) before walking the task's policies[].prefer list top to bottom and picking the first selector that still matches at least one candidate. If nothing in prefer ever matches, the default endpoint's first remaining candidate wins.

A user policy for a task replaces the built-in policy for that task outright rather than merging with it, so writing your own prefer list is the reliable way to pin a task to a specific model — for example, to keep creative turns on one backend model even though a built-in glob like *large* or *opus* would otherwise tag it (and win) differently:

{
  "schema_version": 1,
  "tag_rules": [
    { "pattern": "*my-writer-model*", "tags": ["storyteller"] }
  ],
  "policies": [
    {
      "task": "creative",
      "prefer": [
        { "tags": ["storyteller"] },
        { "tags": ["general"] }
      ]
    }
  ]
}

This tags one specific backend model storyteller and makes it the first choice for creative turns, falling back to any available model (general) if that one is offline. See docs/router.example.json for a fuller starter covering the classifier, routes, and multiple policies together, and docs/endpoints.example.jsonc for the model_overrides syntax used to set capability tags explicitly per backend.

Semantic router classifier

By default Auto classifies each turn with fast, deterministic heuristics only. To classify turns semantically instead, either configure the auxiliary model (CHAT_AUX_*) — which automatically doubles as the classifier via the JSON-schema adapter with no extra setup (disable with CHAT_AUX_ROUTER=0) — or point the router at a dedicated classifier: set CHAT_ROUTER_CLASSIFIER_URL and CHAT_ROUTER_CLASSIFIER_MODEL (and optionally CHAT_ROUTER_CLASSIFIER_API_KEY / CHAT_ROUTER_CLASSIFIER_TIMEOUT), or add an equivalent classifier section to router.json (see the commented example in docs/router.example.json) — env vars override the file per field. A dedicated classifier always takes precedence over the auxiliary model.

The default classifier model is Arch-Router-1.5B, a purpose-built router model. The core auto-detects it and speaks its native protocol: it is handed the task classes (with descriptions overridable via routes above) and returns the single best-fitting one, which maps to a model through the usual tag rules and policies. classifier.format / CHAT_ROUTER_CLASSIFIER_FORMAT forces the adapter — set schema to use a general instruct model with the richer JSON-schema classification instead. Because Arch-Router returns only a route, under it (or with the classifier disabled) the Auto-profile persona is chosen by a deterministic task-affinity heuristic over the profile catalog rather than by the classifier; the schema classifier gives a richer, content-aware persona pick.

On every routed turn the classifier races the heuristics against a deadline (default 3000ms, sized for a CPU-only sidecar classifying a fresh, full-context turn). Whichever finishes first within the deadline wins; a classifier error, timeout, or unset config falls back to heuristics, so a turn never fails or blocks on it. The persisted routing reason records which one won, prefixed semantic: or heuristic: , and the per-message stats hover's decision time includes any time spent waiting on the classifier.

To self-host a tiny classifier model with no extra setup, use the compose.router.yaml (or compose.router-only.yaml for host-run cores) Compose overlay described under Docker Compose deployment; it runs ghcr.io/ggml-org/llama.cpp with the Arch-Router-1.5B GGUF and constrained JSON output.

Third-party model license

Arch-Router-1.5B is a Katanemo model published by DigitalOcean, LLC under the Katanemo Community License, not Omni Chat's own Apache-2.0 license. Built with DigitalOcean. Commercial use of Arch-Router-1.5B requires a separate license from DigitalOcean — see the model's license terms on Hugging Face before using it commercially. Katanemo Models are licensed under the DigitalOcean Community License, Copyright 2026 DigitalOcean, LLC. All Rights Reserved.


Chapter 12

Administration & Deployment

Multi-user isolation

Every owned resource carries an owner, and the store layer filters all queries by the session's user. Identity comes from the session, never the client. Accessing another user's private resource returns 404 — the app reveals nothing about resources you don't own.

Projects have a visibility of private or shared. Marking a project shared makes its chats, documents, and memory readable and writable by members of the group; private projects stay yours alone.

Deployment modes

Shared web instance

Run with CHAT_BIND=0.0.0.0 so household or team members can reach it in a browser and log in. Each gets an isolated workspace; shared projects are the common ground.

Single-machine desktop

The Tauri shell bundles the core as a sidecar on 127.0.0.1 and opens it in a native window. Auto-login for solo use is not built yet, so the first launch shows the setup screen and later launches ask you to log in.

Building the desktop app (macOS)

The shell picks a free loopback port, starts the same core binary as a sidecar with that port, waits for the core's READY port=<n> handshake, then points the webview at it. All logic stays in the core — the Rust layer only supervises the process.

The app is built for Apple Silicon only and declares a minimum system version of 26.0, so the DMG will not install or launch below macOS 26. make build-desktop runs make check-macos-prereqs first, which fails on a non-arm64 host or any DESKTOP_TARGET other than aarch64-apple-darwin.

On top of the usual toolchain you need Rust 1.82+ and the Tauri CLI:

cargo install tauri-cli --version "^2" --locked

make dev-desktop     # run it
make build-desktop   # bundle it

For a locally signed build, create an Apple Development certificate in Xcode, find its exact name with security find-identity -v -p codesigning, then run:

APPLE_SIGNING_IDENTITY='Apple Development: Your Name (IDENTIFIER)' make build-desktop

The macOS signing config supplies the executable-memory entitlement required by the core's SQLite/WebAssembly runtime, while keeping Hardened Runtime enabled. A signed build missing this entitlement can stop with signal 9 before printing anything (CODESIGNING: Invalid Page in the macOS crash report). Rebuild with desktop/Entitlements.macos.plist configured; restoring endpoint settings does not fix this packaging error.

The bundle lands at desktop/target/release/bundle/macos/Omni Chat.app. make build-desktop also builds the user-facing omni-chat CLI into Contents/Resources/omni-chat (not a Tauri sidecar — those get a target-triple suffix). The one-line installer then symlinks that binary onto PATH.

Both targets first run make llama-sidecar, which downloads the llama.cpp release pinned in desktop/llama.json, verifies its SHA256 and keeps llama-server plus the libraries it actually links in desktop/vendor/llama/ (git-ignored, ~23 MB). It needs the network once, then no-ops until the pinned tag changes — the same discipline web/runtimes.json uses for the Pyodide assets. To bump the version, edit the tag, run the target, and paste the checksum it reports.

WhatPath
Config~/.config/omni-chat/ — the same directory the core uses in web mode, so one machine has one config.
Database~/Library/Application Support/omni-chat/app.db — separate from the repo's data/app.db, so the desktop app has its own accounts and chats.
Signing and distribution

Without a signing identity, builds use ad-hoc signing. An Apple Development certificate supports local development; Developer ID signing and notarization are needed for normal distribution outside the App Store. The build accepts Tauri's APPLE_SIGNING_IDENTITY setting; no identity is stored in the repo.

If a macOS build finishes compilation but fails at bundle_dmg.sh, the verbose build output identifies the failing disk-image command. Retry packaging with cargo tauri bundle --bundles app,dmg --verbose --config from desktop/, supplying the same version configuration and signing identity as the original build (see the README example). For Finder/AppleScript errors only, CI=true skips window layout while retaining the Applications shortcut.

Linux AppImage

Linux desktop builds use the same thin Tauri supervisor and Go core as macOS. Prebuilt AppImages are published for x86_64 and aarch64. The one file carries both the GUI and the omni-chat launch CLI; the installer gives that file two names using AppImage's multi-call entrypoint.

On first use, the shell materializes its embedded Go core and CLI as private, content-addressed executables under the user's XDG cache. They still come from the one downloaded and verified AppImage; no system-wide runtime is installed.

curl -fsSL https://geekaholic.gitlab.io/omni-chat/install-linux.sh | bash
omni-chat-app
omni-chat launch pi

Immutable, versioned downloads are available from the project's GitLab package registry. Make the matching Omni-Chat-x86_64.AppImage or Omni-Chat-aarch64.AppImage executable and run it. Endpoint keys use the desktop Secret Service (GNOME Keyring, KWallet or compatible) and are never written to SQLite or plaintext config.

Release builds use Ubuntu 24.04 target-architecture tools, once per architecture with the same explicit BUILD_VERSION. Published compatibility starts at Ubuntu 24.04 and is smoke-tested on Debian 13, current Fedora, and a current rolling distribution. Linux deliberately does not bundle llama-server; the core uses an installed binary or Docker so the runtime can match the host's GPU drivers.

make check-linux-prereqs
make build-linux BUILD_VERSION=2026.9.1120000

GITLAB_DEPLOY_TOKEN=your-deploy-token make publish-linux \
  APPIMAGE_X86_64='/path/Omni Chat_2026.9.1120000_amd64.AppImage' \
  APPIMAGE_AARCH64='/path/Omni Chat_2026.9.1120000_aarch64.AppImage'

Create a project deploy token with the write_package_registry scope and provide it as GITLAB_DEPLOY_TOKEN only for the publish command. A personal, project or group access token with api also works as GITLAB_TOKEN; GitLab CI uses its automatic CI_JOB_TOKEN. AppImages and their checksums go to GitLab's generic package registry; GitLab Pages stores only the small downloads/latest.json manifest used by the installer.

On other Linux distributions, including Omarchy, the bundled Ubuntu 24.04 Docker builder supplies Go, Node, Rust, Tauri and the native libraries. Only Make and a working local Docker daemon accessible by your user are required:

make build-linux-docker
# Optional: pin the release version.
make build-linux-docker BUILD_VERSION=2026.9.1120000

The target checks Docker access, builds/reuses docker/build-linux/Dockerfile, and runs the build stages as your UID/GID without a privileged container or FUSE device. The first run downloads and compiles the tools; later runs reuse layers and caches. AppImages are written to .cache/linux-docker/amd64/target/release/bundle/appimage/ (use arm64 for an ARM64 target). Rust and npm caches are separate from your host's desktop/target and web/node_modules; Go sidecars and web assets still use their normal checkout paths, so run only one build at a time in a checkout. The default target is the host's native architecture. If Docker access fails, start the daemon and check your Docker context and socket permissions before retrying.

To build x86_64 on ARM64 or ARM64 on x86_64, set LINUX_ARCH=amd64 or LINUX_ARCH=arm64 (aliases x86_64/aarch64 also work). The frontend builds in native Node/Go Docker userspace, which also cross-compiles the pure-Go sidecars for the selected architecture. Rust and linuxdeploy's ELF dependency step run in target Ubuntu userspace under QEMU when needed. The native stage then creates the target AppImage from the completed AppDir and target runtime, avoiding the target AppImage plugin on large-page ARM64 hosts. npm caches follow the host architecture in .cache/linux-docker/assets-<host-arch>/; Rust caches and AppImages follow the target architecture. The wrapper checks target execution before compiling. Ubuntu's versioned Rust 1.91 packages avoid the rustup compiler's allocator crash; a Rust compile/link/run smoke test runs before installing Tauri. No special QEMU or Node environment settings are required. If QEMU/binfmt is missing, enable the host executable handlers once:

docker run --privileged --rm tonistiigi/binfmt --install amd64,arm64

This setup changes host executable handlers and is never run automatically by Make. See Docker's QEMU documentation. Emulated compilation and compression can be much slower. Build both targets sequentially with the same version, then publish the two files:

make build-linux-docker LINUX_ARCH=amd64 BUILD_VERSION=2026.9.7032027
make build-linux-docker LINUX_ARCH=arm64 BUILD_VERSION=2026.9.7032027
GITLAB_DEPLOY_TOKEN=your-deploy-token make publish-linux \
  APPIMAGE_X86_64='.cache/linux-docker/amd64/target/release/bundle/appimage/Omni Chat_2026.9.7032027_amd64.AppImage' \
  APPIMAGE_AARCH64='.cache/linux-docker/arm64/target/release/bundle/appimage/Omni Chat_2026.9.7032027_aarch64.AppImage'

Building and self-signing the iPad app

The native iPad target requires an M-series iPad with at least 8 GB RAM, iPadOS 26.0+, and full Xcode 26.6+ on a Mac. The stable reference test point is iPadOS 26.5. Install CMake, CocoaPods, the iOS Rust target, and Tauri CLI, then generate and build:

sudo xcode-select -s /Applications/Xcode.app/Contents/Developer
brew install go node make cmake rustup cocoapods
export PATH="$(brew --prefix rustup)/bin:$PATH"
rustup default stable
rustup target add --toolchain stable aarch64-apple-ios aarch64-apple-ios-sim
rustup component add --toolchain stable llvm-tools
cargo install tauri-cli --version "^2" --locked
cmake --version
pod --version

make check-ipad-prereqs
make build-ipad

The Tauri Xcode project under desktop/gen/apple is tracked and ready after checkout so Xcode Cloud can select it. Run make init-ios only when Tauri asks you to regenerate the mobile scaffolding, then review and commit the generated changes.

Xcode Cloud's post-clone hook installs llvm-tools and rust-src before compilation. The iOS wrapper selects the real stable compiler and propagates that toolchain to child commands, preventing competing rustup downloads during parallel builds.

The iPad build explicitly enables Go modules, even if your global Go configuration has GO111MODULE=off.

The iPad app uses the desktop icon artwork from desktop/app-icon.png. Project initialization and each iPad build regenerate the required icon sizes. To change the icon, replace that artwork, rebuild, and install over the existing app to keep its data.

For a cable-free test, install the iOS 26.5 Simulator runtime under Xcode → Settings → Components, then run make check-ipad-simulator-prereqs and make dev-ipad-simulator. Tauri prompts for a destination; an exact one can be supplied as IPAD_SIMULATOR="iPad Air 11-inch (M4)". No Personal Team or Developer Mode is needed. The Apple-Silicon simulator exercises the embedded Go core and Metal-linked library, but its model speed and memory pressure are not physical-iPad results.

Xcode 26.6 supports iPadOS devices only through 26.5. An iPad running iPadOS 27 beta needs a current Xcode 27 beta supported by the Mac; Developer Mode does not add missing device support. Check Apple's Xcode SDK and device-support table, especially when the device beta is newer than Xcode.

Install CocoaPods with Homebrew. If Tauri reports that the package is missing and falls back to a system-Ruby installation requiring sudo, stop and run brew install cocoapods; do not use the Ruby-gem fallback.

The llvm-tools component supports the pinned upstream swift-rs Xcode 27 compatibility fix. The generated Xcode build phase uses the repository's Rust wrapper, which locates Homebrew rustup and the Cargo-installed Tauri CLI even though GUI applications start with a minimal PATH.

iOS 27 also turns UIKit's missing-scene-lifecycle warning into a launch-time breakpoint at UIApplicationMain. Omni Chat supplies a tracked scene manifest naming tao's delegate, mirrors it into generated Xcode files, and pins tao's merged scene-configuration lifetime fix. After updating the source, rerun make build-ipad before pressing Run in Xcode; continuing past the breakpoint does not fix the lifecycle declaration.

The first llama.cpp compile can take several minutes. Missing Git-repository and OpenMP messages are expected warnings for the pinned source archive: iPad inference uses Metal and Accelerate. The iOS build excludes llama.cpp's standalone executable and web UI and links only its static API server. Keep the make build-ipad terminal open after Xcode launches; a missing-signing-certificate warning is expected until a Personal Team is selected.

In Xcode → Settings → Accounts, add your Apple ID. In the generated Omni Chat target's Signing & Capabilities pane, enable automatic signing and select the resulting Personal Team. Keep bundle id org.geekaholic.omni-chat, select the connected iPad, and press Run. Trust the Mac when prompted. If requested, enable Developer Mode under iPad Settings → Privacy & Security, restart, and Run again. After the first install, approve the Personal Team certificate under Settings → General → VPN & Device Management → Developer App → Trust, then Run or open Omni Chat again. This certificate approval is separate from trusting the Mac and enabling Developer Mode.

A free Personal Team profile normally expires after seven days; reconnect and Run from Xcode to reprovision. This does not produce an App Store/TestFlight or redistributable build. Re-signing over the same installed bundle preserves data; deleting the app removes its local database and downloaded model. The full procedure and troubleshooting are in docs/ipad-self-signing.md.

TestFlight builds come from Xcode Cloud, not from a local archive. Each release begins by bumping CFBundleShortVersionString and CFBundleVersion across the five tracked files that carry the app version, then pushing the branch the workflow watches; a repeated build number is rejected at upload. The procedure is in docs/ipad-testflight.md. desktop/Info.ios.plist declares ITSAppUsesNonExemptEncryption = NO, so builds skip the export compliance prompt.

First launch offers a resumable ~1.75 GB download of the pinned Qwen 3.5 2B MLX snapshot. The Go core checks every file's pinned length and SHA-256 before the MLX helper can load it. It supports private offline text chat and tools; vision is not included. Leave the app in the foreground for downloads and inference because iPadOS may suspend the complete process. Docker services, host coding agents, stdio MCP servers, and SFTP are unavailable in the iPad sandbox; HTTP MCP servers, remote OpenAI-compatible endpoints, and other network features remain optional.

Editing endpoints from inside the app: AI Servers and Models

Omni Chat edits endpoints.json itself — in the desktop app, and for admins in the web app too — so getting running does not mean copying docs/endpoints.example.jsonc and hand-editing JSON. Open the user menu → Settings; every user can open it. Admins (the owner, plus anyone given the Admin switch in Members) edit everything; everyone else sees Integrations and a read-only view of the servers, model assignments and web tools, never their URLs or keys. On the web, saving restarts the server in place and a typed API key is stored in the server's endpoints.json (there is no Keychain); under Docker the ./config mount must be writable by the container user, uid 10001 (the shipped compose.yaml mounts it read-write; on Linux run sudo chown 10001:"$(id -g)" ./config && sudo chmod 775 ./config once; add :ro to lock web saving off). Settings → Integrations connects ChatGPT (sign in, then Use for Chat / Image generation; its model list follows your Codex CLI when installed, and so does each model's context window, so long chats keep their history) and your MCP servers. Local services, Agent handoff and Desktop stay desktop-app-only. Settings splits into AI Servers ("Discover and manage servers") and Models ("Defaults for each job"), followed by Web & tools, Local services, Agent handoff, and Desktop. It opens on AI Servers when you have no saved servers yet, else on Models. On a phone-width screen (820 px or narrower, e.g. an iPhone in portrait) the page list is a slide-in menu: Settings opens with it showing, and the ☰ button beside the page title brings it back. Server rows there fold their status into a dot on the server-type tile plus a word on the address line.

AI Servers holds the servers themselves. Connections lists what you've saved — each row shows a status (Connected / Needs key / Unreachable / Unsaved), the label, URL, its "Used for" chips, and a model count, with Manage expanding the full editor in place — plus an Add server button for typing one in by hand. Below that (above it, when you have no servers yet) is Network discovery: an auto toggle (servers on this machine are always used, whatever the toggle says) and a Scan now button for an on-demand full pass — this machine, named local hosts, and the attached network — regardless of the toggle. Scan now lists every open server it finds with its model count, and a key-locked one as Locked · needs API key. Each server shows a lettermark tile for its detected server type (llama.cpp, vLLM, Ollama, LM Studio, oMLX, LiteLLM, SGLang, OpenAI; a generic server glyph when unknown) next to its model count. A status dot on each result says where it stands: green In use (auto-discovery already offers its models in the picker) or Saved, orange Needs key, grey Available; each result gets its own Add connection button, which opens the editor with the URL and a suggested label already filled in — and, for a locked one, the key field focused so you can type it straight in.

The connection editor itself (Add server / Manage / add-from-scan) has a Label (an internal id is derived once behind the scenes and never shown again — rename the label as often as you like), Base URL (the server address only — no /v1 at the end), Server type (set by detection, or pick it yourself), Authentication (None / API key — ChatGPT sign-in lives in Integrations), and Used for switches — Chat, Embeddings, Image generation, Read aloud, Voice input, at least one required. Turning a use on reveals that job's settings inline (Chat: context window, max output, sampling, thinking, model overrides; Image: model, carrier model, stream, timeout; Read aloud: model, voices, timeout; Voice input: model, language, timeout). Test connection runs one check per enabled use and shows a row per use, then Save & restart applies it. Saving works even when a test fails or was never run, with a warning.

Models holds the default for each job — each select lists only connections with that use turned on on AI Servers. Default chat server picks a server only (no model; the model you used last still wins on a new chat). Embeddings is a connection + model + dimension, with a warning that changing the model re-indexes memory — it's optional; memory works without it. Image generation, Read aloud (plus a default voice), and Voice input each pick a default connection the same way. A job with no tagged connection shows "No servers are set up for this — add one in AI Servers" instead of an empty select. Removing on AI Servers a connection that a Models default still points at is blocked at Save, and the problem message names — and clicking it opens — whichever page fixes it.

Auxiliary model

The Models page's Auxiliary card offers On-device (AppleFoundation, native Apple shells only), Managed (the Local services llama.cpp below), a connection + model id, or Disabled. This small, non-thinking model handles chat titles, background memory jobs and long-chat summaries without occupying the chat model. The Use for Smart Router switch is on by default; turn it off when routing should remain heuristic or use a separately configured classifier. Show in model picker is also on by default. When enabled, the model can be selected for an ordinary chat turn but remains excluded from Auto routing and omni-chat launch. Turn it off to hide new selection without interrupting auxiliary jobs or chats that already use it. Web operators can set aux.show_in_model_dropdown: false in endpoints.json. The app reads that model's live endpoint metadata for prompt budgeting; this includes apfel's reported context limit for apple-foundationmodel. If Apple's framework returns its opaque retryable generation error before emitting any text or tool call, the turn is retried once safely. Apfel can also fall back to writing a tool_calls JSON block as ordinary text; the app recognizes a complete block only when every named tool was actually offered, then executes it through the normal validation and budget loop instead of showing raw JSON in the conversation.

Test sends one real completion using the values currently in the form, including a connection and key that have not been saved yet. A cold local model can take several seconds to load. The test never starts a service or downloads weights.

Image generation, read aloud and dictation

These three jobs are chips you tick on a connection (AI Servers), with defaults you pick on the Models page — so a local backend and a hosted one can sit side by side, and a server serving several jobs (chat, images, TTS) is typed and keyed only once. They live in the same file as the chat connections and are saved by the same Save & restart button.

Image generation's chip settings take a model, and — via the connection's Authentication and Advanced → protocol — either an OpenAI-compatible backend or a Codex subscription through the Responses API. The two want different values: a local image server is something like z-image-turbo on a port on this machine, while Codex serves gpt-image-2 from chatgpt.com with protocol: "responses". The Codex case also asks for a carrier model — the text model that carries the image request.

A sign-in or Responses API server (Codex) can only be used for Chat and Image generation: embeddings, read aloud, voice input and the auxiliary model send just a URL and key, so the editor disables those switches, the Models page leaves such servers out of the auxiliary and embeddings choices, and the core refuses the file if one is named there anyway.

A connection picks its credential explicitly on Authentication: an API key, or one of the sign-in providers the core offers. Choosing Codex preselects Sign in with ChatGPT, because that is how a subscription is authenticated — leaving it on the key field is what used to produce connections with no credential at all, which then failed with a bare 401. You sign in from the editor itself, so image generation no longer requires creating a separate chat connection first — tick both Chat and Image generation on one Codex connection instead.

One sign-in, shared by everything bound to it

A provider sign-in is stored once per connection, and every job section that references that connection rides it. If a Codex connection is ticked for both Chat and Image generation, one sign-in covers both — and signing out drops it for every job using that connection.

Test connection's Image row checks credential and reach without generating, listing the backend's models when it has a listing and falling back to a health check otherwise (which is what the bundled image sidecar does).

The image test does not generate an image

The smallest image the API offers is 1024×1024 — there is no thumbnail mode — which is real money against a hosted backend and a queued GPU slot on a local one. A settings button that quietly spent one on every press would be a trap, so this one checks the connection and says so.

Read aloud's chip works with any OpenAI-compatible speech server, including the bundled voice sidecar. Test connection's Read aloud row synthesizes a short phrase and plays it, which is the point: no status line can tell you the voice is the one you wanted. It also reports the voices the backend lists, with one click to adopt them into the Voices field — leave that field empty and the backend is asked at runtime instead.

Voice input's chip configures dictation. Test connection's Voice input row sends a half-second clip generated on the spot, containing silence. It passes when the backend accepts it, which proves the route exists, the key is accepted and the model id resolves — but not that anything was heard, because there is nothing to hear. Whisper given silence returns either nothing or a small hallucination, so a test that demanded a transcript would fail a backend that works perfectly.

What a save can and cannot remove

connections — like endpoints, search, and fetch — always moves on save, because every client that speaks schema 2 renders it; the other job sections (images/tts/stt/aux/embeddings) move only when the client actually sent them, so an older client leaves a section it doesn't know about alone. Removing your last image connection really does remove it. The Auxiliary card can replace or disable the background-model section written by Local services, while preserving its advanced dedicated-router model field.

Local services

The Settings → Local services page can start two optional sidecars for you, so a fresh install is not stuck with search and background intelligence switched off. Starting one also writes the config that points the core at it, so it asks for a restart afterwards.

ServiceNeedsUnlocks
Web search (SearXNG)Dockerweb_search, search_images
Background models (llama.cpp)Nothing on macOS — llama-server is bundledChat titles, memory extraction, rolling summaries, semantic routing

For the background models you pick what to host. Background model runs one ~700 MB model — not a reduced setup, since that model also serves as the router's semantic classifier by default, so titles, memory and routing all work. Background + router adds Arch-Router-1.5B as a purpose-built classifier; both are hosted by one llama-server in router mode, on one port, loaded on demand. Models swap by default, which is safe on any machine; Keep both models loaded holds them resident and wants roughly 32 GB.

Test background models runs one real request per configured model and reports each as ready in green or not ready in red, with a timing. It is the only honest readiness signal, because Running does not mean loaded: router mode loads a model only when a request names it, and the service probe uses /health, which answers 200 with nothing resident. The first test is therefore slow — it includes reading the weights off disk, a few seconds on a cold machine — while later tests return in milliseconds. The test never downloads anything; that is what Start does. With a dedicated router classifier configured there is a row per model, so a failure says which of the two is wrong.

Managed services survive config saves — and stop when you quit

The core restarts on every config save, so a loaded model deliberately survives those restarts rather than reloading hundreds of megabytes each time you edit a setting. Quitting the app is different: everything the app started — the model server and its omni-managed-* containers — is stopped with it, so nothing keeps eating memory after the app is gone. A service you started yourself by hand still shows here as running (state is discovered by probing), but quitting leaves it alone: only what the app launched is stopped. If the app is force-killed, the leftovers are picked up on the next launch and stopped at that session's quit.

On macOS, an off-by-default Keep running in background setting (under Settings → Desktop) changes what closing the window means: the window hides, the app stays visibly in the Dock, and the services keep running — useful when a model server should stay warm. A Dock click brings the window back; quitting from the Dock or with ⌘Q still stops everything. Note this also applies to a dev run: Ctrl-C on make dev-desktop now stops the services it started too.

The macOS app ships its own llama-server (23 MB, in Contents/Resources/llama/), so background models work on a Mac with neither Docker nor Homebrew installed. A native binary is always preferred over a container because containers on macOS never see the GPU, and the bundled build is Metal-enabled. Resolution order is the bundled binary, then llama-server on PATH, then Docker — so a Linux build with no bundled binary still works through a container.

Managed services can create an optional-service-only endpoints.json; no remote chat endpoint is required first. This is the same valid shape used on iPad, where @on-device is injected at runtime instead of being written to the file.

Image generation, TTS and STT are deliberately not managed: a container cannot reach the GPU on macOS, so a one-click container there would be unusably slow while looking supported. Those keep their native and remote-host paths described below.

Agent handoff

The Agent handoff page sets up Pi, Hola, or both — letting a chat pass a coding task to an enabled agent, which reads and edits files on this machine. It writes <config-dir>/agents.json, a different file from the endpoints config. The owner console stages both documents and applies them with one Save & restart action.

Pick where the agent works first. The default gives every chat its own fresh subfolder under one workspace folder you choose — created on the first handoff, reused by follow-ups in the same chat, and never your existing files or the folder itself. The other mode lets the agent work directly inside folders you list, for real repositories; there the model proposes an existing directory under one of them and you approve the exact path.

Turn Pi and Hola on independently, then pick where each one runs. Installed on this machine runs pi or hola-coder directly — install it from pi.dev or the Hola usage guide and the form finds it. Docker container runs each delegation in a throwaway container instead. The repository supplies only the Pi image, built with make agent-image; for Hola, give a custom image containing Hola 0.6+ and hola-coder. The form discovers and tests each enabled agent separately.

Hola uses a private, temporary 0.6 profile

Hola profile values override provider environment variables, so Omni Chat generates an omni-chat profile for each delegation. It pins the chat model and scoped proxy, enables the coding toolsets and OpenAI tool-call parser, and applies the documented balanced convergence policy. The profile omits api_key, so Hola reads the short-lived token from OPENAI_API_KEY in the child process environment; the token never enters the profile or command line. Omni Chat launches each handoff with --new, so files persist in the workspace but Hola never replays stale model history from hola-coder.db. Personal ~/.hola settings are excluded. A repository-level .hola/profiles.json would take precedence, so that workspace is refused with an explanation rather than silently routing somewhere else. See Hola's profile guide and 0.6 options.

Use the bridge network on macOS

Docker Desktop's host networking does not give a container this machine's loopback, so an agent on the host network cannot reach the model and the run ends without output. On bridge the core rewrites localhost endpoint URLs to host.docker.internal, which works. The form defaults to bridge.

Folders the agent may work in is the safety boundary, and nothing is filled in for you: the agent can only ever read and write inside a folder that resolves under one of these, so it is worth choosing deliberately rather than accepting a guess.

Each Test button reports a row per check — the command resolved to a real path, the Docker daemon answered, the image exists, the folders resolve, and finally the agent itself runs and reports its version. That last row is the point: a binary being present says nothing about whether it works, and a half-installed npm package, a broken node or an image built for the wrong architecture all pass a file-exists check and fail here. It deliberately does not run a real delegation — that needs a model, a chat, your consent and minutes of waiting, and it would write into your own repository.

Turn off deletes agents.json (the format has no off switch, and an empty agent list is invalid) and keeps an agents.json.bak beside it, so setting it up again is not starting over.

A GUI app does not inherit your shell's PATH

An app launched from Finder or the Dock gets launchd's minimal environment, which contains neither /opt/homebrew/bin nor /usr/local/bin — so a plainly installed Pi/Hola command, or Docker itself, can look missing. The core therefore searches the usual install directories after PATH, never instead of it. If your copy lives somewhere unusual (a version manager, for instance), give the full path in the Command field.

Self-hosting SearXNG

The search section of endpoints.json names one provider at a time. For the self-hosted option:

{ "search": { "provider": "searxng", "base_url": "http://localhost:8888", "max_results": 5 } }

Bring the sidecar up on its own — it is not part of the normal Compose stack:

printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.search-only.yaml up -d

That publishes SearXNG on 127.0.0.1:8888. Point the app at it — Settings → Web & tools → Web search → SearXNG in the desktop app, or CHAT_WEB_SEARCH_URL=http://localhost:8888 for web mode — and press Test search. Running the app and the sidecar in one Compose project instead? Use -f compose.yaml -f compose.search.yaml, which sets the URL for you; docs/endpoints.docker.example.jsonc shows that form.

JSON output is the setting that matters

deploy/searxng/settings.yml must list json under search.formats, or SearXNG answers 403 and the core reports "SearXNG rejected JSON output; enable the json search format." That is exactly what Test search surfaces.

The same screen configures web search and web fetch, which used to be environment-only and so unreachable in a bundled app. Search offers Tavily (hosted — paste a free API key, nothing to run) or SearXNG (self-hosted — give it a URL); a Test search button runs one real query, so a rejected key or a SearXNG without JSON output enabled shows up immediately. SearXNG additionally provides image search; Tavily does not. On iPad, Tavily can be saved without adding a remote chat endpoint and is available to the injected @on-device model. Web fetch is a single switch: its default provider runs inside the core behind an SSRF guard, so fetch_url and fetch_image need nothing else running.

Saving does three things in order: writes any new API key to the OS secure store, writes the config file, then restarts the core. The restart is the apply mechanism — the core reads its configuration once at startup — and it ends any answer in progress. The previous version is kept at endpoints.json.bak, and a config the core would refuse to boot on is rejected before the file is touched.

Keys go to the platform secure store, not the file

endpoints.json stores only api_key_env — the name of an environment variable. The shell reads the value from Apple Keychain or Linux Secret Service and injects it when it spawns the core, so the secret never reaches disk. Two consequences: the first launch after storing a key raises a macOS permission prompt (an ad-hoc build gets a new signature on every rebuild, so during development this recurs — the launch is not blocked, the shell waits ten seconds then starts without the stored keys), and a core started outside the shell (web mode, Docker) reading the same file will not have those variables set. Keys already inline in the file keep working everywhere; the UI shows them as stored in plain text and offers to move them.

This is native-shell-only by design. GET/PUT /api/config return 501 not_implemented unless the core was started with CHAT_DESKTOP=1, which only the Linux, macOS and iPad Tauri shells set — a shared web instance never exposes instance config, endpoint keys, or the ability to repoint the app at another backend.

If the core fails to start — most often an invalid file in ~/.config/omni-chat/ — the window stays on the shell's own splash page and shows the core's output instead of a blank screen, with Try again and Restore previous config buttons.

Three desktop differences

Confirmations (Archive, Delete chat, New project, discarding a temporary chat) are drawn by the app rather than by the operating system. The webview does not implement the browser's alert/confirm/prompt, so a native dialog would never appear — the same in-app dialog is now used in the browser too, so both behave identically.

Saving a file — exporting a chat, downloading a code block or an image — opens the macOS save panel, because the webview ignores the browser's download mechanism. Print / Save as PDF is unavailable in the desktop app for the same reason; use Download HTML and print that from a browser.

Links open in your default browser. The webview cannot open a tab, so citations, source cards and links in a reply are handed to the operating system instead — and so is OAuth sign-in, which is why signing in to a provider was impossible in the app before. Only http and https are ever opened, and only to another site: the app's own links, such as an attachment, stay inside, because a separate browser has no session for the local core.

Production build & run

Build the web SPA and the core binary (the SPA is embedded into the binary):

make build
# binary: .cache/bin/omni-chat-core

Run it with a persistent database path and your endpoint:

CHAT_ENDPOINT_URL=http://your-endpoint:port \
  CHAT_BIND=0.0.0.0 \
  CHAT_PORT=8080 \
  CHAT_DB_PATH=/persistent/path/app.db \
  ./.cache/bin/omni-chat-core
Before sharing on a network

CHAT_BIND=0.0.0.0 exposes the instance to the LAN. Only set it when you intend to share, and put it behind appropriate network controls.

Docker Compose deployment

To include the optional self-hosted SearXNG provider, use:

printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.yaml -f compose.search.yaml up -d

The overlay enables JSON output, connects Omni Chat over the private Compose network, and exposes SearXNG only at 127.0.0.1:8888 by default. The generated SEARXNG_SECRET is required, belongs in the ignored .env file, and must not be committed or copied into settings.yml.

For a host-run core (make dev) instead of the full Compose stack, start just SearXNG and point the core at its published port:

printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose -f compose.search-only.yaml up -d
CHAT_WEB_SEARCH_URL=http://localhost:8888 make dev

To include the optional llama.cpp auxiliary-model sidecar (see Auxiliary model above), use:

docker compose -f compose.yaml -f compose.aux.yaml up -d

The overlay wires CHAT_AUX_URL / CHAT_AUX_MODEL to an internal aux-model service running ghcr.io/ggml-org/llama.cpp, which downloads the configured Hugging Face GGUF (default LiquidAI/LFM2.5-1.2B-Instruct-GGUF, ~731 MiB) into a named volume on first run. This one model handles chat titles, memory jobs, and smart-router classification, so on its own it needs no separate router sidecar. It is stackable with the other overlays, e.g. docker compose -f compose.yaml -f compose.search.yaml -f compose.aux.yaml up -d. For a host-run core, start just the sidecar and point the core at its published port:

docker compose -f compose.aux-only.yaml up -d
CHAT_AUX_URL=http://localhost:8091 \
CHAT_AUX_MODEL=lfm2.5-1.2b-instruct make dev

To also run the optional dedicated llama.cpp router-classifier sidecar (see Semantic router classifier below), use:

docker compose -f compose.yaml -f compose.aux.yaml -f compose.router.yaml up -d

The overlay wires CHAT_ROUTER_CLASSIFIER_URL / CHAT_ROUTER_CLASSIFIER_MODEL to an internal router-classifier service running ghcr.io/ggml-org/llama.cpp, which downloads the configured Hugging Face GGUF (default katanemo/Arch-Router-1.5B.gguf, ~1 GiB) into a named volume on first run — allow a few minutes before the sidecar reports healthy and Omni Chat starts. When present, this dedicated classifier outranks the auxiliary model for routing (the aux model then does titles and memory only); omit it and the aux sidecar classifies via the fallback. For a host-run core, start just the sidecar and point the core at its published port:

docker compose -f compose.router-only.yaml up -d
CHAT_ROUTER_CLASSIFIER_URL=http://localhost:8090 \
CHAT_ROUTER_CLASSIFIER_MODEL=arch-router-1.5b make dev

To run the fetch_url tool through a self-hosted reader sidecar instead of the in-core direct provider, use:

docker compose -f compose.yaml -f compose.fetch.yaml up -d

The overlay enables the tool (CHAT_WEB_FETCH=true), switches it to the reader provider, and points Omni Chat at an internal reader service on http://reader:8081 (also published on 127.0.0.1:3001 for debugging). The service is ghcr.io/jina-ai/reader:oss, Jina AI's official self-host image (multi-platform: linux/amd64 and linux/arm64, so it runs unmodified on Apple Silicon) that renders JavaScript in a bundled headless browser; override READER_IMAGE/READER_VERSION to pin or replace it. It is stackable with the other overlays. When to use it: the default direct provider needs no extra container and handles most static or server-rendered pages, but it only does a plain HTTP GET — it cannot run JavaScript, is blocked by anti-bot defenses such as Cloudflare, and cannot read YouTube transcripts. The reader sidecar runs a real browser, so it handles all three, at the cost of a much larger image, more memory, slower fetches (consider raising CHAT_WEB_FETCH_TIMEOUT, up to 60s), and sending target URLs to that container. Prefer direct and switch to the reader overlay only when you need one of those capabilities. For a host-run core, start just the reader and point the core at its published port:

docker compose -f compose.fetch-only.yaml up -d
CHAT_WEB_FETCH=true CHAT_WEB_FETCH_PROVIDER=reader \
CHAT_WEB_FETCH_URL=http://localhost:3001 make dev

To offer the generate_image tool, use the image-generation overlay (see Running Omni Chat, combination g, for the requirements — an NVIDIA GPU with the container toolkit, and a ~12 GB first-run weight download into the imagegen-models volume):

docker compose -f compose.yaml -f compose.imagegen.yaml up --build -d

The overlay builds the bundled Z-Image-Turbo sidecar (docker/imagegen/ — FastAPI + diffusers exposing POST /v1/images/generations) and sets CHAT_IMAGE_GEN_URL=http://imagegen:8000 on the app. Tune it with IMAGEGEN_MODEL_ID (any diffusers text-to-image model id; default Tongyi-MAI/Z-Image-Turbo), IMAGEGEN_STEPS, and — for precise instruction edits via the edit_image tool — IMAGEGEN_EDIT_MODEL_ID (recommended: Qwen/Qwen-Image-Edit-2509, Apache-2.0 but ~20B; alternative: black-forest-labs/FLUX.1-Kontext-dev, 12B, gated + non-commercial — accept its license on Hugging Face and set HF_TOKEN in .env; unset falls back to img2img with the base model, tuned by IMAGEGEN_EDIT_STRENGTH, default 0.7 — raise toward 0.8 for stronger restyles — and IMAGEGEN_EDIT_GUIDANCE, default 1.0, which only helps CFG-capable models as the distilled base can burn above 1) in .env. Tongyi's announced Z-Image-Edit has no public weights as of July 2026. On non-x86 GPU hosts (Jetson/L4T or arm64 CUDA), override IMAGEGEN_BASE_IMAGE with a matching CUDA-enabled PyTorch base and set IMAGEGEN_RUNTIME=nvidia — see the GPU note in Getting started. For a host-run core — or to run the sidecar on a separate GPU machine — start just the sidecar and point the core at its published port (127.0.0.1:8001 by default; set IMAGEGEN_PUBLISH_ADDRESS=0.0.0.0 and an IMAGEGEN_API_KEY/CHAT_IMAGE_GEN_API_KEY pair when it must be reachable over the network):

docker compose -f compose.imagegen-only.yaml up --build -d
CHAT_IMAGE_GEN_URL=http://localhost:8001 make dev

Any other OpenAI-Images-compatible server also works — set CHAT_IMAGE_GEN_URL (plus CHAT_IMAGE_GEN_MODEL/CHAT_IMAGE_GEN_API_KEY as needed) and skip the sidecar.

To use the default in-core direct provider under the base Compose stack (no sidecar), just add CHAT_WEB_FETCH=true to .env — no overlay needed. If you are testing a feature (such as fetch_url) that predates the published latest image, docker compose up -d alone keeps running that stale image; rebuild from your checked-out source with docker compose up --build -d to pick up local changes.

The repository's Compose configuration pulls the non-root, multi-platform image from the GitLab Container Registry, publishes the app only on 127.0.0.1:8080 by default, checks /api/health, and restarts the service unless it is stopped explicitly:

# Optional environment overrides; .env is ignored by Git.
cp .env.example .env
docker compose up -d
docker compose ps

Compose passes variables from .env into the container. Use it for CHAT_ENDPOINT_URL, CHAT_API_KEY, variables referenced by api_key_env, and other supported CHAT_* overrides. Set OMNI_CHAT_PUBLISH_ADDRESS=0.0.0.0 only when the app should be reachable from other machines, and protect the published port appropriately.

The default image is registry.gitlab.com/geekaholic/omni-chat:latest for both AMD64 and ARM64. Every successful main-branch pipeline also publishes an immutable full-commit-SHA tag. Set OMNI_CHAT_IMAGE=registry.gitlab.com/geekaholic/omni-chat:<full-commit-sha> in .env to pin a deployment. Private projects require docker login registry.gitlab.com before pulling. Use docker compose up --build -d to build from the local checkout instead.

Connecting a container to AI on the host

localhost inside a container is the container itself. Use host.docker.internal for an OpenAI-compatible endpoint running on the Docker host. Compose adds the host-gateway mapping needed on native Linux; Docker Desktop provides the same hostname. On native Linux, the AI process must also listen on an interface reachable from Docker's bridge rather than only 127.0.0.1. Bind the AI server to an appropriate host interface and restrict that port with host firewall rules.

Docker data and configuration

SQLite is stored in the named omni-chat-data volume. Operator-managed JSON files are bind-mounted read-write from ./config to /config, so admins can save Settings and the files survive image upgrades. The process runs as uid 10001; on Linux a directory you created is not writable by that user until you hand it over, even with no :ro on the mount. Settings tells the two causes apart: "cannot write to the config directory" means ownership, fixed once with the commands below; "the config directory is read-only" means the mount carries :ro. Append :ro yourself to edit files only on the host.

sudo chown 10001:"$(id -g)" ./config
sudo chmod 775 ./config

Files that Settings saves (endpoints.json and its .bak) are written mode 0600 and owned by uid 10001, because they can hold API keys. After the first save from Settings, edit them on the host with sudo.

cp docs/endpoints.docker.example.jsonc config/endpoints.json   # then delete the // comments
cp docs/profiles.example.json config/profiles.json
cp docs/SOUL.example.md config/SOUL.md
mkdir -p config/skills && cp docs/skills.example/*.md config/skills/
# Edit the files, then reload startup configuration.
docker compose restart

docs/endpoints.docker.example.jsonc is an annotated example (schema 2, connections[]) the app never loads — copy from it, then delete its comments before the result is valid JSON. It points to host.docker.internal:8000. If config/endpoints.json exists, it takes precedence over CHAT_ENDPOINT_URL. Existing JSON configuration is validated at startup exactly as it is for native runs.

Update and replace the container without deleting its data:

git pull
docker compose pull
docker compose up -d
Named volumes contain the database

docker compose down preserves omni-chat-data. docker compose down -v permanently deletes it.

For a consistent backup, stop writes while archiving the volume:

docker compose stop
mkdir -p backups
docker run --rm \
  -v omni-chat-data:/data:ro \
  -v "$PWD/backups:/backup" \
  alpine:3.22 tar -czf /backup/omni-chat-data.tar.gz -C /data .
docker compose start

Docker without Compose

docker volume create omni-chat-data
docker run -d --name omni-chat --restart unless-stopped \
  --add-host host.docker.internal:host-gateway \
  -p 127.0.0.1:8080:8080 \
  -e CHAT_ENDPOINT_URL=http://host.docker.internal:8000 \
  -v omni-chat-data:/data \
  -v "$PWD/config:/config" \
  registry.gitlab.com/geekaholic/omni-chat:latest

The image runs as uid 10001. On Linux, ./config must be writable by that user (sudo chown 10001:"$(id -g)" ./config && sudo chmod 775 ./config). Add :ro to the config mount to forbid saving from Settings. For a local source build, run docker build -t omni-chat:local . and use omni-chat:local as the final argument instead.

Importing from ChatGPT

Bring existing ChatGPT history into Omni Chat with the core import chatgpt CLI subcommand. In ChatGPT, go to Settings → Data controls → Export data; a download link for the export zip arrives by email.

./.cache/bin/omni-chat-core import chatgpt ~/Downloads/chatgpt-export.zip --user alice

<export> accepts the export zip as downloaded, an already-unzipped export directory, or a bare conversations.json. A bare JSON file has no attachments to pull from, so the importer prints a note to stderr and continues without them. Both the classic single-conversations.json layout and the newer sharded layout (conversations-000.json, conversations-001.json, … plus .dat assets with a conversation_asset_file_names.json restoring their original names) are supported.

FlagEffect
--user <username>target omni-chat username (required)
--db <path>SQLite database path (defaults to $CHAT_DB_PATH or data/app.db)
--skip-attachmentsopt out of importing images/files referenced by messages
--skip-memoriesopt out of importing ChatGPT memories
--skip-archivedopt out of importing conversations archived in ChatGPT
--dry-runparse the export and print the report without writing anything

For each conversation the importer takes the active branch only — exactly what was last seen in ChatGPT, not every edited-away alternative. It imports the title, timestamps, user/assistant messages, reasoning ("thoughts") folded into the following assistant message, images/files as attachments, and archived state; every imported chat gets a chatgpt tag. ChatGPT memories (bio-tool writes) are harvested into the user's user-scope memories, deduplicated against what's already there. Tool-call plumbing (code execution, browsing traces, canvas) is not imported and is counted in the summary.

Re-running the import against a newer export is incremental: conversations that haven't changed are skipped, conversations that grew in ChatGPT get only their new messages appended, and conversations whose active branch changed (an edit that switched branches) are skipped with a warning — existing messages are never rewritten or deleted. It is safe to import while the server is running (SQLite WAL). Imported chats have no model/endpoint set, so they pick the user's default model the next time they use them.

Importing under Docker Compose

When Omni Chat runs via docker compose, the same binary is the container's entrypoint, so the import runs as a one-off container that shares the service's database volume and environment:

docker compose run --rm \
  -v "$PWD/chatgpt-export.zip:/import/export.zip:ro" \
  omni-chat import chatgpt /import/export.zip --user alice

docker compose run reuses the service's CHAT_DB_PATH=/data/app.db and the omni-chat-data volume, so no --db flag is needed and the import lands in the same database the server uses — safe to run while the service is up. Bind-mount the export read-only at a fresh path such as /import/… (avoid /tmp, which the service config covers with a small tmpfs). All the flags above work the same way; add --dry-run first for a report before writing anything. If the image predates the import feature, rebuild it with docker compose build.

Quality gate

For contributors: make check runs fmt + lint + test + build and must pass before a change is considered done.


Chapter 13

Troubleshooting & FAQ

No models appear in the picker

The model list comes from the backend's /v1/models. Confirm CHAT_ENDPOINT_URL is set and reachable, or that your endpoints.json is valid. Remember that endpoints.json, if present, makes CHAT_ENDPOINT_URL ignored. A backend that fails at startup is shown as offline and skipped.

No models appear when Omni Chat runs in Docker

Do not use localhost for an AI server running on the host. Set the endpoint to http://host.docker.internal:<port>. If you mounted config/endpoints.json, update its base_url because that file overrides CHAT_ENDPOINT_URL, then run docker compose restart. On native Linux, also confirm the AI server listens on a bridge-reachable interface.

SearXNG exits because server.secret_key is unchanged

SearXNG rejects its bundled ultrasecretkey. Run printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env, then run docker compose -f compose.yaml -f compose.search.yaml up -d again. Compose reports a clear setup error before startup when the variable is missing.

Empty model picker with a LAN endpoint on macOS

If CHAT_ENDPOINT_URL points to a machine on your local network (e.g. http://192.168.x.x:8000 or http://<hostname>.local:port) and the picker stays empty even though curling the same /v1/models URL works, the cause is usually macOS Local Network privacy. Recent macOS releases block a process from reaching local-network addresses until you grant it access, which surfaces here as a "no route to host" error and an offline endpoint; public/internet endpoints are unaffected.

Fix: open System Settings → Privacy & Security → Local Network, enable the toggle for the terminal app you launch make dev from (Terminal, iTerm2, Ghostty, VS Code, …) and for Omni Chat when using the desktop build, then quit and reopen that app. The entry only appears after a LAN connection is attempted. On Sequoia the grant is per-binary, so an ad-hoc go run or unsigned sidecar can stay blocked even after Terminal or Omni Chat is allowed. If you launch through tmux/screen/ssh, start from a plain terminal window so the permission is attributed correctly.

I get a context_overflow error

The input didn't fit the context window. Lower Response Length, shorten the conversation, or raise CHAT_CONTEXT_WINDOW to match your model. The current user message is never trimmed — if it alone exceeds the budget you'll see this error.

Creativity changes nothing in the model's behavior

Sampling parameters are only sent when the backend honors them. Set CHAT_HONORS_SAMPLING=true (or honors_sampling on the endpoint). Otherwise Creativity only changes the style prompt text.

The model wants a different max-tokens field

Some endpoints require max_completion_tokens. Set CHAT_MAX_TOKENS_FIELD=max_completion_tokens globally, or max_tokens_field on the specific endpoint.

No sources are ever returned

Check that an embeddings endpoint is configured and that documents have actually been uploaded and indexed. For web results, check that Search is on for the chat. An empty corpus never produces citations.

Reasoning parameters for thinking models

Use the UI Thinking switch first, and Thinking effort in Settings for models that support it (Qwen 3.8). Omni Chat detects common adapters such as Qwen on llama.cpp and capability-advertising models on OMLX. Configure thinking.on_body and thinking.off_body in endpoints.json only for custom endpoints or models with different request knobs. To opt a non-inferred model into thinking effort, set thinking_overrides[].effort to qwen38.


Chapter 14

Acknowledgements

Omni Chat is released under its own Apache-2.0 license, but it stands on — and interoperates with — a large ecosystem of open-source projects. This chapter credits the projects it bundles, ships as optional sidecars, or connects to. Each project remains under its own license; follow the links for their terms. Thank you to their authors and maintainers.

Models

Third-party model license — Arch-Router-1.5B

The default smart-router classifier, Arch-Router-1.5B, is a Katanemo model published by DigitalOcean, LLC under the Katanemo Community License, not Omni Chat's own Apache-2.0 license. Built with DigitalOcean. Commercial use of Arch-Router-1.5B requires a separate license from DigitalOcean — see the model's license terms on Hugging Face before using it commercially. Katanemo Models are licensed under the DigitalOcean Community License, Copyright 2026 DigitalOcean, LLC. All Rights Reserved.

  • LFM2.5-1.2B-Instruct (Liquid AI) — the recommended default auxiliary "background intelligence" model (chat titles, memory jobs, router-classifier fallback).
  • OmniVoice (k2-fsa, Apache-2.0) — the default zero-shot text-to-speech model, with an Apple-Silicon port from mlx-audio.
  • Whisper (OpenAI) — the speech-to-text model, run via faster-whisper in containers and mlx-whisper on Apple Silicon.
  • Z-Image-Turbo (Tongyi-MAI) — the default image-generation model, served through Hugging Face diffusers. Optional image-editing models include Qwen-Image-Edit-2509 (Apache-2.0) and FLUX.1-Kontext-dev (Black Forest Labs — a gated, non-commercial license; review its terms before commercial use).

Inference & serving backends

Omni Chat bundles llama.cpp (ggml) as the Docker sidecar image (ghcr.io/ggml-org/llama.cpp) that serves the router and auxiliary models. As an OpenAI-compatible client it also interoperates with Ollama, vLLM, LM Studio, LiteLLM, and Apple MLX (mlx_lm).

In-browser code sandbox

  • Pyodide / CPython (MPL-2.0) — Python execution (bundles NumPy and pandas).
  • QuickJS (MIT) — JavaScript/TypeScript execution.
  • Yaegi — the embedded Go interpreter that runs Go snippets.
  • Wasmer (MIT) with WASIX bash and coreutils (GPL) — shell command execution.
  • xterm.js (MIT) — the interactive terminal UI for the sandbox.

Agents

  • Pi — the external coding agent behind the delegate_task handoff (npm package @earendil-works/pi-coding-agent).
  • Codex (OpenAI) — supported as a hosted backend over the OpenAI Responses API ("Sign in with ChatGPT").

Model Context Protocol

Tool integrations use the Model Context Protocol and its official Go SDK. Under Docker, stdio MCP plugins run as sidecars bridged to HTTP by supergateway. Example servers referenced in the docs include @modelcontextprotocol/server-filesystem, figma-developer-mcp, the GitHub MCP server, and Robinhood's MCP server.

The server and web app also build on many open-source libraries — among them Svelte, Vite, marked, DOMPurify, SQLite, and pdfcpu — each under its own license.