The Gateway

Install and Run

This chapter teaches you to get the gateway running on your machine: how to install it, how to start it with a config file and a profile, and how to confirm it is healthy. You do these things every time you bring the gateway up, so they are worth learning well.

Install the binary

The gateway is a single binary named promptforge-gateway. It serves an OpenAI-shaped inference API. Install it with cargo:

cargo install gateway

Confirm the install by printing the version:

promptforge-gateway --version

Start the gateway

Start the gateway by naming a config file and a profile:

promptforge-gateway --config gateway.toml --profile main

The --config flag gives the path to the config file. The --profile flag names the profile to activate. The gateway always starts from one config file and one active profile.

Startup is bind-first. The gateway opens its listener and answers health, status, progress, configuration, and every ready route immediately, then provisions models afterward as one queued boot command. Downloads, local model spawns, and the speech engine load all run inside that command while the gateway is already serving, and you can watch it on the status and progress endpoints. A configured model that is still loading answers 503 with the code model_loading until its provisioning finishes.

You can supply both values through environment variables instead of command-line arguments. The config path comes from --config or from PROMPTFORGE_GATEWAY_CONFIG; the flag wins when both are set. The profile comes from --profile, then PROMPTFORGE_PROFILE, then the sibling state file the gateway keeps beside the config.

You can also start the gateway with no config file at all. When no gateway.toml exists beside the executable, in the working directory, or in the user profile's .promptforge directory, the first run writes a default config there - loopback-only on an OS-assigned port, with a fresh random bearer key and trust_loopback = true so callers on the same machine need no key - and boots from it. The generated file notes the caveat beside that line: on a shared machine any other OS account can then use the gateway, and trust_loopback = false requires the key from everyone. The generated config selects a profile named default, so a bare first boot needs no flags.

The system tray

On a desktop system the gateway's face is the system tray. The icon shows the gateway's state, and its menu lists a status line, a Workshop item that launches the Workshop application when the installer laid it beside the gateway, a Settings item that opens the configuration UI in your browser, a Launch at Login toggle, and Quit. A gateway started at login never opens a browser or a window.

For servers and CI, --no-tray keeps the plain headless loop. In a tray-less environment, --print-url prints the Settings URL to stdout once the gateway is bound. --browser opens the Settings page in your default browser once bound; the installer uses it on a Gateway-only install's first run. Launching promptforge-gateway while one is already running never starts a second copy: it opens the running gateway's Settings page instead.

After every successful bind the gateway writes a gateway discovery file (gateway.json in the run directory under the state directory) that records its port, bearer key, and process id. PromptForge components read that file to attach to the running gateway instead of starting a second one, and a clean shutdown removes it.

Check that it is healthy

Once the gateway is serving, probe its health endpoint:

curl http://127.0.0.1:8081/health

GET /health needs no credentials. It always answers 200 while the gateway is serving.

Every /v1 route is authenticated with the shared bearer key from the config file. The address in this request is the bind value from the [server] section of the config file; 127.0.0.1:8081 is an example bind. A request with a wrong token is rejected with status 401 and error code unauthorized, from any peer:

curl -H "Authorization: Bearer wrong-token" http://127.0.0.1:8081/v1/models

From the gateway's own machine you can leave the key out entirely. With the default trust_loopback = true, a loopback request that presents no credential is admitted:

curl http://127.0.0.1:8081/v1/models

This convenience has one cost: on a shared machine, any other OS account can use the gateway the same way, including reading upstream API keys from the admin config surface. Set trust_loopback = false in [server] to require the key from every caller. The configuration chapter covers the rule in full.

Choose what to build

Build-time feature flags decide which capabilities exist in the binary. The flags local, web-search, stt, and config-ui are on by default. A headless build without local refuses any configuration that declares local models; the refusal happens at startup.

Run it as a service on Linux

On Linux the release archive contains a sample systemd unit. The unit runs the gateway as a service with a fixed config path and profile, and restarts it automatically on failure:

ExecStart=/usr/local/bin/promptforge-gateway --config /etc/promptforge/gateway.toml --profile main
Restart=on-failure
RestartSec=5

The gateway holds vendor credentials, so run it as a dedicated unprivileged user. The sample unit does this with DynamicUser=yes and keeps state in a systemd-managed state directory (StateDirectory=promptforge).

Watch the logs

A serving gateway logs to gateway.log in the logs directory under the state directory (~/.promptforge/logs on a default install) and mirrors the same stream to stdout. Startup rotates the previous run's log aside - gateway.log becomes gateway.log.1 - and keeps five previous runs, deleting the oldest. Every record crosses a redaction pass before it reaches disk: bearer tokens, authorization and cookie header values, and api_key assignments are masked. The log location is never configurable, so a config failure still has somewhere to report itself.

Control log verbosity through the standard RUST_LOG environment filter. The speech library logs at warn level by default, so it stays quiet unless you ask for more.

Startup failures appear on stderr with the full cause chain: one error: line followed by one caused by: line per cause, and the same chain lands in the log file. Once the gateway is serving, the log shows the bound address. If you configured port 0, the log reports the real bound port.

Inspect a failed run

When a gateway run fails before it can serve, promptforge-gateway diagnostics finds the evidence without any config knowledge. It prints a read-only JSON report: the state directory, the resolved config path and whether it exists, the current and retained log paths and which exist, the gateway discovery file, whether a gateway is running, and the version. It never serves, rotates a log, parses a config, or mutates the state directory, and it never prints secrets - no bearer key, environment value, config content, or log content. The generated config points at it in a comment.

Stop the gateway

From the tray, choose Quit. From a script or another PromptForge component, send an authenticated POST to the /shutdown route with the bearer key; it answers 202 and then the server goes down. Under --no-tray, Ctrl-C stops the gateway cleanly. Every path drains in-flight requests before the exit.

The Configuration File

This chapter teaches you the shape of the one file that configures the whole gateway. You will learn the version key, the server section, how to keep secrets out of the file, when a same-machine caller needs no key, and the loopback wall that guards the admin surface. Every other chapter adds sections to this file, so a solid mental model here pays off everywhere.

One file, one version

You configure the gateway in a single version-2 gateway.toml file. The file owns the global settings, the complete model catalog, and the profiles. The file must declare its version on the first line:

config-version = 0

Any other version fails to load. There is no silent upgrade path.

A minimal configuration

A minimal configuration has one [server] section, one or more [[endpoint]] backends, and one or more [[model]] entries that map public names to upstream aliases:

config-version = 0

[server]
bind = "127.0.0.1:8081"
api_key = "${GATEWAY_KEY}"

[[endpoint]]
id = "openai"
protocol = "openai"
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"

[[model]]
name = "gpt-5"
kind = "chat"
description = "GPT-5 via OpenAI"
context = 272000
thinking = "switchable"
upstream = "gpt-5"
endpoints = ["openai"]

The [server] section sets the socket address and the shared bearer key. Every request from another machine must present the key, and a key that is presented is always checked. The key must not be empty. A third field, trust_loopback, controls whether callers on the gateway's own machine may skip the key. It defaults to true, which on a shared machine also admits every other OS account there; set trust_loopback = false to require the key from everyone. The rule is covered in full below.

The model catalog lives in the same file. Remote models are [[model]] entries. Local models are [[local_model]] entries. Speech models are [[stt_model]] entries. Later chapters cover each kind.

Keep the sections in the canonical order: config-version, [server], [workshop], [local], [tools], [[dominion]], [[endpoint]], [[model]], [[local_model]], [[stt_model]], [[profile]]. The order minimizes merge noise when two people edit the file.

Keep secrets out of the file

Reference environment variables in string values with ${VAR} syntax:

api_key = "${OPENAI_API_KEY}"

A literal dollar sign is written $$. Interpolation runs only on string values, after the TOML is parsed, so a variable reference inside a comment or a key is never expanded. An unclosed ${...} fails the load. A reference to an unset variable fails the load with a distinct error that names the variable.

At startup the gateway loads the config-sibling .env file into the process environment before it reads the config. Variables already set in the environment win. A missing variable surfaces later as the unresolved ${VAR} error.

Secrets never serialize. When the gateway renders the configuration, every secret field shows *** instead of credential material. You can view the running configuration rendered as JSON in TOML shape, and you can list which config fields reference each ${VAR} variable; the values are never exposed.

Validation never lets a bad file load

A configuration never loads without passing validation. Unknown keys in any section are rejected, never ignored. Removed layout features, such as include chains or a sibling profiles directory, fail with hard-break diagnostics that name the file, the removed key, the source line, and the replacement layout. Removed legacy keys such as [queue] or an endpoint's concurrency fail at parse time. An old config cannot silently load.

You can classify a load failure into stable kinds: unreadable file, invalid TOML, malformed interpolation, unset environment variable, failed semantic check, removed layout feature, or shadow write failure.

Two field rules are worth memorizing early. A sha256 pin must be exactly 64 hexadecimal characters; uppercase and surrounding whitespace are accepted and normalized to lowercase. And a [[model]] entry without a description or a context is rejected at load.

Loopback trust

By default a caller on the gateway's own machine needs no key. With trust_loopback = true (the default, and what the first-run config writes), a request from a loopback peer that presents no credential at all is admitted on every route, the admin surface included. That is what lets curl http://127.0.0.1:8081/v1/models, the SDK with only PROMPTFORGE_GATEWAY_URL set, and the config UI on its own origin work without a key.

The trust is narrow on purpose. It applies only when the request arrives without an Authorization header: a presented-but-wrong bearer is still rejected with 401, even from loopback, so a stale key is always detected. And it applies only when the request's fetch metadata allows ambient access: no Sec-Fetch-Site header (curl, the SDK, any non-browser client) or a value of same-origin or none (the config UI, a typed URL). A page on another origin sends cross-site, and browsers never let a page strip that header, so a web page cannot use your loopback peer to reach the admin surface. A request with no peer address fails closed and needs the key.

The cost is the shared-machine case. On a machine with more than one OS account, any other account can use your gateway, including reading upstream API keys from the admin config surface. If that describes your machine, set trust_loopback = false to require the bearer key from every caller, or bind the gateway off loopback:

[server]
bind = "127.0.0.1:8081"
api_key = "${GATEWAY_KEY}"
trust_loopback = false

[server] is process-owned, so a change to trust_loopback takes effect on the next restart.

The loopback wall

The admin config endpoints sit behind a loopback wall in every build. A non-loopback peer gets 403 before bearer auth even runs. The wall covers config read and write, the env file, pending state, apply and revert, orphans, system metrics, model info, chat templates, the Hugging Face proxy, profile create and delete, and reveal. The wall fails closed: a request with no peer address is refused. Loopback trust adds a rule to authentication; it removes no wall.

Derived addresses

An unspecified bind IP such as 0.0.0.0 or :: becomes the matching loopback address in derived client URLs. Same-machine consumers, including the workshop, always get a dialable URL.

Remote Models and Endpoints

This chapter teaches you to declare remote backends and the models they serve. You will learn endpoint entries, model entries, and the catalog your callers see. Remote models are the simplest way to get the gateway serving, so they come first.

Declare an endpoint

A remote backend is a [[endpoint]] entry. Start with one:

[[endpoint]]
id = "openai"
protocol = "openai"
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"

Each entry has an id, a protocol of openai, a base_url, an api_key, and an optional dominion binding. A trailing slash on the base URL is trimmed. Endpoint ids must be non-empty and unique. Each base_url must be an absolute http or https URL with a host; values like not-a-url or ftp://example.com fail validation.

Declare a model

A remote model is a [[model]] entry that maps a public name to the alias the backend knows:

[[model]]
name = "gpt-5"
kind = "chat"
description = "GPT-5 via OpenAI"
context = 272000
thinking = "switchable"
upstream = "gpt-5"
endpoints = ["openai"]

Each entry has a name, a kind, a description, a context size, a thinking mode, an upstream alias, a list of endpoints, an optional default_max_tokens, and an optional tool_dialect. The upstream alias is the string the backend knows the model by.

Every remote model must list at least one endpoint, and every endpoint it names must be defined. Model names must be unique across remote and local models, so one name always refers to one model.

Kinds and thinking modes

Every model declares a kind: chat, embedding, classifier, or speech. The kind scopes which fields are meaningful. Chat-only fields such as thinking and default_max_tokens are rejected for non-chat kinds at load time.

Record each chat model's thinking behavior as never, always, or switchable. Switchable means the client may toggle thinking per request.

Tool dialects

The default openai tool dialect forwards tool definitions to the backend verbatim. For a backend without a native tool array, set the emulating dialect on a chat model:

tool_dialect = "gemma3_tool_code"

With this dialect the gateway injects a tool guide into the system prompt, strips the tool fields from the outgoing request, and parses tool fences from the reply.

You can advertise per-model capabilities that surface verbatim on GET /v1/models, so clients can shape requests before sending them:

[model.capabilities]
max_output = 16384
default_temperature = 1.0
images = true
parallel_tool_calls = true
effort_levels = ["low", "medium", "high"]
default_effort = "medium"
adaptive_thinking = true

The capability fields are max_output, default_temperature, images, parallel_tool_calls, effort_levels, default_effort, and adaptive_thinking. They obey cross-field rules at load time. A default_effort without effort_levels fails. A default_effort not listed in effort_levels fails. Effort fields fail when thinking is never. A max_output larger than context fails; an exact fit passes.

Enumerated fields accept a fixed spelling vocabulary. Use the spellings verbatim: protocol openai; thinking never, always, or switchable; tool_dialect openai or gemma3_tool_code; model kind chat, embedding, classifier, or speech.

What the caller sees

Callers observe the catalog at GET /v1/models:

curl -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/v1/models

Each configured model is listed with its caller-facing id, its workload kind, its description, its context window size, its thinking mode, and its capability metadata.

When a caller sends a chat, embedding, or rerank request, the gateway forwards it to the backend paths chat/completions, embeddings, or rerank relative to the configured base URL. The public model name is rewritten to the upstream alias. The caller's bearer token is never sent upstream.

Local Models

This chapter teaches you to run models on your own machine through the gateway. You will learn to declare a local model, how the gateway provisions and verifies it, and how the managed child processes behave. Local models share the gateway's OpenAI routing with remote models, so everything you learned about the catalog still applies.

Declare a local model

A gateway-served model is a [[local_model]] entry. Start with the smallest useful declaration:

[[local_model]]
name = "qwen3-local"
kind = "chat"
description = "Qwen 3 8B, local"
source = "https://huggingface.co/qwen/qwen3-8b/resolve/main/model.gguf"
sha256 = "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
vram_gb = 6.0
context = 8192
thinking = "switchable"

Each entry has a name, a kind, a description, a source, and sizing and serving knobs. The source is an https URL or a local GGUF path. The knobs and their defaults are: parallel default 1, vram_gb, context, thinking, gpu_layers default 99, flash_attention default true, cache_type_k default q8_0, cache_type_v default q4_0, and n_predict default 8192. A local model may also bind to a local dominion with the optional dominion key, and every model bound to the same dominion shares that dominion's concurrency limit. The gateway renders the child's launch flags directly from these knobs: the context size, the generation ceiling, the parallelism, the KV cache types, the GPU layers, the flash attention, and the chat template file.

A model downloaded from an https URL must be pinned by a sha256 digest. A plaintext http source is rejected, even with a valid pin. A local filesystem path may be unpinned, and the path may use ~ expansion. The pin is verified after download and on every cache hit.

The cache directory

Set [local].cache_dir for GGUF files and the pinned llama-server install:

[local]
cache_dir = "~/.promptforge"

The default is ~/.promptforge, or %USERPROFILE%\.promptforge on Windows, where the location inherits the per-user ACL. Models land in <cache_dir>/models, keyed by a hash of the full source URL, so two distinct URLs that share a filename never collide. The llama.cpp runtime installs in <cache_dir>/llama.cpp. The whisper.cpp speech runtime installs in <cache_dir>/whisper.cpp, with one directory per pinned build, such as b4938-windows-x86_64 or b4938-linux-x86_64-cuda. A CUDA build takes about 1.2 GB of cache on Windows and 1.7 GB on Linux, counting the downloaded archive the cache keeps.

On Windows x86-64 you can pick the llama-server build with [local].llama_backend: auto, cuda-blackwell, cuda, or vulkan. The auto setting picks from the machine's GPUs. You can also force an explicit llama-server executable with [local].llama_server_path; it wins over the PROMPTFORGE_LLAMA_SERVER environment variable and the managed download.

What runs underneath

Local inference runs on a pinned llama-server build, b10082. The gateway prefers GPU-enabled archives per platform: Vulkan on Windows and Linux, Metal on macOS. The gateway never compiles native dependencies at runtime; it downloads, verifies, stages, and launches pinned archives. A completed runtime install records its archive pins and a tree digest in a marker file, and a valid install skips re-extraction on later starts.

The gateway runs one managed llama-server child per [[local_model]] in the boot profile's checklist. The set of children is fixed for the process lifetime: it is decided by the profile selected at boot, and changing it means selecting a profile or editing the local catalog and restarting. Children get supervised respawn and deterministic teardown at shutdown. Staged CUDA bundle directories are prepended to the child process's PATH only; the gateway's own environment is never mutated. Local models appear to clients as ordinary routed models under their configured names.

A local model's kind selects the child's serving mode: embedding models serve embeddings, and classifier models serve reranking. A speech kind has no local serving mode and is refused at launch: local speech models are not yet supported. The parallel key sets both the child's concurrency and its admission limit. The thinking setting changes the child's sampling preset: thinking models sample at temperature 1.0 and top-p 0.95, while non-thinking models run with reasoning switched off and sample at 0.7 and 0.8.

Chat templates

A local chat model needs a chat template. The gateway resolves one through a fixed precedence:

  1. An explicit chat_template_file path.
  2. A chat_template_file = "builtin:<family>" setting.
  3. A known-override match.
  4. The GGUF embedded template.

A model with no usable template refuses to launch, and the error names the model and the fix. The bundled catalog has twelve template families: ChatML, Llama 3, Llama 3.1, Qwen 2.5, Qwen 3, Gemma 3, Gemma 4, Mistral, Phi 3, Phi 4, GPT OSS, and Zephyr. Family names accept documented aliases, and case and surrounding whitespace are ignored. The gateway also recognizes 181 revision-pinned Hugging Face repository IDs and maps each to its family automatically. Models with a known-broken embedded template are silently repaired with a bundled corrected template. The configuration UI can show the effective template source and a plain-language reason before the model is downloaded.

Reading the GGUF header

The gateway reads the architecture, the layer count, the parameter count, and the embedded chat template straight from each GGUF header, without loading tensor data. A malformed or hostile GGUF is rejected with a typed error instead of a crash or an unbounded read.

Companion artifacts

Attach a speculative-decoding drafter to a chat model with a [local_model.speculative] sub-table:

[local_model.speculative]
type = "draft-mtp"
source = "https://huggingface.co/qwen/qwen3-8b/resolve/main/drafter.gguf"
sha256 = "abcdef0123456789abcdef0123456789abcdef0123456789abcdef0123456789"
draft_max = 4

The only type is draft-mtp, and draft_max is bounded to 1 through 16. Attach a multimodal projector with a [local_model.multimodal_projector] sub-table that sets a source and a sha256 pin; a model with a projector accepts image inputs.

Companion artifacts follow the main-model source rule: an https URL must be pinned, a local path may be unpinned, and plaintext http and empty sources are rejected. Companions on a non-chat model kind fail validation. Companions are provisioned and pin-verified before the child launches, and any failure aborts the launch.

Downloads and verification

Artifact downloads are bounded. The connect timeout is 30 seconds, the whole-request ceiling is 2 hours, and a single artifact is capped at 256 GiB. Cache lookups refuse path traversal and absolute paths before any file is read, so a crafted model path cannot escape the cache root. An interrupted download resumes from the partial file's offset when the source URL still matches. A partial download from a different source restarts from zero. A pin mismatch on a cached blob is repaired by re-downloading. Once a blob passes its pin check, later runs skip re-hashing. When a runtime download fails and an older verified install exists, the gateway uses the cached install with a warning. Bundled runtime assets, including the chat templates, are written into the cache only after a SHA-256 verification pass, and a cached copy whose bytes have drifted is repaired from the bundled copy.

Authenticate gated Hugging Face downloads with the HF_TOKEN or HUGGING_FACE_HUB_TOKEN environment variable. The token is attached only to HTTPS requests to huggingface.co and its subdomains.

Models downloaded from Hugging Face get a metadata sidecar file beside the cached GGUF. The sidecar records the source URL, the fetch time, the chat template, and an optional model card excerpt.

Startup and supervision

Startup reports a structured progress tree under the boot load's stages: loading-profile, downloading-models, starting-models, and loading-speech. One subtree covers the llama-server runtime, and each local model gets download, verify, and ready stages. Progress renders as tracing log lines on every stream and on the live progress stream.

Startup is best-effort. Every model that launched keeps serving, and each model that failed is reported by name with its error. One bad model never blocks the rest. Startup failures are classified as plausibly transient or permanent, and the classification annotates the respawn diagnostics you see in the logs.

Each child server listens only on loopback, and each launch uses a fresh random alias and bearer key, so other processes on the machine cannot reach the local endpoint. Responses still report your configured model name. Startup waits up to 180 seconds for a child to become ready, and a port collision retries on a fresh port up to four times.

A child that dies is transparently respawned on the same port, alias, and key, with a 3 second cooldown between attempts so a crash loop cannot storm. Only transport-level deaths trigger a respawn, and an explicitly shut-down child is never respawned. Shutdown cancels and terminates even an in-flight respawn. Teardown is bounded to 5 seconds, so shutdown never hangs.

Child stdout and stderr are captured into bounded tails with the credential redacted. You can pull the tails per model as diagnostics; they include the CUDA device report and per-model GPU offload lines. At startup the gateway also probes each local chat model to detect native tool-call support and picks the correct tool-calling dialect from the evidence.

On Windows the child runs at below-normal priority with no console window, so weight loading and inference yield to interactive desktop use.

What callers can do

Local chat completions accept deterministic sampling parameters such as temperature, seed, presence_penalty, and max_tokens. They also accept tool definitions and, with a projector, image inputs. Chat completions on a speculative-drafted model expose decoding statistics in the response's timings extension: draft_n and draft_n_accepted.

Speech-to-Text

This chapter teaches you the gateway's transcription surface: how to declare speech models, how the interim and final roles work together, and how to use batch and Realtime transcription. Speech builds on local models, because speech models are provisioned and cached the same way.

Declare speech models

A speech-to-text model is a [[stt_model]] entry:

[[stt_model]]
name = "whisper-base-en"
role = "interim"
source = "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.bin"
sha256 = "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
vram_gb = 1.0

Each entry has a name, a role of interim or final, a source, an optional sha256 pin, a vram_gb estimate, and an optional dominion binding. The interim role transcribes while a take is still recording. The final role crystallizes completed audio.

A profile may select at most one interim and one final STT model. A final model requires an interim partner. Interim-only is a supported degraded mode. You can restore a built-in recommended pair at any time: whisper-base-en for interim and whisper-small-en for final, both with canonical whisper.cpp URLs and SHA-256 pins.

Tune push-to-talk capture

Tune the pipeline in the optional [stt] section:

[stt]
window_seconds = 15
interval_ms = 500
vocabulary = ["MCP", "GGUF", "Lua"]

The window_seconds key sets the seconds of trailing audio transcribed per pass (default 15), and interval_ms sets the milliseconds between passes (default 500). Each must be at least 1; a zero value fails startup. The vocabulary lists domain terms that bias both transcription workers toward those terms. An empty list disables biasing. A vocabulary that exceeds the model's prompt budget is truncated, and a warning is logged. The whisper_backend key chooses the whisper.cpp runtime build: auto (the default), cpu, or cuda, as The runtime describes.

Version 2 accepts only the canonical [stt] section. Legacy [workshop.stt] input is rejected as an unknown workshop field whether it appears alone or beside [stt], and saved configuration uses only [stt].

Batch transcription

With the default-on stt feature the gateway serves OpenAI-compatible audio transcription at POST /v1/audio/transcriptions. The multipart form accepts file, model, language, prompt, temperature, response_format, and the repeated field timestamp_granularities[].

Uploads are capped at 25 MiB; an over-limit upload is answered with "audio file exceeds the 25 MiB limit". Only 16 kHz mono WAV audio is accepted. Other sample rates or channel counts are rejected with a message naming what was received.

Two response shapes are offered. The json shape returns text only. The verbose_json shape returns task, language, duration, text, segments, and words. Segment timestamps are on by default, and word timestamps are always empty. A transcription request for a model not loaded in the active profile is rejected as an unknown model. A caller-supplied temperature must be a finite non-negative number. The prompt hint is accepted but ignored by the current English whisper workers.

The runtime

Speech-to-text runs on a separately pinned whisper.cpp library bundle, b4938. A library that does not match the pinned layout fails to load, and only 64-bit targets are supported. The gateway serves first and loads speech second: after the listener is bound, the queued boot command downloads and verifies the model artifacts and the runtime into the configured cache directory, with progress on the status and progress endpoints. Each model file is prewarmed and then loaded, with progress per model. Speech routes answer as unavailable until the load completes, and the model catalog advertises speech models only once the speech engine is ready.

Windows x86-64 and Linux x86-64 each have two builds of the runtime, one for the CPU and one for CUDA. The [stt].whisper_backend setting chooses between them: auto, cpu, or cuda. The default auto downloads the CUDA build when nvidia-smi reports an NVIDIA GPU, and the CPU build when it reports none or cannot run. On Linux the CUDA build needs NVIDIA driver 570 or later, the floor for CUDA 12.8, so auto also takes the CPU build when nvidia-smi reports an older driver or a version it cannot read. The Linux CUDA build also needs a C++ runtime from GCC 12 or later: auto takes it only where the machine's libstdc++.so.6 defines GLIBCXX_3.4.30, so RHEL 9, its rebuilds, and Amazon Linux 2023 keep the CPU build, and an explicit cuda there fails to load. On both platforms nvidia-smi ignores CUDA_VISIBLE_DEVICES, so auto reads it as well: a value that hides every GPU from CUDA, such as -1, an empty value, or a device index past the last GPU, counts as no GPU and takes the CPU build. On Linux a first entry that is a GPU- or MIG- identifier counts as a visible GPU without being matched against the machine's GPUs, so an identifier that names no GPU, or an abbreviated one that names several, still takes the CUDA build. On Windows auto counts a first entry that is not a device index, a GPU- or MIG- identifier included, as hiding every GPU and takes the CPU build; set cuda to restore GPU decoding for a valid identifier. The cpu and cuda settings download their own build without probing. An explicit cuda is honored below the Linux floor, and loading that build can end the gateway. On Windows with every GPU hidden, an explicit cuda decodes on the CPU and ends the gateway at a graceful stop. Every other platform has one build and ignores the setting.

On Windows the CUDA build needs NVIDIA driver 580 or later, the floor for CUDA 13, and it carries native code only for GPUs of compute capability 8.6, 8.9, 12.0, and 12.1. So auto takes the CPU build when nvidia-smi reports an older driver, a version it cannot read, or any GPU of another compute capability. An explicit cuda is honored on such a machine, with two outcomes: a GPU without native code under a driver older than CUDA 13.3's ends the gateway at the first transcription, and where CUDA finds no usable device the build decodes on the CPU and a graceful stop then crashes the gateway instead of exiting cleanly. Set auto or cpu to recover from either. Where the driver can compile it, a GPU without native code, such as a Linux GPU at compute capability 7.5, 8.0, or 9.0, compiles the build's PTX at its first decode, which took about 20 seconds with an empty driver JIT cache.

Every x86-64 build, whatever the setting, uses the SSE4.2, AVX, AVX2, BMI2, FMA, and F16C instructions. On a CPU that lacks any of them, provisioning the whisper library fails with an error that names the build, the required extensions, and the missing ones, and speech stays unavailable while the gateway keeps serving.

STT startup failures are named by stage: opening the artifact store, provisioning the whisper library, provisioning a named model, a missing interim partner, an unsupported role, or speech engine load. Library load failures name the failing path or symbol in the logs. A failed boot load never stops the gateway and is never retried in-process: speech stays unavailable, the failed boot command shows on the queue and progress surfaces, and a restart is the recovery. A stop during the speech load cancels the whisper library and speech model downloads at their next chunk, and the next start resumes them; an extraction already under way finishes first. A graceful stop aborts running transcriptions after their current encoder pass or decoder step, their requests fail, and speech retires within the gateway's existing shutdown bounds. A stop while a speech model is still loading can outlast those bounds, and the gateway then exits without retiring speech.

How a take is transcribed

A recorded take is split into speech segments at silence boundaries. A segment closes only after 2 seconds of trailing silence, so sentence-internal pauses survive. Speech bursts shorter than 250 ms are discarded as clicks. Audio quieter than -60 dBFS is treated as silence and never sent to the model, and fragments shorter than half a second are gated out.

With a final model configured, completed speech segments are re-transcribed in the background while the take still records, and each segment's text is reported as it finishes. Without a final model, the stop falls back to the interim model. Silent or very short fragments are skipped so the model does not invent text for them. Transcription is pinned to English, and translation is disabled.

Realtime transcription

The gateway serves authenticated Realtime transcription at WS /v1/realtime?intent=transcription. The query is exact: missing, duplicate, malformed, unsupported, or additional parameters are rejected before upgrade. Native clients may omit Origin; browser clients must send an HTTP loopback Origin.

The server creates a transcription session for the logical model realtime-transcribe. Clients may send session.update, input_audio_buffer.append, input_audio_buffer.clear, and input_audio_buffer.commit. Audio appends are canonical Base64 containing signed little-endian mono PCM16 at 24 kHz. The gateway preserves an odd trailing byte across appends, continuously resamples to 16 kHz, flushes the resampler on commit, and resets the whole input on clear.

Only null noise reduction and turn detection are accepted. Session updates may change the transcription prompt and negotiate the PromptForge extension item.input_audio_transcription.hypothesis. Standard clients receive OpenAI-shaped session, item, transcription delta, completed, failed, and error events. Extension clients also receive revisioned replacement snapshots with the complete transcript and its finalized, agreed, and tentative regions; completion remains authoritative.

A continuous Realtime recording remains one provisional item and one logical take for arbitrary duration. Continuous speech forces an accurate boundary every 10 seconds. Each successor repeats the preceding 8 seconds for text reconciliation, so every accurate decode contains at most 18 seconds. Finalized text and exact lifetime duration survive source-buffer compaction, and every hypothesis remains a complete replacement snapshot for that same item.

The 30-second limit is retained ownership, not recording duration. It includes resident, queued, and actively decoding 16 kHz PCM. Arbitrary-duration capture therefore requires steady-state final throughput at least equal to capture. If decoding falls behind until that retained budget is exhausted, the append receives too_much_unfinalized_audio; previously accepted audio and text remain valid and may still be committed. One append decodes to at most 15 MiB, committed audio must be at least 100 ms, one connection may have four committed items finalizing concurrently, and the service admits at most eight Realtime sessions. Queue and capacity overloads return explicit errors instead of waiting without limit.

The desktop Workshop exposes the same /v1/realtime path on its own origin. Its server authenticates the fixed upstream target and relays payloads without parsing them, so the webview never receives the gateway credential.

Speech loads exactly once per process, from the profile active at boot. Switching the active profile or applying a new configuration persists a changed speech selection but never loads, reloads, or unloads the running speech engine; the new selection takes effect on the next start. The configuration UI raises a restart toast when an apply changes the speech tuning, the speech model catalog, or the active profile's speech membership.

Speech Synthesis

This chapter teaches you the gateway's speech synthesis surface: how to declare a speech model, how to call the synthesis route, and how to enumerate voices. Synthesis builds on remote models, because the gateway routes speech to remote providers only; a [[local_model]] with kind = "speech" is refused at launch.

Declare a speech model

A speech synthesis model is an ordinary [[model]] entry with kind = "speech", backed by an ordinary [[endpoint]]:

[[endpoint]]
id = "together"
protocol = "openai"
base_url = "https://api.together.xyz/v1"
api_key = "${TOGETHER_API_KEY}"

[[model]]
name = "orpheus"
kind = "speech"
description = "Orpheus 3B conversational speech synthesis"
upstream = "canopylabs/orpheus-3b-0.1-ft"
endpoints = ["together"]
context = 8192
voices = ["tara", "leah", "jess", "leo", "dan", "mia", "zac", "zoe"]

The entry takes the usual remote-model fields, and the kind scopes which of them are meaningful. Chat-only fields such as thinking, the effort knobs, default_max_tokens, and tool_dialect are rejected on a speech model at load time. The speech-only voices list declares the voices the model offers: setting it on any other kind fails at load, entries must be non-empty and unique, and an empty or omitted list means the model exposes no fixed voice list, so the route accepts any voice name. The catalog advertises the kind and the voice list verbatim on GET /v1/models, so clients can shape requests before sending them.

Synthesize speech

The gateway serves OpenAI-shaped speech synthesis at POST /v1/audio/speech:

curl -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "orpheus", "input": "The quick brown fox.", "voice": "tara"}' \
  -o speech.mp3

The request takes model and input (both required; the input is non-empty and capped at 4096 characters), voice (required; a plain name or the OpenAI object form {"id": "tara"}), and four optional fields: response_format from the closed set mp3, opus, aac, flac, wav, pcm; speed between 0.25 and 4.0; instructions, the gpt-4o-mini-tts dialect's style-control string; and stream_format, sse or audio. An omitted response_format resolves to mp3 before the request leaves the gateway: OpenAI defaults to mp3 while Together defaults to wav, so the pin lives in the wire type and every forwarded body includes it. Fields the gateway does not name pass through to the provider verbatim, so provider extras such as Together's sample_rate go out in the same request, and angle-bracket emotion tags such as <laugh> in the input reach the provider untouched.

Authentication runs before the body is parsed, so a bad key earns 401 even for a malformed body. Shape failures earn 400: an empty or over-cap input or an out-of-range speed as malformed_request, a non-speech model as kind_mismatch, and a voice outside the model's declared list as invalid_voice naming the valid voices, judged before queue admission so the rejection never burns a queue slot. A full queue earns 503 with code queue_full. An upstream 429 comes back as 429 with code upstream_rate_limited and an upstream 503 as 503 with code upstream_unavailable, so an OpenAI client sees a retryable rate-limit or server error rather than a generic failure.

The response is the provider's audio bytes streamed through unread: the gateway forwards the upstream Content-Type (when missing or invalid, the framing selector first, so an SSE stream is labeled text/event-stream and never an audio type, then the requested format's MIME type), never sets Content-Length, and emits Transfer-Encoding: chunked. No JSON error can follow 200 plus audio bytes, so a mid-stream upstream failure surfaces as a truncated body and the client's read fails.

List voices

GET /v1/audio/voices answers the union of the active profile's speech voices, deduplicated and sorted:

{"voices": [{"id": "dan", "name": "dan"}, {"id": "jess", "name": "jess"}, {"id": "leah", "name": "leah"}]}

Each entry is an object with id first and name mirroring it, because the catalog configures voices as bare strings with no separate display name. OpenAI has no voice-list route; the OpenAI-compatible ecosystem converged on this one, and clients such as Open WebUI read the id key, so the entry shape is a compatibility surface. Tolerant clients also accept the plain-string form some servers answer with. A profile with no speech models returns an empty list, and non-speech models contribute nothing.

The stream_format caveat

stream_format = "sse" is OpenAI's selector for event-stream framing, and the gateway forwards it verbatim like any other field. Provider dialects differ: Together spells its streaming mode stream=true, and only with response_format = "raw", which the wire enum rejects, so Together SSE cannot be requested in phase 1 and Together speech is non-streaming. The forwarding is forward-looking, aimed at SSE-capable OpenAI-compatible providers, and the gateway does no reframing or decoding: whatever framing the provider answers with passes through untouched, so a client that receives an event stream owns decoding it itself.

Profiles and Selection

This chapter teaches you profiles: named checklists that decide which local models the gateway loads, how the selection is stored, and how you change it. Profiles are how one config file serves a work machine, a travel laptop, and a demo box without editing a single model entry.

Define a profile

A profile is a [[profile]] entry that owns only a name and a models list:

[[profile]]
name = "work"
models = ["qwen3-local", "whisper-base-en", "whisper-small-en"]

[[profile]]
name = "travel"
models = ["qwen3-local"]

A profile is a checklist of local and speech-to-text models. Membership alone decides which local models spawn and which speech models load. Every name a profile lists must be a [[local_model]] or [[stt_model]] entry, and each must exist exactly once. Naming a remote [[model]] in a profile fails validation with an error saying the model is remote: remote models are never gated by a profile, because every [[model]] in the catalog routes all the time. Duplicate profile names and duplicate members also fail validation.

Profile names must be a single safe path component: no surrounding whitespace, not empty, not . or .., and no path separators. One spelling works in URLs, state files, and labels.

Every profile is validated at load. Names are unique and legal, every listed model exists, each profile selects at most one interim and one final speech model, and the local and speech subsets are checked against dominion VRAM budgets. The gateway never boots into an invalid profile.

Where the selection lives

The selected profile lives in a sibling state file, not in the config. A gateway.toml maps to a gateway.state.toml holding one canonical key:

active_profile = "work"

The selection survives restarts. An absent state file is the persisted form of "no profile": the gateway boots, serves every remote model, and loads no local or speech models. "No profile" is a selectable state, not an error.

At startup the profile is chosen by precedence: the --profile command-line flag, then the PROMPTFORGE_PROFILE environment variable, then the sibling state file. The flag and the variable are ephemeral; they never write the state file. With none set, the gateway boots with no profile.

A state file naming a profile the config no longer defines does not stop the boot. The gateway logs a warning naming the stale value and the defined profiles, then boots with no profile. The stale name stays in the state file until you select something else, and the configuration UI shows it as a stale selection. A --profile flag or PROMPTFORGE_PROFILE value naming an undefined profile is still a startup error, because an operator typed it for this run.

The local model set is fixed at boot

The gateway loads its local models once, at boot, from the profile it started with. After the listener is bound, one boot command downloads the profile's local model artifacts, spawns the llama-server children, publishes each into the routing table as it becomes ready, and performs the process's one speech engine load. While a local model is still downloading or spawning, a request for it gets 503 with code model_loading and Retry-After: 5, GET /admin/status lists it under loading_models, and GET /v1/models lists only routable models. When only some local models start, the boot reports which loaded and which failed, and the ones that loaded keep serving.

Nothing after boot changes the set of local models. There is no live switch, no drain of in-flight requests, and no stop-and-spawn of children while the gateway serves. Remote models are the exception: every [[model]] routes from boot, and an applied edit to the remote catalog reloads routing live, as the next chapter explains.

Select a profile

Select a profile over HTTP:

curl -X POST -H "Authorization: Bearer $GATEWAY_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "travel"}' \
  http://127.0.0.1:8081/admin/switch-profile

The request body holds name: a profile name, or null to select no profile. The gateway checks a named profile against the loaded catalog and refuses an undefined name with the list of defined profiles. It then writes the state file (or deletes it for null) and answers plain JSON:

{"profile": "travel", "restart_required": true}

restart_required is true when the selection differs from the profile the process is running. The selection persists at once; the running gateway keeps serving its boot profile until it restarts. Selecting the profile that is already running answers restart_required: false and changes nothing. Selection uses the in-memory catalog; the config file is never re-read from disk.

Restart the gateway to load the selection. A gateway the Workshop supervises is restarted by the Workshop when you pick a profile from its Model menu; a gateway you run yourself restarts by hand, and the configuration UI shows a banner reading "Restart the gateway to apply these changes." until the new process comes up.

The selection is not part of the config edit surface. PUT /admin/config refuses a document that sets active_profile, and GET /admin/config-dirty never reports it. GET /admin/config-pending reports the persisted selection under profile.active_profile, read from the real state file, so a client can show a selection that differs from the running profile or names a profile the config no longer defines.

Dominions and Queues

This chapter teaches you dominions: named compute pools that cap concurrency, park or reject excess callers, and schedule waiting clients fairly. Dominions are how you keep one busy model from starving the rest.

Declare a dominion

A dominion is a [[dominion]] entry:

[[dominion]]
id = "pool-r"
kind = "remote"
max_concurrency = 4
max_queue = 100
policy = "queue"
fair_scheduling = true

Each entry has an id, a kind of remote or local, a max_concurrency, a max_queue defaulting to 100, a policy of queue or reject defaulting to queue, a fair_scheduling flag defaulting to true, and a vram_gb budget for local pools.

Bind an endpoint or a local model to a dominion by name:

[[endpoint]]
id = "openai"
protocol = "openai"
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"
dominion = "pool-r"

Endpoints bind to remote dominions, and local models bind to local dominions. A wrong-kind or undefined binding is rejected. An endpoint or local model without a dominion binding is unlimited; it behaves as when no cap is set at all.

Budget VRAM

A local dominion can declare a vram_gb budget, and each profile's selected models must fit within it:

[[dominion]]
id = "gpu0"
kind = "local"
max_concurrency = 2
vram_gb = 24

An overflow fails validation with an error naming the dominion and the excess. Fractional estimates such as 1.22 are accepted. Zero, negative, NaN, and infinite estimates fail.

Choose a full-capacity policy

The default queue policy parks callers up to the depth limit. The reject policy turns the caller away immediately, and the gateway answers 429:

policy = "reject"

You can distinguish admission failures by status code. A full waiting queue answers 503 with code queue_full. A fail-fast rejection answers 429 with code queue_rejected. A queue torn down while the caller waited reports the queue as unavailable.

Schedule fairly

Turn on fair scheduling so waiting callers are served in per-client round-robin order, keyed by the X-PromptForge-Client request header:

fair_scheduling = true

The header is a self-asserted hint. Values over 64 bytes or outside the alphanumeric, dash, underscore, dot, and colon charset fold into the shared default bucket. At most 32 distinct client labels are tracked.

How slots behave

A streaming request holds its dominion concurrency slot for the stream's whole lifetime, so a second request waits until the first stream ends. A cancelled queued request frees its waiting slot, and capacity recovers without a restart.

Editing Configuration Safely

This chapter teaches you the safe-edit surface: how the gateway stages edits in shadow files, how you preview and apply them, and how you recover when an edit is wrong. Editing through this surface means a bad config can never take down a running gateway.

Shadow files

Pending admin edits are staged in one shadow file, gateway.toml.next, beside the real config. No save touches a real file until promotion.

Stage a full config edit with PUT /admin/config. The request takes the same JSON shape that GET /admin/config returns. Secrets left as the redacted marker *** are restored from the current values, and a marker with no existing value fails validation. The merged result is validated like a real load before any shadow is written. The reply names the shadow file that was written.

The profile selection is not a config key. A document that sets active_profile is refused with a validation error pointing you at POST /admin/switch-profile, which the previous chapter covers.

Preview before you apply

Preview the merged pending configuration with secrets still redacted:

curl -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/admin/config-pending

The pending envelope also reports the persisted profile selection under profile.active_profile, read from the real gateway.state.toml: null when no profile is selected, otherwise the stored name even when the running profile differs or the config no longer defines it.

Poll a cheap dirty report of pending shadow files and changed sections:

curl -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/admin/config-dirty

Apply

Applying a pending edit is an explicit promote step:

curl -X POST -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/admin/config-apply

The real file is replaced atomically. On platforms where rename cannot overwrite, a backup-and-restore fallback preserves the old file. The reply reports applied, reloaded, and restart_required.

What an apply does depends on which sections changed. The remote-facing sections - [[model]], [[endpoint]], [[dominion]], and [tools] - reload live: the gateway rebuilds the remote routing table from the applied config, keeps the running local models under it, and swaps the routing table in one write. Nothing drains, nothing stops, and no child process starts. The boot-owned sections - [server], [workshop], [[profile]], [[local_model]], [[stt_model]], and [stt] - promote to disk but take effect at the next start, so the reply sets restart_required: true; an env shadow does the same. The gateway's local model set is fixed for the process lifetime, so an edit that adds, removes, or changes a local or speech model, or changes a profile's checklist, always needs a restart. One apply can do both: reload the remote catalog now and report a restart for the rest.

An apply that changes the config runs as a command on the gateway's command queue, the same queue that runs the boot load. The queue is one serialized pending deque with no fixed capacity: debounce decides what stays pending, and a single worker runs the surviving commands in order. The request waits for the command's outcome, so the call above still returns when the apply is done. While the command runs, GET /admin/status reports it as the active command named apply-config, and its one applying-config stage streams on the live progress stream; the config UI's Apply overlay follows it and shows a Cancel button. POST /admin/queue/cancel stops it; the request then answers 503 with error code apply_cancelled. An apply requested while the boot load is still running waits behind it, so a reload never races the boot's publication of local models. An apply that touches only the env file, or only boot-owned sections, needs no reload and runs inline without a command.

Promotion happens at the end. The shadow is read into memory when the apply is requested, the new routing table is built, and only then are the captured bytes written to the real file and the shadow removed. A cancelled or failed apply therefore promotes nothing: the shadow stays on disk, the pending count stays where it was, and the next Apply runs the whole thing again. A save that lands while an apply is in flight is kept as the next pending change, never silently lost and never half-applied.

Revert

Discard every staged edit without touching the real files:

curl -X POST -H "Authorization: Bearer $GATEWAY_KEY" http://127.0.0.1:8081/admin/config-revert

The reply names the deleted shadow files. Deleting the shadows is the whole revert. A revert never touches the profile selection, because the selection is never staged.

The .env file

Read and stage the gateway's global .env file over the same surface. GET /admin/env returns the file with plaintext values and shows which config fields reference each variable. PUT /admin/env stages a .env.next shadow that takes effect after restart. Variable names must use letters, digits, and underscores, and must not start with a digit. Values must round-trip through the dotenv parser.

Failure behavior

You are protected from half-applied state. Saves, revert, profile selection, and the apply's snapshot and commit steps serialize on one lock, and applies serialize with the boot load on the command queue. An invalid pending config is never promoted; the request fails before any command exists. A failed or cancelled apply leaves every shadow on disk for correction, retry, or revert. A revert issued during an apply cancels the apply first, so the apply's commit never writes over files you just reverted.

The Configuration UI

The gateway serves a browser UI for configuration: you reach it over HTTP, sign in with your API key when the gateway asks for one, and edit every part of the configuration through its views. The UI works through the safe-edit surface from the previous chapter, so everything you do there moves through pending shadows and Apply.

Reach the UI

The gateway serves the configuration UI at /config on its own port; there is no second listener. The UI is an optional feature you compile in with config-ui, which is on by default. GET /config redirects permanently to /config/. Five asset endpoints live under /config: the index page at /, the bundled script at /app.js, the stylesheet at /app.css, and the program icon at /icons/promptforge-icon.png with its high-DPI render at /icons/[email protected].

The UI pages need no bearer token, but every asset route answers 403 Forbidden to any peer that is not loopback. The UI is reachable only from the gateway machine itself, and the check fails closed.

Sign in

With the default trust_loopback = true, the UI opens straight into the shell: it runs on the gateway's own machine, and the gateway admits a loopback caller that presents no key. On a shared machine that same trust admits every other OS account, so an operator there sets trust_loopback = false; the UI then asks for the key. On first load without a stored key, you see a "PromptForge Gateway" sign-in card with a labeled API key password field and a Connect button. A wrong key shows "Invalid API key". An unreachable gateway shows "Gateway unreachable". The verified key is stored for the browser session. Any later 401 from the gateway clears the stored key and returns you to the key prompt.

Get oriented

You navigate six top-level views from the tab bar: Settings, Discover, Local, Remote, Profiles, and Secrets. Every view has a bookmarkable hash URL, including a specific model's detail page and a specific Settings section. An unrecognized hash is rewritten to #/local.

A connection dot in the tab bar shows whether the gateway is reachable. The tab bar shows the running UI's version as a muted label, or vdev for a development build. Notifications appear as toasts that dismiss themselves after four seconds. Destructive actions require confirmation in a modal dialog that names the target; focus lands on Cancel as the safe default, and Escape or a backdrop click cancels. Every dropdown works entirely from the keyboard, including typing the first letters of an option to jump to it. The UI is always dark, the reduced-motion system preference disables essentially all animation, and byte sizes appear in human-readable units.

The three states of an edit

Edits move through three states: unsaved edits held in the browser, saved pending shadows on the gateway, and the applied running configuration. When pending changes exist, the tab bar shows an Apply button labeled with the pending file count beside a Revert All button. When a previous session left unapplied changes, a banner offers Review, Apply, and Revert All.

Pressing Apply opens a progress overlay that follows the gateway's live progress stream until the apply finishes or fails; a remote-catalog reload shows as one applying-config stage. The overlay shows a Cancel button; pressing it stops the apply on the gateway, and the overlay reports that the apply was cancelled and your pending changes are still staged. A failed stage holds on the error message for a moment before the overlay closes. When an applied configuration requires a restart - a change to [server], [workshop], [stt], a profile, a local model, or a speech model, or an env edit - a banner reads "Restart the gateway to apply these changes." and clears itself once the gateway comes back on a new config generation. The same banner is raised by a profile selection that differs from the running profile.

Open the Review dialog to list every pending configuration change as a table of path, running value, and pending value. Secret values are never displayed.

Profiles

The tab bar shows the selected profile. Its menu lists "No profile" first and then every defined profile, with the persisted selection checked. Choosing an entry selects it on the gateway at once through POST /admin/switch-profile; the selection is not a pending change and needs no Apply. When the selection differs from the profile the gateway is running, the restart banner appears, because the gateway loads its local models once at boot. A refused selection surfaces an error toast and leaves the current selection unchanged.

In the Profiles view you edit each profile as an ordered subset of the local and speech-to-text catalog through Available and Chosen shuttle listboxes; remote models never appear, because every remote model routes regardless of profile. The listboxes support multi-select, roving focus, typeahead, selection counts, and per-pane search. The profile saves in local catalog order, not click order. A new profile starts Empty or as a Copy of an existing profile. Two pills mark the rows: Active is the profile the gateway is running, and Selected is the persisted choice when it differs. You cannot delete either. The view leads with a "No profile" row that has its own Set Active. Each row's Set Active persists that selection immediately and raises the restart banner when a restart is needed; the selected row's button reads "Selected". When the state file names a profile the config no longer defines, the view shows a Stale pill with the missing name, and Set Active on any row replaces it.

The Profiles view shows an Estimated VRAM summary that sums declared model weights. Per-dominion budget rows warn at 80 percent and error when over. KV cache grows with context length, so 20 percent headroom is recommended.

Discover

The Discover view searches Hugging Face. The search box accepts keywords, a user/repo form, or a pasted hub URL, and keystrokes collapse into one search after a 300 ms debounce. The GGUF filter chip is locked on because the gateway serves GGUF inference only. Chat is the default workload filter; the filters cover Chat, Embedding, Reranker, STT, Image, and TTS. Result rows show the publisher avatar, a parameter-count pill, compact download and like counts, and a relative updated time. Sorts are Most downloads, Trending, and Newest.

A model's GGUF files are grouped into named quantizations with exact summed byte sizes and the LFS SHA-256 for single-file quants, listed smallest first. Each quant shows a fit badge computed against the gateway's system snapshot: Fits GPU, Partial offload, CPU only, or Too large. One Recommended star marks the largest quant that fully fits free VRAM. A multi-part GGUF cannot be downloaded as one model; the button is disabled with an explaining tooltip. You can read model cards in the view, rendered as sanitized HTML so embedded scripts and event handlers cannot execute.

A Download click stages a pending model entry with the hub resolve URL, the LFS digest, and the listing size as vram_gb; the transfer happens at the next boot. Staging a discovered model also adds it to the selected profile's checklist, so the restart after Apply provisions and serves it. When no profile is selected, the model is added to the catalog alone and a toast says so; choose it in a profile to run it. The staged entry prefills a mapped built-in chat template when the server-side catalog matches the repo. An STT-filtered download stages a first-class stt_model entry with the interim role. Without a configured HF token you see a banner linking to the Secrets view instead of search results.

Local and Remote

The Local and Remote views show the gateway's own catalog subsets: Local lists your local and speech-to-text entries, and Remote lists your remote entries. STT entries show a Mic badge so you can pick them out at a glance. Filter chips narrow the list to All, Chat, or STT, a search box filters the rows after a short debounce, and a sort dropdown orders the list by Name, Size, or Kind.

Each model row shows a running-status dot, a kind badge, and capability pills. A quant badge read from the GGUF filename names the quantization, such as Q4_K_M. A model that exists only as a draft gets an "unsaved" badge until you save it.

Secrets

The Secrets view manages the one global .env file. Variables appear as masked password rows with per-row reveal and delete. Save stages a pending shadow that takes effect only after Apply plus a gateway restart. New variable names must use letters, digits, and underscores, and must not start with a digit. Each variable shows "used by" annotations naming the configuration entries that reference it. A dedicated Hugging Face card configures HF_TOKEN, and its Test Connection probes the token the running gateway holds and reports Not set, Valid, Invalid, or Connection failed.

Settings

The Settings view has seven sections: System, Gateway, Workshop, Dominions, Endpoints, Tools, and About. You land on System by default.

The System panel shows live metric tiles: CPU, RAM, VRAM with the GPU name, and disk usage with the cache path. The tiles refresh every 5 seconds, and a failed refresh keeps the last snapshot. Metric bars recolor by load: warning in the 70 to 89 percent band and danger at 90 percent or more.

The Gateway card edits the bind address, the API key, and the Trust loopback connections switch. The switch is on by default and admits callers on this machine that present no key; its help text states the cost, that on a shared machine any other OS account can then use the gateway, and turning it off requires the key from every caller. A note says the boot configuration cannot hot-reload, and changing the API key warns that the new key will be required after restart. The typed key leaves the DOM once saved. Stored secrets render as a masked readout with a Change button; leaving the input empty keeps the existing key, and an Eye toggle reveals and re-hides the secret.

The Dominions and Endpoints cards show used-by chips that count dependents, and a delete confirmation names them. A local-kind dominion reveals the vram_gb budget field, and switching the kind to remote hides it. An endpoint binds to a dominion from a dropdown offering only remote-kind dominions plus None. The endpoint protocol dropdown is locked to openai, and the endpoint API key stays redacted through saves until Change reveals the input.

The Tools section configures web search with the provider locked to Brave and the defaults documented on the card. The Storage card edits the cache directory beside live cache-drive usage, with a warning that changing the directory does not move existing files.

The About panel shows the medallion, the baked version or "dev", and the Boost Software License link. The Config UI card reports the UI as compiled in by the config-ui feature, served on the gateway's own port, loopback only, with the URL derived from the bind. The Workshop card edits the [workshop] section's one live content, the STT capture tuning - the gateway serves no workshop listener, so the section's old bind and open_browser settings are inert and stay out of the editor. Adding the tuning seeds window_seconds 15, interval_ms 500, and an empty vocabulary.

Editing a model

You edit a local model through sections for GPU, generation, source, and capabilities in the model detail view. An unconfigured optional section offers an Add button. The chat template control offers Auto, a built-in template family, or a custom .jinja path, with a read-only summary naming the effective source, the detected family, and the reason.

The model name edits inline in the detail header, and the header shows the model's status: Unsaved, Running, or Stopped. Each edited field shows a dirty dot and a per-field reset. Each saved-but-unapplied field gets a pending chip whose tooltip shows the running value. Deleting a model confirms a dialog naming the model and every affected profile, and the save removes every dangling profile reference in the same payload. A downloaded model shows its cached size and path with a Delete file action. A path source gets a reveal-in-folder button; URL sources get none. Capability pills show images and thinking mode, and the images pill is implied and locked when a multimodal projector is configured. The gpu_layers slider readout shows the GGUF layer total, and typing "Max" maps to the maximum.

The controls follow the shape of the value. Numeric settings pair a slider with a typed readout, typed values clamp to the allowed range, wide-range settings such as the context window use a logarithmic scale, and some sliders offer a rightmost "Max" detent. List-valued settings such as a model's endpoint list are edited as removable chips. Fields with a fixed choice set accept only the listed values. Boolean settings use an on/off switch. A setting can be disabled until a sibling field holds a required value, or hidden until a predicate passes, so you only see applicable controls.

Retyping a field's original value clears its unsaved edit, and you can reset one field or a whole entry. A new model entry starts as an unsaved draft, and name collisions get auto-suffixed. Every settings save sends the complete single-file configuration, so one section's save never erases another staged section.

An orphan section lists unconfigured files on disk with Adopt and Delete actions per file; Delete is disabled when the file has no verified digest. The UI shows whether a model's source file is already downloaded. On gateways built without local-model features, missing orphan and chat-template endpoints degrade to empty lists instead of breaking the UI. You can restore the recommended speech-to-text model pair, digest-pinned, over the existing STT catalog entries from the UI.

Panel mode

The configuration UI runs in two modes. Standalone mode runs in a browser tab. Panel mode embeds the UI inside the Workshop with ?mode=panel. In panel mode your API key never enters the frame; every gateway call goes through a postMessage bridge to the Workshop, and the panel only talks to a loopback workshop origin. Bridged calls fail after a 30 second reply deadline rather than hanging. Apply and Revert actions are announced to the workshop's status bar, and the workshop pushes its theme and an initial route into the embedded panel once the bridge is up.

Serving and Observing

This chapter teaches you the running gateway: the HTTP endpoints it serves, the tools it can run, and the health, logs, and observability surface you operate day to day. You already run a configured gateway with a profile and its models.

Enable the built-in web-search tool with a [tools.web_search] section:

[tools.web_search]
provider = "brave"
api_key = "${BRAVE_API_KEY}"
default_count = 10
max_count = 20
max_per_host = 2
strip_tracking = true

The provider is locked to brave. The base_url defaults to the Brave Search endpoint and must be an HTTP(S) URL. The default_count must not exceed max_count. Freshness and safesearch defaults are closed vocabularies, not free text. The gateway calls the Brave Search API at {base_url}/web/search with the configured API key sent in the X-Subscription-Token header.

Callers run a web search through POST /v1/tools/web_search. The request body takes a query and optional count, freshness, country, search_lang, safesearch, include_domains, and exclude_domains. Unknown fields are rejected. The query is trimmed, rejected when empty, and capped at 512 characters. Caller knobs are validated before any provider call: freshness must be pd, pw, pm, py, or a date range; safesearch must be off, moderate, or strict; country is a 2-letter code; the search language is a 2 or 3 letter code; each domain entry must be a bare valid domain. The count defaults to default_count and clamps into 1 through max_count. The gateway over-fetches up to three times the requested count, capped at max_count, so post-processing filters still yield enough results. Omitted freshness and safesearch fall back to the configured defaults.

Results include title, url, site_name, and extra_snippets. Result text is sanitized and capped, results are diversified by host at max_per_host, and a result whose URL is not navigable or is over 2048 characters is dropped. When strip_tracking is on, known tracking parameters such as utm_*, fbclid, gclid, mc_cid, and mc_eid are removed from result URLs. Include and exclude domain lists match the host itself or any subdomain.

When no [tools.web_search] section is configured, the route answers 404. The route exists only in builds compiled with the web-search feature. Search provider failures surface with a web_search: prefix on the error, so you can distinguish search upstream errors from other gateway errors. The search service is built from the [tools.web_search] section and is replaced live when an applied edit changes it. The provider credential never appears in logs.

The deprecated [workshop] section

The gateway and the workshop run separately: the desktop application embeds the workshop server itself. A boot config left over from an older version may still declare a [workshop] section with the inert bind and open_browser settings, which produce a deprecation warning at startup. Speech pipeline tuning belongs in [stt]; legacy [workshop.stt] input is rejected as an unknown workshop field whether it appears alone or beside [stt].

Manage the cache

Manage the blob cache through the gateway's cache routes. GET /v1/cache lists entries with source URL, path, SHA-256, and size. Only blobs with a .meta.json sidecar appear in the listing, and listing reads the sidecar metadata only; it never re-hashes the blobs. POST /v1/cache downloads a blob with an optional pin and streams progress events ending in a ready event. DELETE removes one blob by digest. Cache downloads validate the source URL and the pin before any network access. A cache download lands in the same slot layout that local model provisioning uses, so a cache download is a provisioning cache hit for the same URL, and vice versa.

GET /admin/orphans lists cache files that no [[local_model]] or [[stt_model]] declared in the catalog references, so leftovers can be adopted or deleted. GET /admin/model-info reports a GGUF file's header summary (architecture, layer count, parameter count, and chat template) without loading the model; only files inside the artifact cache can be inspected, and escaping or missing paths are refused. POST /admin/reveal opens the machine's file manager at a model or config file; reveal requests are confined three ways: loopback-only, bearer key required, and the path must canonicalize to strictly inside the artifact cache.

The gateway restricts the cache root to your own account at startup and refuses to run when it cannot, failing with a cache-not-private error.

Status, progress, and metrics

GET /admin/status reports the running profile (null when the gateway booted with no profile), the local and speech models that profile lists as model_allowlist, the models the gateway exposes, and a config generation that changes when the gateway restarts. It also reports the command queue: the active command's name, progress fraction, and start time, plus the pending commands, so the boot load and applies are visible while they run. With the STT feature it also includes generic speech facts: whether speech is configured, whether the boot-time speech engine load has completed and speech is ready, and whether its backend reports GPU acceleration. A featureless build omits the speech object. GET /admin/profiles lists the profiles in the loaded catalog.

GET /admin/progress streams every long-running operation in the process as one server-sent event stream. A fresh subscriber first receives live operations replayed, then every event. Heartbeat comment lines arrive every 15 seconds while idle.

Download progress renders as tracing log lines on every stream.

GET /admin/system reports machine metrics: CPU, RAM, the cache drive, and the first NVIDIA GPU's VRAM. The GPU field is absent, never an error, when no capable driver is present. You can also pull bounded captured stdout and stderr tails for each running local model as diagnostics.

GET /admin/chat-templates returns a bearer-authenticated catalog of chat template families, known model-to-family mappings, and each pending local model's effective template decision.

You can search Hugging Face and read model details and READMEs through the gateway's hub proxy. A missing or invalid HF_TOKEN surfaces as a distinct "set HF_TOKEN" error. Hub search queries are validated against a closed allowlist before any upstream call, and repository paths must be an exact owner/name pair of hub-legal segments.

Errors and limits

Every request failure reaches the client in the OpenAI error envelope: an object with message, type, and code under error, with a stable HTTP status. Examples: 401 unauthorized, 404 model_not_found, 400 malformed_request, 400 kind_mismatch, 429 queue_rejected, 503 queue_full, 503 model_loading, 503 partial_start, 422 config_write_rejected, and 422 model_info_error.

Outbound calls to any backend have fixed timeouts: 10 seconds to connect and 120 seconds for a whole non-streaming request. Streaming connections are bounded only by the connect timeout. Response bodies the gateway reads are capped: 64 KiB for error bodies and 4 MiB for success JSON bodies.

Malformed client requests are rejected at the boundary. An empty model name, an empty messages array, an unsupported message role, or a message with neither content nor a tool call all fail validation. Request fields the gateway does not name pass through to the backend verbatim, while the reserved keys model, messages, and stream may not be smuggled in twice. Embeddings requests accept one string or a batch of strings, with an optional encoding_format of float or base64; an empty batch is rejected. Rerank requests take a query, a document set, and an optional top_n limit; an empty query or document set is rejected.

Reading failures

The error code distinguishes a connection that never reached the provider from a mid-flight failure. The first is safe to retry; nothing was billed. The second is not safe to retry blindly. A backend's own client-error status, for example 429, passes through to the caller with code upstream_client_error instead of a generic 502. A model of the wrong kind is refused with 400 kind_mismatch before any upstream call. A request for a workload the resolved model cannot serve is rejected with 400 model_unavailable.

When the gateway recovers from a malformed tool fence in an emulated tool dialect, the response message includes a gateway_warning extension field. The turn never fails, and protocol junk never appears as final text. Streaming clients still receive tool calls from an emulated-dialect model: the gateway buffers one upstream round trip and re-emits the rewritten response as synthetic chunks, with a trailing summary chunk holding usage and timings.

A malformed upstream stream chunk is logged and skipped without ending the stream. A mid-stream transport failure ends the stream with an error. An upstream error status fails a streaming request before any chunk is delivered, returned as a JSON 502, never as a stream that dies mid-flight. A client disconnect cancels the upstream request.