Skip to content
Reference

Self-hosting

One binary, one data directory. Everything you need to run, secure and back up open-server, the free VAKYN server, on your own machine.

open-server is open source (MIT) on GitHub. It serves the same API as VAKYN MAX, so code written against one runs against the other once you change the base URL and the key. It brings its own dashboard for keys, usage and a playground, and a server page that shows its status.

Install

MethodWhere
One-line installer (curl on macOS and Linux, PowerShell on Windows)Below; it picks the right build and checks it
Prebuilt binaries (macOS Apple Silicon, Linux x64 and arm64, Windows x64)On every GitHub release; see Release archives
Docker images (CPU and CUDA)On GHCR with every release; see Docker
Build from sourceSee Building from source
# Linux and macOS (Apple Silicon)curl -fsSL https://github.com/vakyn-ai/open-server/releases/latest/download/install.sh | sh # Windows (PowerShell)irm https://github.com/vakyn-ai/open-server/releases/latest/download/install.ps1 | iex

install.sh picks the build for your system: on Linux x64 with an NVIDIA GPU and driver R570 or newer, the CUDA build for the GPU's family (from nvidia-smi); otherwise the CPU build, and it says why. It downloads it from the latest release, checks its SHA-256 and installs vakyn to ~/.local/bin (/usr/local/bin as root). On Linux with systemd it then asks Install vakyn as a service and start it? [Y/n]: yes by default on a terminal, no when there is none. install.ps1 installs vakyn.exe to %LOCALAPPDATA%\vakyn\bin and adds it to your user PATH. Running an installer again upgrades in place and restarts an installed service.

install.sh flagEnvironmentMeaning
--version 0.2.0VAKYN_VERSIONInstall this release instead of the latest (install.ps1 too)
--dir DIRVAKYN_INSTALL_DIRWhere the vakyn command goes (install.ps1 too)
--cpuVAKYN_CPU=1The CPU build even with an NVIDIA GPU
--target TVAKYN_TARGETA given archive target, e.g. linux-x64-cuda-sm90
--serviceVAKYN_SERVICE=yesInstall and start the service without asking
--no-serviceVAKYN_SERVICE=noDon't install the service
VAKYN_SERVICE_ARGSvakyn serve flags for the service, e.g. --model /srv/VAKYN-4B-UQFF
GITHUB_TOKEN / GH_TOKENDownload with this token (a private fork, or GitHub API rate limits)

Flags go after sh -s --, for example curl -fsSL …/install.sh | sh -s -- --version 0.2.0 --no-service. The Linux builds need glibc 2.35 or newer; on Alpine and other musl systems use the Docker image.

Release archives

Each release has one archive per target, named vakyn-<version>-<target>, with a .sha256 file next to it and all checksums in SHA256SUMS. Every binary includes the model engine.

TargetForEngine
linux-x64-cpuLinux x86-64 with glibc 2.35+ (Ubuntu 22.04, Debian 12 and newer)CPU
linux-arm64-cpuLinux arm64 with glibc 2.35+CPU
linux-x64-cuda-sm80NVIDIA Ampere and Ada: A100, A10, A40, RTX 30xx, RTX 40xx, L4, L40SCUDA, flash attention
linux-x64-cuda-sm90NVIDIA Hopper: H100, H200CUDA, flash attention
linux-x64-cuda-sm100NVIDIA Blackwell data center: B200, GB200CUDA, flash attention
linux-x64-cuda-sm120NVIDIA Blackwell RTX: RTX 50xx, RTX PRO 6000CUDA, flash attention
macos-arm64-metalmacOS on Apple SiliconMetal
windows-x64-cpuWindows x64 (.zip)CPU

Flash attention is compiled for one GPU architecture at a time, so there is a CUDA build per GPU family. nvidia-smi --query-gpu=compute_cap --format=csv shows yours: 8.x is sm80, 9.0 is sm90, 10.0 is sm100, 12.0 is sm120. The CUDA builds use CUDA 12.8, need an NVIDIA driver R570 or newer, and carry the CUDA runtime libraries they use in lib/ next to the binary, so no CUDA toolkit is needed.

tar -xzf vakyn-<version>-linux-x64-cpu.tar.gzsha256sum -c vakyn-<version>-linux-x64-cpu.tar.gz.sha256./vakyn-<version>-linux-x64-cpu/vakyn serve

The web app (front page, server page and dashboard) is built into the binary; nothing else needs to be installed or served at runtime.

Building from source

You need Rust (stable) and Node.js 22 for the web app that is embedded in the binary:

git clone https://github.com/vakyn-ai/open-server.gitcd open-server(cd web && npm ci && npm run build)cargo build --release# the binary: ./target/release/vakyn — put it on your PATH./target/release/vakyn --version

Commands

CommandWhat it does
vakyn serveRun the server in the foreground (logs to stderr; Ctrl-C stops it)
vakyn pull REPO [--quant Q6K]Download a release from Hugging Face into the cache without serving it; see Downloading releases
vakyn startRun the same server in the background, with the same flags; logs go to <data-dir>/vakyn.log
vakyn statusShow whether it is running, with pid, address and uptime (exit code 3 when it is not)
vakyn stop [--timeout 30]Stop the background server gracefully; kill it if it has not stopped after the timeout
vakyn keys create --name NAMECreate an API key and print it once
vakyn keys listList keys: name, prefix, seed, created, last used, revoked
vakyn keys revoke NAME_OR_PREFIXRevoke keys by name or by prefix (vk_ plus 8 characters)
vakyn db backup [--out PATH]Write a consistent copy of the database; safe while the server runs
vakyn db restore PATHReplace the database with a backup (server stopped; the current database is backed up first)
vakyn db reset [--yes]Delete the admin account, keys and usage after taking a backup (server stopped)
vakyn service install [FLAGS]Install vakyn as a systemd service, enable it at boot and start it (Linux); see Service
vakyn service uninstall [--purge]Stop and remove the service; --purge also deletes its data directory
vakyn --versionPrint the version

Every command accepts --data-dir, so several servers can run side by side on different ports and directories. Keys can also be created, listed and revoked on the dashboard's API keys page after signing in.

Per-key seed

Each API key gets a random 64-bit inference seed when it is created (keys from earlier versions get one when you upgrade). It never changes, even after the key is revoked, and every request made with the key passes it to the engine. Playground runs use one fixed seed of their own. The seed is shown read-only on the API keys page and by vakyn keys list.

Today the model computes its answers without sampling, so the seed does not change any numbers: the same request gets the same probabilities with any key. The seed applies wherever the engine uses randomness.

Service (Linux)

vakyn service install --model /srv/VAKYN-4B-UQFF        # user servicesudo vakyn service install --model /srv/VAKYN-4B-UQFF   # system servicevakyn service install --dry-run --port 9000             # print the unit onlyvakyn service uninstall                                 # keeps the datavakyn service uninstall --purge                         # deletes it too
  • With sudo (or as root) it writes /etc/systemd/system/vakyn.service, which runs vakyn serve as a dedicated vakyn user on /var/lib/vakyn. Otherwise it writes a user unit in ~/.config/systemd/user/ on your own data directory; a user service runs while you are logged in, and sudo loginctl enable-linger $USER starts it at boot. --user and --system choose explicitly.
  • Flags after install are vakyn serve flags for the service; they are checked, and relative paths are made absolute. --model vakyn/VAKYN-4B-UQFF works too: the service downloads the release on its first start (put HF_TOKEN=... in the env file below for a private release; --hf-token is refused so the token never lands in the unit). An --admin-password (or VAKYN_ADMIN_PASSWORD) is applied at install time and never written into the unit; a generated admin password is printed once by service install. Other settings (RUST_LOG, ...) can go in /etc/vakyn/vakyn.env or ~/.config/vakyn/vakyn.env.
  • Installing again with other flags updates the unit and restarts the service; with the same flags nothing changes. It refuses while vakyn start runs on the same data directory.
  • Status and logs: systemctl status vakyn and journalctl -u vakyn -f (add --user for a user service). API keys for the system service: sudo -u vakyn vakyn keys create --name app --data-dir /var/lib/vakyn.
  • macOS and Windows services are not supported. There, vakyn service says so, and vakyn start runs VAKYN in the background.

Server flags

For vakyn serve and vakyn start:

FlagEnvironmentDefaultMeaning
--hostVAKYN_HOST127.0.0.1Address to listen on. Use 0.0.0.0 to accept connections from other machines
--portVAKYN_PORT8080Port; 0 picks a free one (printed at start)
--data-dirVAKYN_DATA_DIRper OS, belowWhere the database, log and backups live
--admin-userVAKYN_ADMIN_USERadminDashboard username
--admin-passwordVAKYN_ADMIN_PASSWORDgenerated onceDashboard password; when given, it is set on every start
--max-queue32Requests admitted at once, running or waiting; more get 429
--session-ttl604800 (7 days)Dashboard sign-in lifetime, in seconds
--login-attempts5Failed sign-ins from one address before it has to wait (1 s, then doubling, up to 15 minutes); 429 with retry-after meanwhile
--login-global-attempts50Failed sign-ins across all addresses before everyone has to wait (up to 1 minute)
--trust-proxyoffTake the client address for sign-in limiting from X-Forwarded-For; turn on only behind a reverse proxy
--site-urlVAKYN_SITE_URLhttps://vakyn.comThe VAKYN website the web app links to for docs, skills and the FAQ
--enginemistralrs with --model, else fakeWhich engine answers: the model (see Model and tuning) or a deterministic test engine
--fake-latency-ms0Test engine only: delay added to every request
--logAlso write the server log to this file

Log verbosity follows RUST_LOG (default info). At start the server logs its effective settings; the public server page (/server on your server) shows the same values.

Model and tuning

VAKYN releases are on Hugging Face, for example vakyn/VAKYN-4B-UQFF. Name one with --model and the server downloads it on first start, or point --model at a release folder you already have:

vakyn serve --model vakyn/VAKYN-4B-UQFF             # downloads, then serves Q8_0vakyn serve --model VAKYN-4B --quant Q4K             # short name; or Q6Kvakyn serve --model ./VAKYN-4B-UQFF                  # a local release folder# or with a config file holding the same keys (flags win)vakyn serve --config vakyn.toml
FileWhat it is
VAKYN-4B-Q8_0-0.uqffThe weights at 8 bits: the default, closest to the reference answers
VAKYN-4B-Q6K-0.uqff6 bits: smaller, answers stay close; --quant Q6K
VAKYN-4B-Q4K-0.uqff4 bits: smallest, answers drift more; --quant Q4K
config.json, residual.safetensorsWhat the engine needs next to the UQFF files
vakyn_head.safetensorsThe pointer head that scores the options
calibration.jsonTemperatures that calibrate the probabilities
tokenizer.jsonThe tokenizer

The server downloads and loads only the quantization you choose.

The model loads in the background: the server page (/server) shows loading and the API answers 503 with retry-after until it is ready. If the model cannot load (for example the context does not fit in GPU memory), the server stops with a message saying what to change. The effective settings are logged at start and shown on the server page.

Flag / config keyDefaultMeaningllama-server analogue
--ctx-size65536KV cache size in tokens (the default fits the largest request the API accepts)-c
--gpu-memory-fractionSize the KV cache as a share of GPU memory instead
--gpu-memory-mbSize the KV cache in MB instead
--cache-typeautof8e4m3 halves the KV cache (CUDA)--cache-type-k/v
--max-seqs32Sequences the engine runs at once-np
--max-batch-tokensengine defaultTokens per scheduler step-b
--prefill-chunkengine defaultPrompt tokens per prefill chunk-ub
--prefix-cache-n6Documents kept for reuse; each cached 32k-token document holds about 1.2 GB of GPU memory (0 disables)--cache-reuse
--paged-attnauto (off)Paged attention; off by default because it is slower here, see the note below
--cpuoffRun on the CPU even with a GPU
--device-layersautomaticLayers per GPU: 0:20 1:12, or one number for GPU 0; layers not placed run on the CPU (same answers, several times slower)-ngl
--dtypeautoActivations: f16 on GPUs, the engine's choice on CPU
--load-timeout900Seconds for loading, warm-up and the fit check before startup fails with the step it was on (0 waits forever)
--seedSeed for the engine's startup checks; API and playground requests use their key's seed (see Per-key seed)-s

Each request reads its document once and then answers all its questions in one batched pass; requests from different clients run side by side. On a GPU the server checks at startup that --prefix-cache-n + 1 maximum-length documents fit, so an oversized setting fails then and not under load. Before loading on a single NVIDIA GPU, it also estimates the memory the model needs at --ctx-size. If the GPU has less free, startup fails with that number and what fits (a lower --ctx-size or --prefix-cache-n, --quant Q4K or a smaller model) instead of quietly running part of the model on the CPU; --device-layers N runs the remaining layers on the CPU on purpose. Requests longer than --ctx-size (or the API's 32,768-token limit) get the same 400 as any over-limit request.

Downloading releases

export HF_TOKEN=hf_...                                # private releases onlyvakyn pull vakyn/VAKYN-4B-UQFF                       # download Q8_0, print the foldervakyn pull VAKYN-4B --quant Q6K                      # another quantvakyn serve --model vakyn/VAKYN-4B-UQFF              # serves from the cache, no network
  • What is downloaded: the weights of one quant and the files the server needs next to them (config.json, residual.safetensors, tokenizer.json, vakyn_head.safetensors, calibration.json, generation_config.json, chat_template.jinja, manifest.json). Other quants are not.
  • Token: --hf-token, else the HF_TOKEN environment variable, else the token saved by hf auth login. It is never printed or logged. With vakyn start, prefer HF_TOKEN: flags show up on the process command line.
  • Cache: the standard Hugging Face cache (~/.cache/huggingface/hub, or HF_HUB_CACHE / HF_HOME/hub), shared with hf download. --model-dir DIR (VAKYN_MODEL_DIR, or model_dir in the config file) keeps the files in DIR instead, with the same layout.
  • Progress: a bar per file and one for the total on a terminal; a log line every 10 seconds under Docker or systemd. While it downloads, the server already listens: GET /api/status reports "status": "downloading" with download.bytes_done and download.bytes_total, the server page shows "Downloading 1.2 / 4.5 GB", and the API answers 503 with retry-after.
  • Resume and checks: an interrupted download resumes where it stopped (run the same command again), and every file is checked against the SHA-256 in the release's manifest.json; a file that doesn't match is downloaded again.
  • Offline: once a quant is complete in the cache, vakyn serve uses it without the network. vakyn pull always checks Hugging Face for a newer release.
  • Errors say what to do: a private release without a token (set HF_TOKEN or run hf auth login), an unknown release or quant (the available quants are listed), a full disk, or a lost connection (retried with back-off; then run the command again to resume).

Building with the model engine

Prebuilt releases include the engine. From source, choose the backend with a cargo feature:

(cd web && npm ci && npm run build)cargo build --release --features cuda,flash-attn   # NVIDIA (set CUDA_COMPUTE_CAP, e.g. 120 for RTX 50xx)cargo build --release --features metal             # Apple Siliconcargo build --release --features engine-mistralrs  # CPU only

Flash attention matters on CUDA: long documents are about three times slower without it.

Data directory

SystemDefault
Linux~/.local/share/vakyn
macOS~/Library/Application Support/ai.vakyn.vakyn
Windows%APPDATA%\vakyn\vakyn\data
FileContents
vakyn.dbSQLite: the admin account (password hash), API keys (hashes and seeds), usage events, dashboard sessions, the playground seed
vakyn.lockHeld by the running server so two servers never share a directory
vakyn.pidThe running server's pid and address, for status and stop
vakyn.logOutput of vakyn start
backups/Copies written by db backup, db restore and db reset

Keys and passwords are stored only as argon2 hashes; a key is shown once, when it is created. Documents and answers are not stored, only per-request counts (time, key, input tokens, status, request id).

Backup and restore

# while the server runsvakyn db backup                         # -> <data-dir>/backups/vakyn-YYYYMMDD-HHMMSS.dbvakyn db backup --out /mnt/backup/vakyn.db # restore: stop firstvakyn stopvakyn db restore /mnt/backup/vakyn.db   # the current database is backed up firstvakyn start

vakyn db reset asks for confirmation (skip it with --yes), backs up, then deletes everything; the next start creates a new admin account. Back up before upgrading.

Network and HTTPS

  • By default the server only listens on 127.0.0.1, so only the same machine can reach it.
  • There is no built-in TLS. To serve other machines, put a reverse proxy that terminates HTTPS in front (Caddy, nginx, Traefik) and keep VAKYN on localhost behind it.
  • Behind a proxy every request comes from the proxy's address, so start VAKYN with --trust-proxy: failed sign-ins are then counted per client, using the last X-Forwarded-For entry (the one the proxy adds). Without a proxy leave it off, since clients could forge the header.
  • The front page and the server page are public; the dashboard needs the admin sign-in; the API needs a key.

Example: Caddy

vakyn.example.com {	reverse_proxy 127.0.0.1:8080}

Docker

Every release pushes one image per engine build to the GitHub Container Registry. All of them run as a non-root user (uid 10001), listen on 0.0.0.0:8080 and keep the database and the model cache in the /data volume (/data/.cache/huggingface/hub), so a restart doesn't download the model again.

TagFor
cpuAny x86-64 or arm64 machine, no GPU
nvidia-sm60Pascal: P100
nvidia-sm70Volta: V100
nvidia-sm75Turing: T4, RTX 20xx
nvidia-sm80Ampere and Ada: A100, A10, RTX 30xx, RTX 40xx, L4, L40S
nvidia-sm90Hopper: H100, H200
nvidia-sm100Blackwell data center: B200, GB200
nvidia-sm120Blackwell RTX: RTX 50xx, RTX PRO 6000

Each tag also exists per release, for example 0.1.0-nvidia-sm70. To find your GPU family, run nvidia-smi --query-gpu=compute_cap --format=csv: 8.6 means nvidia-sm80, 7.0 means nvidia-sm70. A CUDA image refuses to start on another family and names the right tag. Images from sm80 up use flash attention.

docker run -d --name vakyn --gpus all -p 8080:8080 -v vakyn-data:/data \  -e VAKYN_MODEL=vakyn/VAKYN-4B-UQFF -e VAKYN_QUANT=Q8_0 \  ghcr.io/vakyn-ai/open-server:nvidia-sm80 # A release folder you already have, mounted under /modelsdocker run -d --name vakyn -p 8080:8080 -v vakyn-data:/data \  -v "$PWD/VAKYN-4B-UQFF:/models/VAKYN-4B-UQFF:ro" -e VAKYN_MODEL=/models/VAKYN-4B-UQFF \  ghcr.io/vakyn-ai/open-server:cpu docker logs vakyn                                # the admin password generated on first startdocker exec vakyn vakyn keys create --name app   # create an API key
  • Every vakyn serve flag has a VAKYN_* environment variable: --ctx-size is VAKYN_CTX_SIZE, --max-seqs is VAKYN_MAX_SEQS, and so on. The full list is in BUILDING.md in the repository.
  • Set the admin password with -e VAKYN_ADMIN_PASSWORD=…, or read the generated one from docker logs on the first start. Pass -e HF_TOKEN for private models.
  • A named volume (vakyn-data above) needs no setup. A host directory mounted at /data must be writable by uid 10001 (sudo chown 10001 /srv/vakyn).
  • The CUDA images need the NVIDIA Container Toolkit and driver R570 or newer (R580 for Pascal).
  • The port is published on all host addresses by -p 8080:8080; use -p 127.0.0.1:8080:8080 to keep it local behind a reverse proxy (with --trust-proxy).

Building the images

docker build -t vakyn:cpu .docker build --build-arg VARIANT=cuda --build-arg CUDA_COMPUTE_CAP=70 -t vakyn:nvidia-sm70 .

A CUDA build compiles the kernels for the one GPU family you choose: tens of minutes, or hours from sm80 up because of flash attention. BUILDING.md in the repository covers every platform.

Loading the docs…