Skip to content
Vertex Lake
Menu

Trust · from the assurance pack

Hardening guide

Published from docs/assurance/hardening-guide.md as it ships with the software: the same text an auditor receives. A reference to another document of the pack or to a runbook is named, not linked; they ship beside it.

The controls in the mappings hold when the deployment does. This is the short list that makes it so; each item says why.

Network and transport

Identity

Data at rest

Audit evidence custody

Platform

Modules and deployment locks

VX Drive is a core that is always on and modules an administrator switches (ADR-0016). What the deployment must never offer is fixed above the administrator, in configuration:

Set the same values on the server and the worker. The server publishes its locks and the worker applies them together with its own; where they differ, off wins. A lock that names something unknown, contradicts another (on and off; on, while something it needs is off), or that this build cannot yet honour stops the server at startup with the reason: a typo in a hardening setting never passes silently. A change of locks is on the audit trail (module.locks_changed), and so is every attempt to switch against one (module.switch_refused).

Never set VXDRIVE_MODULES_PREVIEW on a real deployment: it treats planned modules as built so tests can exercise their switches, like the mock scanner.

The legal modules: directories, sealed data, egress, operator-supplied things

Directories the worker reads. Provenance, Legal discovery and Legal forensics take their inputs from directories the operator fills — never through the API — and list them into the database every minute: VXDRIVE_HASH_SETS_DIR (NSRL RDS, hash lists, VICS JSON), VXDRIVE_DISCOVERY_DIR (load-file packages), VXDRIVE_BIOMETRIC_SETS_DIR (labelled sets a calibration is fitted on) and VXDRIVE_FORENSICS_IMAGE_ROOTS (forensic images, read in place, colon-separated roots that must exist when the worker starts). Mount each read-only into the worker (deploy/compose.yaml shows the lines): the worker only reads them, an image is evidence that nothing may write, and the sandbox around every helper gets the roots as read-only paths. Keep who may write to those directories to the people who handle the material; the application never can. VXDRIVE_MODELS_DIR stays writable by the worker alone (it installs catalog models there); an operator's own model file is placed there by hand and registered from Settings.

Sealed data and the instance key. Beside source credentials, TOTP enrolments, OIDC client secrets and keychain keys, the instance key seals SIEM bearer tokens, backup destinations' repository passwords and credentials, and biometric templates (each template bound to its workspace, sample and model, so a copy opened elsewhere fails); the backup signing key is derived from it, so a reseal changes the key every later manifest is signed with (re-pin it wherever a restore is verified). vxdrive-server --reseal-secrets re-encrypts every sealed table under a new key; the worker needs the key file for the hosted stages, SIEM forwarding, biometric work and every backup run (compose and the appliance mount it into both processes). A key that is lost leaves templates unopenable — destroyed in all but name — and the sealed backup secrets with them: "If the instance key is lost" in docs/runbooks/backups.md says what to do then.

Egress, in one place. Every outbound connection the product can make, with the host to allow; nothing else leaves, and every sandboxed tool runs without network (the processes outside the per-file sandbox — the model engines, which run under their own Landlock confinement, restic and the database dump — are named below):

Connection Process Switch What leaves
Hosted AI providers (ADR-0012) server and worker a stage pointed at a provider and the workspace allowing it text, pictures or audio of the documents and questions the policy allows; the audit chain records sizes, never content
Keychain verification server (daily check: worker) adding or testing a key the key, to the provider's own host
Time-stamping authority (tsa_url) worker Settings → Forensics a SHA-256 digest inside an RFC 3161 query; no content
SIEM destinations worker Settings → Forensics → SIEM destinations, with the activity feature on the audit chain's event lines and hashes; no document content
DNS lookups for email authentication (external_lookups) worker Settings → Forensics, off by default DKIM selector, SPF and DMARC names of examined messages, to the resolver the worker uses — point it at a resolver you run, or leave lookups off (verification then uses answers kept earlier and reports the rest unverifiable)
Model catalog downloads worker installing a catalog model nothing; the files are fetched from huggingface.co and from the vendor's storage account (vlstoree5183944.blob.core.windows.net, the public-read models container) and checksum-verified; an air-gapped instance places them by hand
Backup destinations (ADR-0022) worker (restic, outside the sandbox; the proxy settings are not passed) Settings → Backups the encrypted, deduplicated snapshots, to the repository hosts the administrator named
A web address a contributor adds (Documents → Add from a web address) server the Export and import module a GET to the address the person typed; restrict egress to what you allow
Ingestion sources, scanners, the identity provider, the object store server, worker the operator's own configuration as those protocols require, to hosts the operator named

Two settings shape that table. VXDRIVE_URL_IMPORT_DENY_PRIVATE=1 makes the server refuse a web address whose final peer is a private, loopback or link-local address: set it wherever the server can reach internal services a contributor should not be able to fetch through it (the check is on the address the connection reaches after redirects). And the HTTP clients honour HTTPS_PROXY/HTTP_PROXY/NO_PROXY — hosted providers, the keychain's checks, the time-stamping authority, SIEM destinations, web addresses, the identity provider and the object store go through a proxy you name — except the inference clients (the engines are on the private network), model downloads and restic, which connect directly; an instance behind a proxy places model files by hand and names backup destinations it can reach without the proxy.

The server runs two programs of its own outside the worker's sandbox, for LAN protocols that need the host's network: smbclient (network shares, read-only, credentials in a 0600 file) and scanimage (SANE scanners). Each runs with a scrubbed environment — the path, a private home, the locale, and for scanimage the SANE_* variables — so the database password and the engine token, which the server holds in root-only environment files (0600) on the appliance and under compose, never reach them.

The worker runs two more outside its sandbox, for the backups (ADR-0022): restic, which needs the network to reach the repositories (and runs ssh or rclone for those backends), and the PostgreSQL client (pg_dump, pg_dumpall), which needs the database. Each starts with a scrubbed environment that carries only what the run needs — a path, a home of its own for restic's cache, the repository, the destination's password and backend credentials opened from the sealed backup_secrets row for that run (a copy run carries the primary's as well, since restic reads from it), and the database connection as libpq's PG* variables, never on a command line; nothing of the worker's own environment and no proxy settings. Credential lines are checked by name when they are saved — capital letters, digits and underscores; the names the worker sets (PATH, HOME, TMPDIR, DATABASE_URL) and the prefixes RESTIC_, LD_, PG, VXDRIVE_, SSH_, SSL_ and OPENSSL_ are refused with 422 — and the worker sets its own variables last, so a credential line can never choose what the backup tools run. The worker restores nothing: a backup is proven from the copy, by whoever restores it, with the public signing key.

Operator-supplied executables and models. A perceptual matcher (VXDRIVE_TOOL_PERCEPTUAL_MATCHER), an ONNX detector or embedder registered from the models directory, a YARA ruleset entered on the settings page, a validation set: each runs in the worker's sandbox with the one file it is given and no network, and each is the operator's licensing and fitness decision, recorded with who added it. Ghostscript is the one tool probed by name and never shipped (AGPL): an operator who holds a licence adds it in a derived image and points VXDRIVE_TOOL_GS at it. Nothing else optional is integrated today (pffexport and plaso are named in the ADRs as candidates only).

Regulated features. Lock off what your instance must never offer, whatever an administrator does: VXDRIVE_MODULES_LOCKED_OFF=legal_forensics.voice_comparison,legal_forensics.face_comparison,legal_forensics.restricted_hash_sets (or the module as a whole). Where a feature is on, its notice was accepted by a named administrator at a recorded text version; state the jurisdictions the instance operates in (Settings → Modules) so the notices carry the right notes, and know that changing them makes every acceptance stale.

Evidence mode and the superuser. The gate at every byte site, the two-person rule and the custody chain bind every database role, as the audit chain does; a PostgreSQL superuser on the host can still drop them, which is why the off-host copies — a verified backup, an exported custody record, a SIEM capture, a timestamp token — are what make a rewrite detectable (docs/security/audit-chain.md).

The appliance (ADR-0021, ADR-0023, ADR-0024)

Things NOT to do

Inference engines (semantic search, Ask)

The worker starts one llama-server per role (embeddings on VXDRIVE_INFERENCE_PORT, 8089; the answer model on VXDRIVE_INFERENCE_GENERATOR_PORT, 8090; the re-ranker on VXDRIVE_INFERENCE_RERANKER_PORT, 8091; the describer — a vision model — on VXDRIVE_INFERENCE_VISION_PORT, 8092, which only the worker calls; the visual model on VXDRIVE_INFERENCE_VISUAL_PORT, 8093, reached by the server through VXDRIVE_INFERENCE_VISUAL_URL for the question's vectors; and the local answer models the Ask profiles pin beside the instance's own on VXDRIVE_INFERENCE_LOCAL_PORT_BASE + n, 8100 onwards, one slot per model within the instance's Local answer models loaded at once setting, VXDRIVE_INFERENCE_LOCAL_SLOTS slots at most, which the server reaches on the generator URL's host at the port the worker reports on the model's row) and the server reaches the first three through VXDRIVE_INFERENCE_URL / VXDRIVE_INFERENCE_GENERATOR_URL / VXDRIVE_INFERENCE_RERANKER_URL with the shared VXDRIVE_INFERENCE_TOKEN[_FILE]. Keep every engine port on the compose network only — deploy/compose.yaml publishes neither — and rotate the token like any other secret. The engines run outside the per-file sandbox (they need a listening socket) under a Landlock confinement of their own: an engine reads its own directory, the models directory and the system's paths (/usr, /lib, /lib64, /etc, /opt, /proc, /sys), writes only under the worker's work directory, uses the device nodes under /dev in place (the GPU's, their ioctls included), binds one TCP port — its own — and connects nowhere; with a descriptor limit, no core dumps, no new privileges and a scrubbed environment (the shared token reaches the engine in that environment, as LLAMA_API_KEY, never on its command line: the worker clears the environment it starts the engine with and sets the home, the path and the key alone). The worker's start line says confined=true when the kernel enforces Landlock (best effort by ABI: a kernel without network rules leaves the socket unconfined); VXDRIVE_INFERENCE_CONFINE=0 turns the confinement off for a GPU stack that lives outside those paths, after which an engine sees whatever the worker's user can. The engines die with the worker (PR_SET_PDEATHSIG), and a worker that finds its port already answered by another engine refuses to start one and says so in its log — kill the stray llama-server before restarting. Models are installed by the worker from deploy/models.json with a checksum check — from Hugging Face, or from the vendor's storage account (vlstore<suffix>.blob.core.windows.net, the models container, public read) for the files the vendor converts itself — or placed in VXDRIVE_MODELS_DIR by hand on air-gapped instances (the checksum is still verified); keep that volume writable by the worker only. An operator may also register a model file of their own from that directory (POST /inference/models/register, Settings → Models): it is hashed and recorded as operator-registered with the licence note the operator entered, it passes the same gates before it serves (the re-ranker's scoring self-test, the visual model's numerics gate), and it is the operator's licensing decision, never part of what the catalog vouches for. The release publishes the worker image in two variants of one build (ADR-0021 §9): cpu, and cuda under the tag suffix -cuda — the CUDA build of the same engine (the archive pinned as LLAMA_CUDA_URL, with the CUDA runtime libraries inside it and their EULA in the image's notices); it needs the NVIDIA container toolkit to see the GPU and nothing else — no extra ports, no network, the same sandbox around every other tool — and it serves on the CPU when no GPU is present, which vxdrive-workers devices tells apart (the engine's own device listing). Leave "Ask your documents" off until the instance's records rules for conversations (retention days, who may ask) are decided; an answer is never a record.

Keys for hosted providers (ADR-0012). The keychain (Settings → Keychain) stores API keys sealed under the instance key (or as env:/file: references the processes resolve at use); the application role cannot read the sealed table, the API never returns a value, and every change is audited by id and hint. Adding or testing a key makes one outbound request to the provider's own host — the key, no content — so an egress allowlist needs api.openai.com, generativelanguage.googleapis.com, api.cohere.com, api.voyageai.com, api.jina.ai or the operator's own endpoint as the keys require, and nothing else; an air-gapped instance leaves keys unverified or adds none. Mount the same instance.key into the worker (compose does) so its daily check can open sealed keys, and rotate the instance key with vxdrive-server --reseal-secrets (runbook hosted-ai.md).

Hosted stages (ADR-0012). Every AI stage except text recognition and the visual index can be pointed at a provider from its Runs on block; nothing is until an instance admin does it, and then only documents whose workspace allows hosted AI (off by default, per workspace, never under an information barrier) and questions or searches whose scope allows it go there. To keep an instance local: add no provider, or leave the instance policy off and every workspace on Follow the instance default; the Runs on blocks then show "On this instance" and the departures report stays empty. When you do use a provider: allow egress only to its host (the server and each worker call it), set a daily spend ceiling and a data-handling note recording the zero-data-retention term you agreed, leave the per-stage fallback switch off unless continuity matters more than knowing where each unit of work ran, and review the departures report (GET /ai/departures, the AI providers page) as part of the audit review: one ai.call event per call, with sizes and outcome and never content. Held documents are processed like any other; the report says which held documents' content left.

SIEM destinations (ADR-0018 part F). With Legal forensics' Activity analysis on, an administrator may name SIEM destinations under Settings → Forensics: the worker then POSTs the audit chain to each — the export's own lines with their hashes, no document content — over https (http only to a loopback collector) with an optional bearer token sealed under the instance key. Allow egress from the worker to the destination's host and nothing else for it; a capture at the receiver verifies with vxdrive-server --verify-audit-export, and a receiver that stops taking the chain is reported to the administrators, never skipped.

Records housekeeping. Superseded derived copies are purged after superseded_retention_days (30 by default; 0 keeps them): only rows a newer copy replaced, never originals, current copies or anything under a hold, one audit event per row — set the period to your retention schedule, not to disk pressure. Migration 0041 (search in Chinese, Japanese and Korean) rewrites the two search tables and rebuilds their indexes under lock when the server first starts on it: on a large library take a maintenance window and raise maintenance_work_mem for that session. Graph pins are a person's own view preference (where they left a document in a graph), stored per user, readable by nobody else and not audited — like a search query, not like a record.

Transcription is different: whisper-cli is not a service. It runs once per recording under the per-file sandbox like Tesseract (no socket, no network, the models volume read-only), on speech ffmpeg extracted, with a wall clock scaled to the recording's length. The transcriber model and its voice-activity companion are checksum-verified like every other model; a transcript is machine output and is labelled so.

The visual index's pages have the same shape: vx-media-embed, this build's own program over llama.cpp's library, runs once per document under the per-file sandbox (no socket, no network, the models volume read-only; ADR-0011 section 22), and only the questions go to the visual engine. When the visual model is set to use the GPU, that run alone is granted what the engines' confinement grants them — the device nodes under /dev with their ioctls and /sys to read — so the card is used; every other sandboxed tool sees five standard device files and nothing else under /dev.