Skip to content
Vertex Lake
Menu

Trust · from the assurance pack

Security questionnaire answers

Published from docs/assurance/questionnaire-answers.md as it ships with the software: the same text an auditor receives. A reference to another document of the pack or to a runbook is named, not linked; they ship beside it.

Honest, reusable answers to the questions vendor-security reviews (SIG Lite, CAIQ, one-off spreadsheets) actually ask about VX Drive. First person = the deploying organization where the question is about operations; "the product" where it is about shipped code.

Deployment model — Self-hosted, single-tenant: one instance per organization, on the organization's own hardware or cloud account. No vendor access to instance data. The vendor operates its customer portal, its administration and the release registry alone (ADR-0025, ADR-0027, ADR-0031): they hold the account and serve releases, never an instance's data.

Data segregation — Between customers: physical (separate instances). Within the organization: workspace-level need-to-know and information barriers enforced by PostgreSQL row-level security that binds every database role, with 404-not-403 masking so restricted workspaces' existence does not leak.

Authentication — Argon2id password hashing (OWASP parameters); optional TOTP MFA with single-use recovery codes; OIDC SSO (authorization code + PKCE); SCIM 2.0 provisioning with deprovision-as-disable. Sessions are opaque tokens stored hashed, 30-day TTL, immediately dead on disable.

Authorization — Roles: owner, admin, member. Admins administer without reading restricted content through the document routes and the pages: need-to-know, barriers, quarantine and evidence mode bind owners and administrators there. Three administrator-only paths carry everything, and each is an audit event: the instance export (everything but evidence-mode bytes and quarantined documents), the backup catalogue (GET /api/v1/backups/inventory: every document's title, workspace, filename and hash), and the backup destinations an administrator names, which receive every byte, evidence-mode and quarantined documents included — so keep administrators to named people and read their audit trail (hardening guide). Time-boxed emergency access with mandatory reason, announced to all admins, never crossing an information barrier. Per-workspace access report available on demand.

Audit logging — Every attributable action, with exactly-one-actor attribution, in an append-only SHA-256 hash chain with periodic anchors. Verification: in-app (/api/v1/audit/verify), and offline against exports with either the shipped binary or an independent stdlib-Python verifier. Readers of a document see its full history (access transparency).

Data retention & deletion — Direct hard deletion is impossible (database triggers bind every role). Deletion is a reversible trash state; destruction happens only through admin disposition — never under a hold, always producing an immutable certificate, sparing content shared with surviving documents. Retention schedules feed a human review queue; nothing auto-destroys. The one automatic removal is of derived copies (thumbnails, reading copies, archival copies) that a newer copy of the same document superseded: they are kept for a configurable period (30 days by default; 0 keeps them), then purged by an hourly sweep that never touches an original, a current copy or a held document, spares bytes anything else still references, and writes one audit event per removed row.

Encryption in transit — TLS 1.2/1.3 (rustls) served directly with HSTS, or at the organization's proxy. An appliance obtains its certificate one of four ways (ADR-0021): a Let's Encrypt certificate for a tailnet name, a private authority the installer makes for a LAN name, the organisation's own certificate, or Let's Encrypt for a public domain through the ACME DNS-01 challenge; a daily timer renews it, the server reloads it live and audits the reload, and an administrator is notified fourteen days before expiry. A FIPS build variant runs the aws-lc-rs FIPS 140-3 validated module for TLS and self-asserts at startup; /healthz reports the posture. No FIPS image or artifact is published: the switch is proven by CI on every change, and a deployment that needs it asks the vendor for that build (the source is not public). Restic's repository encryption, the Ed25519 backup manifests, the P-256 licence verification and the sealing of stored secrets are outside the validated module.

Encryption at rest — Stored credentials and secrets (source passwords, OIDC secrets, TOTP seeds, keychain keys, SIEM tokens, backup destinations' passwords and credentials, biometric templates) are sealed with ChaCha20-Poly1305 under a file-held instance key, from which the backup signing key is derived as well. Bulk content encryption at rest is the platform layer (volume/storage encryption), per the hardening guide.

Backup / DR — Backups are a product function (ADR-0022), not a script an operator must remember: on the schedule the administrator sets (the recovery point, 15 minutes at the shortest) the worker takes one snapshot of the database dump, the files and a readable tree of the originals into an encrypted, deduplicated restic repository (the primary), and feeds every other destination (S3, B2, Azure, GCS, SFTP, a REST server, rclone — any region) from it, each with its own retention. Every backup can be opened without us: stock restic restores it; catalogue.jsonl inside it lists every document in plain JSON (title, workspace, state, the original's hash and type, the pages, the relations, where its file is); the readable tree is a directory of the originals named by workspace and title; vxdrive.pgdump restores into any PostgreSQL 18 with pgvector; and manifest.json (schema version, the audit chain's head, every custody record's head, counts, the digests of the blob list and the catalogue) is signed with an Ed25519 key whose public half is on the Backups page and in the recovery kit, so openssl pkeyutl -verify proves the manifest is the instance's; the copy's wholeness is the verifier's read-back against the digests the manifest carries. Repository integrity is checked continuously: restic check with data read back (a share daily, everything weekly) per repository. Restores are verifiable by anyone with the key and stock tools: vxdrive-server --verify-restore proves a restored copy (the schema version, the audit chain and every custody record to the heads the manifest names, the counts, the catalogue against the database, the readable tree's links and the blobs hashed to their names — a random 25 of each unless --all is given —, and the signature under the pinned key: with a key given, a missing signature file fails; without a key the provenance is printed as not proven and the exit code is still 0), deploy/drill-restore.sh does the whole restore on another machine with --all, and GET /api/v1/backups/inventory answers the live catalogue so a restored copy can be compared with the source. The product runs no restore rehearsal of its own: a rehearsal is the operator's to schedule, with those tools, from a machine that is not the appliance. A failed check, an overdue backup or a destination that fails three runs raises an administrator notification and marks the page. The recovery kit made at install (the instance key, the environment files, the TLS material, the licence, the install kit's files and the signing key) must live off the appliance; each destination's repository password and credentials are typed later, sealed under the instance key and never shown again, so the operator adds them to their copy of the kit — without them no backup can be opened. Backups keep what disposition or erasure removed until the snapshots holding it are forgotten by the retention policy (the defaults keep 48 hourly, 30 daily, 12 weekly, 24 monthly and 5 yearly snapshots), and the appliance's pre-update rollback dumps are plain pg_dump files under /var/lib/vxdrive/backups/pre-update-*, the last three kept. RTO/RPO are the operator's schedule; the runbook prescribes an hourly snapshot, at least one copy in another place and a restore of one's own on a schedule.

Vulnerability management — Coordinated disclosure (SECURITY.md; 2-day acknowledgement, 30-day critical fix target), EU CRA reporting runbook (24h/72h/14d), five-year declared support period, dependency audit + OpenSSF Scorecard in CI, SBOM (SPDX + CycloneDX) and Sigstore-signed artifacts on every release.

Subprocessors — None for a self-hosted instance unless the operator adds a hosted AI provider (ADR-0012), which is then the operator's own subprocessor under the operator's terms with it; the product records what was sent to it and never asserts what it does with the data. The vendor's own services (the portal, its administration and the registry, ADR-0029) use Stripe and Microsoft Azure and no one else; a backup destination the customer chooses holds the customer's encrypted snapshots and is the customer's choice, not ours.

Penetration testing — [Organization to attach the current letter; the vendor publishes product pen-test summaries as they complete per the certification roadmap.]

Certifications — The product ships evidence (this pack); operation certificates (SOC 2/ISO 27001) belong to whoever operates: your organization for your instance, the vendor for the portal and the registry. No FedRAMP/CJIS/HITRUST claims.

Formats, OCR and archival copies

Do you convert documents to an archival format? Optionally, per rule: an archival target (PDF/A-2b by default, others per media type) is produced by the worker and verified by veraPDF; the original is always kept unchanged and is itself the master when it already conforms. The veraPDF report is stored with every verified copy.

How is email preserved? A message ingested from a mailbox is a document of its own (the RFC 5322 file, titled by its subject); its attachments are documents related to it and replies are related to the message they answer, so a thread is navigable and every part is under the same holds and retention. The archival copy is the message itself when it parses, an EML rebuilt from an Outlook item (original embedded), or — per rule — an EA-PDF: a PDF/A-3u with the message (or the whole mailbox, behind an index page) rendered from its own parts without any network access, the source and attachments embedded as associated files, the EA-PDF metadata and extension schema in the XMP, and the DPart structure that maps each message's headers, body and attachment list onto its pages. The PDF/A-3u is verified by veraPDF; the EA-PDF structure is self-checked by the worker (no independent EA-PDF validator exists) and the copy records both. Remote images are never fetched; they are removed and the copy says so.

How are audio and video preserved? By what their codecs are: uncompressed, lossless and obsolete material is normalized to FFV1 (video, in Matroska) or FLAC (audio) and every such copy is proven by decoding both the source and the copy and comparing the per-frame and PCM digests — a copy that does not decode to the same streams is never recorded as verified. Lossy delivery codecs are kept as received once they prove to decode, with their decoded-stream digests on record. MediaConch's implementation checks run when installed. Originals are never altered.

Does search use AI, and does content leave the instance? Semantic (meaning-based) search uses an embedding model served by an inference engine that the instance's own worker runs, on the deployment's private network, with a shared token; the vectors live in the instance's PostgreSQL. Models are installed from a catalog pinned by checksum and licence (MIT/Apache-2.0) and listed with the build's software bill of materials. By default no content leaves the instance. An instance administrator may point a stage (semantic search, answers, re-ranking, picture descriptions, transcripts) at a hosted provider (OpenAI, Google, Azure OpenAI, Cohere, Voyage, Jina, or an endpoint of your own) under a key held in the instance's keychain; even then, only documents whose workspace allows hosted AI (off by default, per workspace, never under an information barrier) and searches or questions whose scope allows it go there, every call is an audit event with sizes and outcome and never content, the reader is told what was sent, and a report lists what left per workspace and stage. When the engine is off, search is word-based and says so.

Is there a generative AI assistant, and what does it see? "Ask your documents" is off by default. When an admin switches it on, questions are answered by a small open-weights model (Apache-2.0/MIT, from the same checksummed catalog) that the instance's own worker runs — or, where an admin pointed the answers stage at a hosted provider and the question's scope allows it, by that provider, with passages from workspaces that do not allow hosted AI withheld before any hosted re-ranker or generator sees them, counted, and named in the answer's footer beside what was sent; a question whose every passage must stay on the instance is answered by the instance's own model, or refused when none is active; an earlier answer made from passages that must stay never travels in a later turn's history (ADR-0032). The model only ever sees passages retrieved through the asker's own database session, so a person cannot obtain, through an answer, content they could not open themselves; answers carry numbered citations to the passages used, validated against them, and the model refuses when the documents do not answer. Answers are not records: they are kept per person for a configurable time, are frozen while a cited document is under a hold, are deleted when a cited document is disposed of (audited), and every answer and every cited document is written to the audit chain. The model has no tools and no network, so text inside a document can at most degrade the one answer it appears in. Since the second generation (ADR-0032): a profile may name an auxiliary model for the cheap calls (the rewrites, the running summary of a long conversation, the claim-by-claim check) under the same policy as the answer's; conversations are trees of versions a person may share with named people or groups as viewers or participants, who see each source only where their own sight reaches (a source they may not open stays a number); the answer model may write machine summaries of documents under a processing rule — kept beside the text, labelled as a machine's wherever they appear, never the document's own words — which search and Ask find like any passage; research runs several searches, each read for findings, then a report, with a working stop and a time limit, every pass on the record; a monthly token allowance and a daily question limit are enforced before a question runs, and every call is counted by person, workspace, route and model; administrators read conversations only through an audited route and see feedback and evaluation runs on their own pages.

How is OCR handled? Text recognition is independent of archival conversion: pages are classified from their structure (born-digital text, existing OCR layer, scan, and so on) and OCR runs only where the rule says; the recognized text, its confidence and per-page evidence are stored, and a reading copy carries an invisible text layer that is checked for alignment before it is used. Archival copies include the recognized text only when the rule says so and never trigger recognition.

Can a policy change rewrite existing copies silently? No. A change updates what is expected; the worker adds what is missing and supersedes what is stale, all attempts remain as history, and the change itself is audited. Admins see the cost before saving.

Is any of this processing sent to a third party? Not by default: every per-file tool runs inside the worker container in a sandbox without network access; the model engines run under a Landlock confinement of their own (their directory, the models and the system's paths to read, one TCP port to bind, nothing to connect to); the backup client (the network to the repositories you named) and the database dump run outside the sandbox, each with a scrubbed environment. The only exception is a hosted AI stage an instance administrator pointed at a provider (above), which is opt-in twice, egress-limited to that provider's host, audited call by call and reported; text recognition, copies, hashes and the visual index never have a hosted route.

Can a document be edited? Never in place. A correction is a new version: a new document with its own bytes and record, related to the previous one, in a chain that cannot fork or loop. Search shows the latest version by default and the older ones on request; an answer never cites a replaced version; the chain is listed on every member's page and every step is audited. Notes written in the application (Markdown) go through the same path and are records like any upload.

Which languages does search support? Word-based search stems 29 languages (the PostgreSQL configurations, chosen per document by language detection) and falls back to exact words for the rest; Chinese, Japanese and Korean are indexed by character pairs inside the same index, so two or more adjacent characters of a name or sentence find it without a dictionary (a pair inside a longer word also matches it, which the runbook says). Meaning-based search uses a multilingual embedding model from the catalog. OCR languages follow the Tesseract packs installed in the worker image and the rules' language list.

How are archives (zip, tar, gzip, 7z) handled? They are opened only inside the worker's sandbox, by the worker's own extractor, under caps that make a hostile archive harmless (at most 500 members, 2 GiB per member, a total bounded by the archive's own size, names that cannot escape the extraction directory; encrypted entries refuse the archive). Each member becomes a document of its own — filed with the archive, related to it, audited as created — so it is identified, searched, retained and disposed of like any upload; the archive itself is kept as received. What was refused or left out is on the record and in the review queue.

Modules, provenance and the legal modules (ADR-0016 to ADR-0020)

Can parts of the product be disabled? Yes, at three levels. The product is a core (sign-in, people, workspaces and their access controls, documents, the audit trail) and modules an instance administrator switches for the whole instance; the deployment locks what it must never offer in configuration (VXDRIVE_MODULES_LOCKED_OFF, or VXDRIVE_MODULES_KERNEL_ONLY=1 for the core alone), above the administrator; and some modules and features are allowed workspace by workspace. Off refuses the module's operations (403 module_off), stops its background work and hides its pages; it never deletes or hides a record. Every switch, refusal, notice acceptance and lock change is on the audit trail. The core-only configuration is exercised by its own browser run on every commit.

What outbound connections does the product make? None that you did not configure. The server reaches PostgreSQL and the object store; the worker the same, plus the inference engines it runs itself on the private network. Everything else is the operator's own choice and host: ingestion sources (IMAP, SMB, SFTP, watched folders, the FTP inbox and printer as LAN listeners), scanners on the LAN, an OIDC identity provider, hosted AI providers (opt-in per stage and per workspace, ADR-0012), an RFC 3161 time-stamping authority (Legal forensics), SIEM destinations (Legal forensics), DNS lookups for email authentication (Legal forensics; off by default — with it off, DKIM, SPF and DMARC are verified from answers kept earlier and otherwise reported as unverifiable), a web address a contributor asks the library to fetch, model downloads from the pinned catalog at install time (from huggingface.co and the vendor's storage account, vlstoree5183944.blob.core.windows.net; an air-gapped instance places the files by hand; checksums are verified either way), and the backup destinations an administrator names (ADR-0022; restic, outside the sandbox, with the proxy settings not passed). Every per-file tool the worker runs — including the operator's own executables — runs in a sandbox without network; the model engines run under their own Landlock confinement with one port to bind and nothing to connect to; restic and the database dump run outside the sandbox with a scrubbed environment. An appliance also reaches the vendor's registry and the storage account it redirects layer downloads to, and the vendor's portal for the daily licence refresh and check-in (five fields: the instance id, the version, the channel, the licence's digest, whether the last refresh installed one). The hardening guide lists each with the host to allow.

Can our SIEM receive the audit log? Yes. With Legal forensics' activity analysis on, an administrator names https destinations (a bearer token, if the receiver wants one, is sealed under the instance key and never shown again); the worker forwards the audit chain from its beginning, at least once and in order, as the export's own lines — seq, the canonical event, prev_hash, event_hash — so a capture at the receiver verifies with the shipped verifier (vxdrive-server --verify-audit-export) or the independent Python one. A receiver that stops accepting is reported to the administrators and retried; nothing is skipped. Document content is never forwarded.

How is chain of custody handled? A workspace can be put into evidence mode (Legal forensics): from then on a document's bytes leave the server through the document routes and the export only under a recorded release (a reason, or a second person's approval under dual control; the product's backups are the stated exception — they carry every byte to the destinations an administrator names, audited), removal, moving and leaving evidence mode take two people and the database refuses otherwise, and every registered item keeps a per-item custody record that is append-only, hash-chained by the database, pinned into the audit chain, exported as JSONL and verified without VX Drive. Supplied hashes are compared with the ones the instance computes, at registration, on demand and whenever a report is built; an RFC 3161 timestamp from an authority you name can cover the audit chain's anchors, the record's milestones and every report; a report is PDF/A-3u with its facts embedded as canonical JSON and validated by veraPDF.

Is redaction irreversible? A redacted copy is a fresh, image-only PDF rendered from the pages with the marks painted into the pixels, and it is refused unless two extractors and OCR find no covered word, every marked region is uniformly black, and the file has one %%EOF (no earlier version inside). The words that were covered are kept only for that verification, are never exported, and the original is never altered. What is handed over (Legal discovery's outputs) uses the verified redacted copy and replaces the covered words in the text with [REDACTED].

Does the product process biometric data? Only if a named administrator switches on voice or face comparison under a responsibility notice written for where the instance operates, and then only in workspaces where it is allowed. Every subject needs a stated legal basis and a retention date; templates are sealed under the instance key, bound to their workspace, sample and model, and never indexed — there is no gallery to search and no route that accepts a probe; a comparison is between two samples of one workspace and reports a similarity or, only under a calibration the operator fitted on a labelled set of their own, a likelihood ratio in words — never an identification; templates are destroyed at the retention date, when the workspace closes, or on a custodian's word, with a certificate. No biometric models ship: the operator registers their own with a licence note and a statement of limitations.

Does the product detect or handle illegal material? It matches files against hash sets the operator imports (nothing is bundled: NSRL, Project VIC and CAID material is the operator's under their own entitlement) and, for a set marked restricted — a regulated feature under its own notice — quarantines what matches: the document is visible only to the workspace's designated officers, its derived copies are removed and not remade, its bytes are served to an officer only under a recorded reason, it is excluded from exports and from every hosted AI route (the product's backups carry its original to the destinations an administrator names, and the backup catalogue lists its line, both audited), and the officers and administrators are told. The product says matched, never more, and the operator follows their own legal procedure. No perceptual matcher ships; an entitled operator may plug in their own.

Are the forensic methods validated or accredited? Each analyser is a versioned method with a validation pack (its method statement with limitations, its cases, a change note per version) that runs as a test on every change and again on your own worker, leaving a PDF/A-3u record; by default a method whose pack has not passed on your toolchain does not run. A finding is an indicator with its numbers, thresholds and the method's limitations beside it, never a verdict. That is evidence that the method behaves as documented; it is not accreditation under the FSR Code of Practice or ISO/IEC 17025, which remain the operator's.

Where do hash sets, models and rules come from — what do you ship? Nothing of that kind. Hash sets, YARA rulesets, detector and embedder models, validation sets for calibration, forensic images and load-file packages are the operator's, placed in directories the worker reads (read-only mounts) or entered on the settings pages; each is recorded with who added it, and a model with the licence note and limitations the operator entered. The catalog of language and vision models for search and answers is the one exception, pinned by checksum and permissively licensed, as before.

Which third-party components does the product run, and under which licences? The licence register lists them: the Rust and TypeScript dependencies (permissive, with copyleft transitive dependencies named), the tools bundled in the worker image — run as separate sandboxed processes, or loaded into the sandboxed helper process (PDFium and the ONNX runtime; the Python helpers import pikepdf, pyHanko and pytsk3), never into the server — the model catalog, the fonts and the test fixtures, and what is the operator's own to license. GPL and LGPL components are unmodified Debian packages or upstream releases at pinned versions; the register says where their source is.

What is the time source for evidence? The host's clock (run NTP) for every record, and, where the operator names an RFC 3161 authority, a timestamp token verified against that authority's chain and kept verbatim, so an auditor verifies it offline with OpenSSL.

The appliance, updates and licences (ADR-0021, ADR-0023, ADR-0024)

How is the product deployed and kept current? On one machine as three podman containers under systemd (PostgreSQL, the server, the workers), installed by one script on Debian 13 or later, Ubuntu 26.04 or later, Fedora, RHEL 10 or its rebuilds, or Arch (SELinux enforcing and firewalld honoured), reachable at one HTTPS address from every browser on the network. Updates are the host's: a timer checks the release channel daily, and applies a release only inside the hours the policy names, never held, never a pre-release unless allowed, never younger than the delay the policy sets (a fleet staggers itself behind a first machine), never a lower version, never one published after the licence's maintenance date — after a database dump, under a signature policy that requires the vendor's release key for the two VX Drive images (the PostgreSQL image is pinned by digest in its unit and never pulled by the update step; the digest-pinned ACME client image is accepted as published), with an automatic rollback when the new container does not start. Every check is recorded for the product to show; an administrator may ask for an update now. The host's own scripts follow the release too: after an applied update the release's appliance tarball is fetched from the vendor's portal with the appliance's credential, verified against the same release key with OpenSSL, and its installer re-run with every choice kept and nothing restarted (a site may hold them and refresh by hand). Air-gapped sites take each release as one signed image bundle carried in by hand — verified whole against the same release key with OpenSSL, applied with the same rollback point, health check and put-back; one bundle per engine variant, the cpu and the cuda workers image — or mirror the images and verify them with the same key through the remapIdentity policy entry the installer writes for a mirror. A machine with an NVIDIA GPU runs the engines on it by the operator's choice (vxdrivectl gpu on: the cuda variant of the workers image, the host's driver handed to the container through the container device interface).

How are releases signed, and can we verify them? Every release artifact is signed keylessly — a Sigstore bundle bound to the release workflow's identity, with a transparency-log entry — and every image three times: with the vendor's release signing key in the registry-native layout podman verifies, with the same key as a Sigstore bundle for cosign 3, and keylessly, with its SBOM attested. That key is held in the vendor's Key Vault (HSM-backed, never exported; ADR-0026) and used under a federated identity — no signing secret exists in the source-control system. The release workflow verifies its own publication in the same run — the three signatures with cosign and a pull of both images with podman under the appliance's policy — before it moves any channel tag, so a failed run leaves every channel where it was; every rehearsal proves the podman path against a registry of its own. What you can verify yourself: the images (cosign 3, with the registry credential from your install kit; docs/security/release-verification.md gives the commands and prints the public key, and the trust page prints the same key), and the appliance installer the portal serves, which comes with its .sha256 and its .sigstore.json beside it. The binaries', the web tarball's and the SBOMs' bundles and SHA256SUMS sit on the vendor's private GitHub release, which a customer cannot fetch; each release ships an SBOM for the binaries and one per image.

Where do the images come from, and who may pull them? From the vendor's own registry (registry.vertexlake.com, ADR-0025), which authenticates every pull against a credential the vendor's portal issued to the customer (three per customer by default, revocable one by one). Nothing is distributed through a personal account or a public code host; a mirror takes a credential and the same signature policy, an air-gapped site the signed image bundle the portal serves beside the installer. Non-renewal never revokes a credential: the release you run stays yours to pull and reinstall.

What does the licence control, and what happens when it lapses? A licence is a signed file the vendor issues (format 2: ECDSA P-256 through a key held in its vault; format 1 files are refused), of kind commercial or evaluation, verified offline with the public key built into the software; nothing needs to reach the vendor. It names the major version, the premium modules (Provenance, Court bundles, Legal discovery, Legal forensics) and the maintenance date. It gates two things: switching a premium module on, and applying a release published after the maintenance date. A commercial licence that verifies never stops anything: when maintenance lapses, everything that is on keeps working and can still be switched off — the software holds that as a tested invariant. An instance with no licence that verifies — never licensed, a free evaluation that ended (it serves through its last day and locks the next, UTC), or a file signed by a key or in a format the build no longer trusts — is locked to sign-in and the licence page until one is installed; while locked its workers claim nothing, so no backups run (and no overdue notice is raised), no fixity checks, no SIEM forwarding and no biometric-template destruction; nothing in it is changed or deleted by the lock. The audit trail records every licence installed (customer, modules, dates; never the text).

What does the appliance send to the vendor? With the install kit installed, the daily update step fetches the current licence for the customer and checks in with five fields — the instance's id, the running version, the release channel, the licence's digest and whether the last refresh installed one — never a document, a user, a workspace or an audit line. Without the kit (an air-gapped site, a mirror) nothing is sent and the licence file is installed by hand.

The vendor's services: the portal, the registry, Stripe (ADR-0025 to ADR-0029)

What does the vendor hold about us? The account: the organisation's name, reference, billing address and tax id; its people's names and email addresses, and the user agent of each session; the licences issued and their history; the hashes of the credentials issued (never the secrets); each appliance's check-ins (the instance id, the version, the channel, the licence's digest, whether the last refresh installed one — five fields); support conversations and enquiries from the site (an enquiry's sender address is kept thirty days, then cleared); the email address and network address of every sign-in, checkout and enquiry attempt for two days (the rate limits' memory); the mail log (which template went where with what parameters, never a rendered message); a mirror of what Stripe holds about orders, subscriptions and invoices, and the full payloads of the events Stripe sends; notes an administrator wrote about the account; and the portal's own audit chain of every act (an address that started a checkout and never bought stays in the chain after its record is swept). There is no route by which a document, a user, a workspace or an audit line of your instance could reach the vendor (ADR-0027 §4).

Where is it held, and by whom? In the vendor's Azure resource group in West US 2 (ADR-0029): a PostgreSQL Flexible Server (seven days of backups, not geo-redundant; at most two portal replicas with pools of 3 + 9 connections and the administration at 2 + 4, ADR-0029 §3b); the portal and its administration as two container apps under identities of their own, the administration IP-restricted on the platform's own name; the registry as a third app; a storage account for the registry's blobs, the installer bundles and the public-read models; Communication Services for the mail; secrets in a Key Vault that no file ever holds; the licence and release signing keys in the vault's HSM, and the key that signs the registry's five-minute pull tokens software-protected in the same vault (ADR-0026). The privacy notice at vertexlake.com/legal/privacy says the same in the customer's words.

Who takes our card? Stripe. The portal opens Stripe's hosted Checkout and hosted Customer Portal; the card never touches the vendor's systems; the vendor keeps Stripe's customer and subscription identifiers and the invoice records Stripe sends (ADR-0028).

Must we create an account to buy? No. The pricing page takes one address and the organisation's name and opens Checkout; the account is made by buying and its owner claims it from the confirmation's link, which verifies the address. What the portal then holds is the same as for any customer: the organisation, its people's names and addresses, the orders and what Stripe sends back. An address that already has an account buys for its own organisation without signing in; the start is refused — 409 sign_in_to_continue, whose message says the address has an account — only when that organisation has an open checkout, a live subscription or an earlier major (ADR-0028 as amended).

Can the vendor see our documents, or sign in to our instance? No, by construction: the instance has no route to the vendor beyond the daily licence fetch and check-in, and the vendor has no credential to the instance. Support is by conversation and by what you choose to send (the support bundle never carries document bytes).

What if the vendor disappears? Your appliance keeps running: the licence is a file on your machine verified offline, the release you run is yours to pull and reinstall for as long as the registry answers, and the perpetual fallback is in the licence agreement. What ends is updates, support and the ability to buy more. Source escrow is on the vendor's named list of things not yet offered.

Is the portal itself accountable? Every act by a customer, an administrator or a machine is on the portal's hash-chained audit trail, anchored daily and exportable in the same format as the product's, verifiable with the same tool. Administrators sign in by emailed link and a mandatory second factor; the administration is a separate table that no customer key can point at.

Where is the vendor's administration, and who can reach it? On a host of its own (ADR-0031) — the platform's default name for that container app, not published under the vendor's domain, so a scan of the domain or of the certificate logs finds nothing to try — a separate container app the same binary runs in its admin role: its ingress answers the vendor's own networks alone (every other address is refused before a request reaches the service) and the application refuses them again on the address the ingress vouches for. The public portal at vertexlake.com registers none of the administration's routes, publishes none of its operations or schemas in its API document, and its web bundle carries no administration code, markup or style — a guard on every build and the live site checks say so. The administration runs no background workers and never receives Stripe's webhooks.