Report · evidence

Acceptance evidence: our published test results.

We would rather show you results than adjectives. Every item below lists its method, its result, the date it was verified, and the hardware it was measured on—so the claims are reproducible at your own acceptance rather than taken on faith.

Position

Why we publish acceptance evidence.

A marketing number with no method behind it is decoration. These are the results we are willing to stand on, each traceable to how it was produced.

Private-AI marketing is full of numbers with no lineage: a throughput figure with no model named, a “never leaves your network” claim with no capture behind it, a “disaster-recovery ready” badge for a backup that has never once been restored. We publish the opposite. Each result on this page states what was tested, the method in a line or two, the result itself, the date it was verified, and—just as important—the hardware and context it was measured on, because a number without its hardware is meaningless.

None of this is meant to be taken on our word. The point of an evidence page is that the same procedures are re-run at your acceptance, on your boundary, with your people watching. What follows is what we have actually run so far, reported honestly, including a plain section on what we are not yet willing to publish. Where a figure is not in our own records, it is not on this page.

The results

What we ran, and what it produced.

Each item is a validated result with its own verification date. Benchmarks were driven by a stdlib load generator that streams chat completions with a roughly 500-token prompt and 256-token completions, excludes warmup, and counts every error.

Zero-egress rehearsal — clean capture, zero packets

Method. After staging a system and flipping it to its offline configuration, we logged every new outbound connection to a public address at the host while exercising the stack—model chat, a retrieval data path, a service restart, and a short load test. Two captures were taken, 240 seconds and 100 seconds.

Result. Zero outbound connections to any public address from the AI stack across both captures; the second capture was completely silent—0 packets, 0 bytes of payload. Ordinary internal name resolution stays inside the network and is handled within the standard build. This was a staging rehearsal on our own network, not a customer deployment. The full method is in the zero-egress verification write-up.

Verified 2026-07-16 · rehearsed on our staging network, on a rented A10-class GPU.

Single-GPU serving benchmark — small model (A10-class)

Method. A Qwen2.5-7B model served on the reference stack, swept by the load generator from concurrency 1 through 32 with coherence checked before publication.

Result. 30.6 tokens/sec at concurrency 1, rising to 883.3 tokens/sec aggregate at concurrency 32, with a time-to-first-token p95 of 0.267s at concurrency 32. Zero failed requests across the entire sweep.

Verified 2026-07-16 · measured on a rented A10-class GPU (24 GB), a proxy for the entry configuration.

Single-GPU serving benchmark — large quantized model (H100-class)

Method. A Qwen2.5-72B AWQ-quantized model on a single accelerator, swept by the same load generator from concurrency 1 through 32, coherence checked before publication.

Result. 43.6 tokens/sec at concurrency 1, rising to 1,262 tokens/sec aggregate at concurrency 32, with a time-to-first-token p95 of 0.20s at concurrency 32. Zero failed requests across the entire sweep.

Verified 2026-07-17 · measured on a single rented H100-class GPU (80 GB).

Document Q&A, end to end

Method. A document containing a distinctive fact was ingested through the interface, then the same fact was requested back through retrieval.

Result. The system returned the fact with a citation to the source document—an ingest-to-cited-answer path exercised end to end, not a component check.

Verified 2026-07-17 · exercised during staging on a rented H100-class GPU.

Backup and restore drill — full cycle, PASSED

Method. Ingest a document and confirm a cited answer; take a verified backup; destroy live data by deleting all vector-store collections; restore from the backup after its integrity manifest verifies; re-query.

Result. PASSED. After restore, the collections were byte-identical, administrative login worked, and the retrieval query returned the same fact with its citation. A backup that has never been restored is a hope, not a backup—so the drill proves the restore with a real query, not just service health. The procedure is in the backup and restore runbook.

Verified 2026-07-17 · rehearsed on a rented H100-class GPU during staging.

Stack validation suite — health plus answer coherence

Method. An automated suite run against a freshly provisioned stack: service health checks plus an answer-coherence assertion, because a stack can report every service healthy while still returning low-quality output.

Result. The first end-to-end reference run passed 6 of 6 checks on rented A10-class hardware. The suite also adds a monitoring-scrape verification and the output-coherence assertion, so acceptance confirms the model answers coherently, not merely that services are up.

Verified 2026-07-17 · initial 6/6 run on a rented A10-class GPU (2026-07-16); coherence and scrape checks verified in the current stack revision.

Honesty

What we don't publish yet.

The absence of a number here is deliberate. We would rather have a gap than a figure we cannot defend.

Multi-GPU benchmark results. We publish these once they are validated on NVLink-class production hardware. Until then, no multi-GPU figures appear anywhere on this page.

Customer-site results. Acceptance results from a customer deployment are published only with that customer’s written release. Everything above was produced on our own rented and staging hardware with synthetic data, and is labeled as such.

Freshness

How current is this page?

Evidence is re-verified when the stack materially changes, not on a calendar. Each result carries its own verification date and the stack revision it reflects.

We do not re-run these procedures on a schedule for the sake of a fresher date. They are re-verified when the underlying stack changes in a way that could affect the result—a change to the serving configuration, the telemetry controls, the backup tooling, or the validation suite. Each item above shows the date it was last verified, so you can see exactly how old each claim is rather than trusting a single “last updated” stamp for the whole page.

Evidence reflects stack revision 89744f5 (2026-07-17). SkilakSpool is designed to fit within a customer’s CMMC, HIPAA, ITAR, or other security program. Skilak Consulting does not certify or confer compliance.

Go deeper

The full evidence pack.

The results above are the summary. The raw material behind them is available in a briefing.

If you are evaluating seriously, the summary is not enough and it should not be. The full evidence pack—the capture logs behind the zero-egress rehearsal, the drill record for backup and restore, and the raw benchmark results at every concurrency level—is shared in a briefing so it can be read in context and matched to your boundary.

Request the full evidence pack

Prefer to watch it live?

Make us reproduce it at acceptance.

Every result here is designed to be re-run on your boundary, with your team watching. Bring the workload and we'll show you the proof.