Comparison · buying

Four ways to run AI on data that cannot leave.

Most comparison pages are written by a vendor who wins every row. This one is not. Three of the four options below are better than us for some buyer, and this page says which buyer and why.

Position

The honest version of this page has us losing some rows.

Every one of these four is the right answer for somebody. The mistake is choosing on vendor enthusiasm instead of on which boundary your requirement actually describes.

If you hold CUI, ITAR, PHI, or privileged material, you have four realistic ways to put a language model in front of it. You can license Microsoft 365 Copilot inside a GCC High tenant. You can buy a system that runs on hardware you own. You can have your own team build one. Or you can put per-seat AI software on each person’s laptop.

So this page states what each option genuinely does well, sources the facts about it to the vendor’s own documentation rather than to our characterisation of it, and then says plainly where we are the wrong purchase.

One thing to settle before the table makes sense. These four do not differ mainly in features. They differ in what kind of boundary they give you. Copilot in GCC High gives you a contractual and configurable boundary, backed by real federal accreditation and real personnel screening. A per-seat laptop tool gives you a device boundary, one per user. A system on your own hardware gives you a physical boundary you can put a packet capture in front of. Which of those satisfies your obligation is a question for your security program, not for a vendor.

The comparison

Eight rows that decide it.

Claims about other vendors are quoted from their own documentation. Claims about us are traced to published results, and where we have not measured something we say so.

 SkilakSpoolCopilot in GCC HighDIY buildPer-seat local
Where the data livesOn hardware you own, inside your boundary. Prompts, retrieval indexes, and answers stay on the system.In Microsoft's US government cloud. If web grounding is turned on, Copilot sends a generated search query to Bing, which Microsoft says “operates separately from Microsoft 365 and has different data-handling practices.”On hardware you own — if every telemetry default in the stack was found and switched off.On each user's laptop, and one copy of the corpus per device.
What you must trustA packet capture you run at your own firewall while the system is exercised, plus a written egress mode in the SOW. You re-run it; you do not take our word.Microsoft's contractual commitments, its personnel screening, and your own tenant configuration. Microsoft states GCC High support “isn't included in the service accreditation boundary,” and that Customer Lockbox does not protect against requests from law enforcement or other third parties.Your own team's thoroughness. The two worst defects we found building ours were invisible to every health check.The vendor's statement that the software makes no outbound connection — unless you capture the traffic yourself.
Per-seat AI cost over timeNo per-seat AI licence. One caveat that applies to us exactly as it applies to a DIY build: Open WebUI requires its branding to remain visible above 50 users in a rolling 30 days without an enterprise licence — a line item to settle, not one to pretend away.Recurring per user, indefinitely: a qualifying base licence plus a Copilot add-on. Microsoft does not publish a GCC High rate; the commercial enterprise list is $30 per user per month on an annual commitment.No AI licence fee — it is the same open-source software. The cost is staff time.Per device or per user, and the floor is genuinely $0. LM Studio became free for commercial use in July 2025; Ollama, AnythingLLM, GPT4All and Jan are free. Paid tiers exist — Msty lists $149 per user/year.
What happens disconnectedRuns. The core workload is built to operate with no public internet path. The exception is cloud SSO — see below.Stops. It is a cloud service.Runs, if it was staged correctly. Staging that stops short leaves failures that only surface later, in front of a user.Runs. This is the category's real strength.
Who operates itYou do, with runbooks, operator handoff, and a support contract. Support access is a designed, revocable path — no permanent remote agent or unmanaged support account.Microsoft operates the platform; you configure it. Microsoft: “Customers remain responsible for configuring Microsoft 365 and Microsoft 365 Copilot to meet their specific regulatory obligations.”You do, entirely — drivers, accelerator state, backups, restores, and the first-boot behaviour of a large model, which does not look like success while it is happening.Each user, on their own machine. In the standalone per-seat deployment there is no central administrator.
Evidence you get at acceptanceA live packet capture at your firewall, a validation suite that asserts answer coherence rather than service health, a restore drill that ends in a real query, a port and flow matrix, and a documented responsibility split.Platform accreditation you build on: NIST SP 800-53 at a FIPS 199 High categorization, which Microsoft describes as “equivalency to IL4.” It reduces some of your assessment burden; it does not remove it. Note this is not IL5, and you cannot pilot first — “Trials aren't available at this time.”Whatever you produce yourself. Nothing ships with the software.The vendor's product claims. Nothing runs at your firewall by default.
Model size and speedMeasured on rented hardware: 43.6 tok/s single-stream on a 72B-class model on one H100-class GPU, rising to 1,262 tok/s aggregate at concurrency 32, TTFT p95 0.20s. On A10-class with a 7B: 30.6 and 883.Microsoft's hosted frontier models. Capacity is Microsoft's problem, not yours — a genuine advantage.The same models on the same hardware. We have not benchmarked other runners and will not claim a comparison we have not run. Concurrency depends on the serving architecture you choose.Bounded by laptop silicon. The NITRO technical report (Cornell, Dec 2024 — an arXiv preprint, not peer-reviewed) measured Llama3-8B at about 2.9 tok/s on an Intel Core Ultra NPU; the same laptop's integrated GPU was faster. One 2024 generation, one silicon path — newer laptops are faster and we have not measured them.
Where it shows up in the workA private browser interface with document Q&A and citations. Not inside Word, Excel, or Outlook.Inside Word, Excel, Outlook and PowerPoint, reasoning across Microsoft Graph. Though in GCC High specifically, Microsoft lists Copilot in Teams, Copilot in SharePoint, and SharePoint agents as “Not currently available” (as of July 2026) — while all three are available in the DoD cloud.Whatever you build.A desktop application on the user's own machine.

Microsoft 365 Copilot in GCC High

It is a serious product, and pretending otherwise would be dishonest.

Copilot reached general availability in GCC High in December 2025. It is not a thin government port of a commercial toy, and a comparison that treated it as one would deserve to be ignored.

What it genuinely gets right, in Microsoft’s own words rather than ours: prompts, responses, and Microsoft Graph data “aren’t used to train foundation models.” Copilot “respects your identity model and permissions, inherits your sensitivity labels, applies your retention policies, supports audit of interactions” — meaning an organisation that has already invested in Purview gets governance an on-premises stack has to rebuild from scratch. Microsoft staff have no standing access to GCC High production, and temporary elevation requires screening against US citizenship, a seven-year criminal history check, an FBI fingerprint check, and the OFAC, Bureau of Industry and Security, and DDTC Debarred Persons lists. That is a substantive ITAR personnel control, and anyone implying Microsoft is careless here is not worth listening to.

There are three things in Microsoft’s own documentation a CUI or ITAR buyer should read before signing, and none requires an adversarial reading.

The support channel is outside the boundary. Microsoft’s service description states that GCC High and DoD support “isn’t included in the service accreditation boundary and doesn’t provide FedRAMP, DOD SRG, ITAR, IRS 1075, or CJIS data handling compliance assurances,” and in the same section reminds customers “not to share any controlled, sensitive, or confidential information with customer support personnel as part of your support incident when using Office 365 GCC High/DOD, at least until you confirm the support agent’s authorization to view or access such data.” Read it as written: the accreditation does not extend to the support channel, and clearing an individual agent is a step your people take themselves, mid-incident, at the moment something is broken and they most want to paste in what they are seeing.

Web grounding is a policy setting, not an architectural constraint. It is off by default in government clouds — correctly. But an administrator can enable it, after which individual users get a toggle that defaults on. Microsoft documents that the generated search query then goes to Bing, that Microsoft acts as “an independent data controller” for those queries rather than a processor, that the Data Protection Addendum “doesn’t apply to the use of generated web search queries,” and that the query “may be informed by data within a Microsoft 365 document.” The whole file is not sent; themes drawn from it can be. The question is not whether the default is safe — it is whether your boundary can depend on a setting one administrator can change.

GCC High is not IL5. Microsoft describes it as assessed against NIST SP 800-53 at a FIPS 199 High categorization, demonstrating “equivalency to IL4.” Only the separate Office 365 DoD environment is certified at SRG L5, and only DoD entities may purchase it. Buyers who assume GCC High equals IL5 are mistaken, and it is better to learn that here than in an assessment.

Microsoft itself concedes the lag: “Feature availability may differ; Release timing typically lags behind commercial environments; Third-party integrations are more restricted in higher-isolation environments.” And the isolated tenant carries a collaboration tax Microsoft documents plainly — GCC High users can share only with other GCC High organisations.

When GCC High is the right answer, and you should buy it instead of talking to us: you are already in a GCC High tenant, you have no air-gap requirement, and the value you want is Copilot inside Word, Excel, Outlook and PowerPoint reasoning across your mail, calendar and files. In that situation Copilot is one add-on line item — no capital purchase, no refresh cycle, no facilities work, no operations headcount. We cannot offer Office-native authoring or Graph-wide organisational context, and we are not going to pretend the browser is the same thing.

Building it yourself

The software is free. The boundary work is not.

There is a version of this argument we refuse to make: that Ollama is the wrong tool. Our own Solo bill of materials lists the serving layer as vLLM or Ollama. The honest argument is about concurrency, evidence, and deadline.

Where DIY genuinely wins. If you have one or two habitual users, real Linux and GPU competence in house, and no assessor waiting on an artifact, you are not obviously better off buying. Build it. We would rather tell you that than sell you a rack you did not need.

Where it stops being comparable. The serving architecture is what shows up under load — vLLM attributes its throughput to continuous batching, chunked prefill, prefix caching, and PagedAttention memory management. That is the difference between 43.6 tok/s for one person and 1,262 tok/s aggregate across 32 of them on the same GPU. If more than a handful of people will use this at once, the choice of stack stops being a preference.

And then there is the part nobody prices. These are defects we hit building our own system, on our own rented hardware — not stories about customers, because we do not have customers yet.

During our zero-egress rehearsal, the only finding was invisible to every health check and every UI test. It was not AI traffic and it carried no prompt, document, or answer — just a trickle of metadata leaving a system every dashboard said was healthy. Packet capture found it. Nothing else did. Separately, on a virtualised multi-GPU slice, the model emitted token soup at full speed while passing every health check: the maths underneath had quietly gone wrong. That run is marked invalid in our records and its numbers appear nowhere on this site. It is also why our validation suite asserts that the model returns a coherent completion, not merely that the service is up.

Neither defect would have been caught by the tests a competent DIY team would think to run. Health checks pass. The system is still wrong.

The rest is work you can price against your own team’s time, once you know it exists. Several components in a stack like this reach outbound by default, each for its own reason and each disabled in its own place — a build inherits every one of those defaults unless somebody knows to go looking. Image versions have to be pinned, or routine start-up itself reaches out to a registry, which is exactly the traffic an acceptance test is meant to prove does not happen. Backups have to cover every stateful part of the system together or the restore is not coherent. And the restore has to be drilled, because a backup that has never been restored is a hope, not a backup.

Then there is the paperwork the box creates rather than solves. If the system holds CUI, NIST SP 800-171 3.12.4 obliges you to produce a system security plan describing boundaries, environments of operation, how requirements are implemented, and connections to other systems. A DIY GPU box creates that obligation and supplies nothing toward satisfying it.

One last thing that is not an argument for us at all. SentinelLABS and Censys published a joint investigation in January 2026 finding 175,108 unique publicly exposed Ollama hosts across 130 countries over 293 days of scanning, with a persistent core of roughly 23,000. Their assessment was that tool-enabled endpoints executing privileged operations, “when combined with insufficient authentication and network exposure,” create “what we assess to be the highest-severity risk in the ecosystem.” Ollama binds to localhost by default and documents no built-in authentication — so the single step every DIY team takes to let colleagues reach the box is the same step that produced that population. Access control has to come from something in front of it. That is true whether you buy from us or not.

We will not tell you how long a DIY build takes. No credible source for that number exists, and we are not going to invent one. Price it from the list above.

Per-seat AI on laptops

Right for a small number of people. Structurally wrong for a shared corpus.

For roughly five to fifteen people summarising their own documents on hardware they already own, this is the correct purchase and we would say so in the room.

There is no server, no rack, no power and cooling review, no deployment, and no acceptance test. The market floor is genuinely $0 — LM Studio dropped its commercial-use licence requirement in July 2025, and AnythingLLM, GPT4All, Jan and Ollama are free and open source. Commercial products exist above that; Msty publishes $149 per user per year. Any page arguing that per-seat local AI is expensive would be wrong.

Two things bound it, and both are structural rather than fixable by a better release.

Silicon. Laptop silicon bounds you to small models. Our measured record is 43.6 tokens per second on a 72B-class model on an H100-class GPU. That is a different class of work.

Shared retrieval. This matters more than speed. In a standalone per-seat deployment, each laptop holds its own copy of the corpus. There is no shared index, no central permission model, and no single audit log. You cannot answer “who queried which document, when” from one place. You cannot enforce document-level permissions centrally. And you cannot revoke a departing employee’s embedded copy of the document set. A shared system has one index, one permission model, one log — which is exactly what a CMMC, ITAR, or HIPAA assessor asks about. Some products offer an optional in-network relay; the limitation above describes the standalone deployment, which is how most are actually bought.

We are also not going to claim a per-seat vendor leaks. We have not captured their traffic and neither, probably, have you. The distinction we will make is about method of proof: their zero-egress story is a stated property of the software, and ours is a test run at your firewall with your network team reading the output. If a per-seat vendor will run that test at your site, that is a good sign and you should take it.

Where we lose

The gaps, stated here rather than discovered later.

A comparison page that concludes the author wins everything is worthless. Here is what is missing on our side.

We have no customers yet. Every failure mode on this page was found on our own rented hardware while building the system, not at a customer site. That is first-hand and documented, but it is not field experience and we will not dress it up as such.

We do not publish multi-GPU throughput. The one multi-GPU run we have is marked invalid for the reason described above. Numbers appear on this site after they validate, not before.

Our boundary has one documented exception. If you want cloud identity, the system needs controlled outbound access to your identity provider for sign-in and token operations. The honest target in that mode is zero outbound AI or data traffic with the identity destinations captured, logged, and listed as a named exception. Literal zero-packet operation requires local identity. That decision is recorded in the statement of work before anything is configured — not discovered on acceptance day.

We do not certify compliance. SkilakSpool is designed to operate inside your CMMC, HIPAA, or ITAR program. Skilak Consulting does not certify or confer compliance; that determination belongs to you and your assessor.

Ask all four of us

Three questions that are fair to every vendor here.

None of this is proprietary. Whoever you are evaluating — us included — these are reasonable and non-negotiable.

  1. Which egress mode does the system run, in writing?

    Fully disconnected, or cloud identity with documented exceptions. Get it in the statement of work. A vendor who cannot commit to a mode has not thought about your boundary.

  2. Will you run a live capture at our firewall, during acceptance?

    Evidence you can watch beats any assurance you are asked to accept. Ours is described in full on how we prove nothing leaves.

  3. What leaves, and to where?

    If anything does — cloud identity, updates, support, web grounding — demand the list of destinations and the reason for each. Honesty about exceptions is the tell of a serious offer.

See what we have actually measured

Still deciding?

Bring us the option you are leaning toward.

Tell us your boundary, your user count, and your deadline. If GCC High or a build of your own is the better answer, we will say so — and tell you what to ask them for.