Buyer's guide · pillar

On-premises AI: a buyer's guide.

A practical, hype-free guide to buying private AI you run yourself—what it is, the workloads that work today, what stays inside your network boundary, how to size it, and what evidence to demand before you accept a system.

Definition

What on-premises AI is—and when it's justified.

On-premises and private AI both describe the same commitment: the models, the data they read, and the answers they produce stay on infrastructure you control.

On-premises AI—also called private AI—means the language models, retrieval indexes, prompts, and generated answers are designed to remain inside a customer-controlled environment rather than being sent to a public AI service. A SkilakSpool deployment is a physical appliance that lives on your premises: usually a locked server room, network closet, or other controlled IT space with adequate power, cooling, and physical access control. Depending on the tier it ships as a professional tower or a rackmount server.

The approach is justified when the value of the work depends on information you cannot comfortably hand to an external model—proprietary business context, controlled records, client material, or regulated data—and when you can name the users, the data owner, the system owner, and the approval authority for that work. If a workflow only touches public information, a public tool may be enough. On-premises AI earns its keep when sensitive data, governed access, and a defensible boundary all matter at once.

It is also the right model when you want the option to run entirely offline. The core AI workload is built to run fully disconnected from the public internet; connectivity, where it exists at all, is a deliberate choice rather than a default. That property is what makes air-gapped and zero-egress operation possible instead of aspirational.

Capability

The workloads that work today.

Private AI is most useful on a focused set of proven workloads—not an open-ended promise. Start with the work that has a clear reviewer and a clear source of truth.

A few workloads are dependable enough to build a first deployment around, and they map directly to what a local model plus governed retrieval does well:

  • Private chat and drafting. A general assistant for authorized staff to think through problems, summarize, and draft—without sending the context to a public model.
  • Document Q&A with citations. Optional retrieval turns approved internal knowledge into traceable answers, where every response can be checked against the source documents it came from.
  • Code assistance. Local models help engineering teams read, explain, and write code inside the boundary, which matters when the codebase itself is sensitive.
  • OCR and vision. Specialized models extract and interpret text and imagery so scanned records and documents become searchable knowledge.

A workload is a good candidate when it depends on information the organization is permitted to use, when a reviewer can define what a good answer looks like and check the citations or output quality, and when the process is frequent or costly enough that faster retrieval or synthesis actually matters. Our right-sized private AI solutions describe how these capabilities are packaged for teams, departments, and enterprise sites.

Boundary

What actually stays inside the boundary.

A private system is only private if you can describe exactly where it sits and every path that reaches it. Four things define that boundary.

Network placement. The user interface is published only to approved internal segments over private DNS and HTTPS. The model and retrieval services stay on restricted backend networks and are not exposed directly to the internet, so there is no public inbound service to defend.

Identity. Only the users and administrators you authorize can reach the system. Deployments use local accounts or approved enterprise identity such as Microsoft Entra ID, everyday use is kept separate from administrative access, and access is aligned to least privilege with reviewable access events.

Remote access. Remote and work-from-home staff connect through your own managed remote-access path—typically a managed device, MFA, and your existing VPN or zero-trust network access (ZTNA). Skilak designs that path with your IT and security teams rather than opening a separate tunnel of its own. The companion customer-site network and access guide maps each user path and who owns it.

Offline operation. The core workload runs with no internet connection. Optional functions such as cloud SSO, software updates, monitoring, or remote support may need narrowly scoped outbound access, and every one of those destinations is documented and explicitly approved before it is enabled. The target state is no public AI or customer-data egress.

Explore the full deployment and security model

Acceptance

What to verify before you accept a system.

Acceptance is evidence, not a promise. Before you sign off, insist on demonstrations—not assurances—for each of these.

A responsible on-premises AI deployment is proven, not just delivered. These are the checks worth demanding before hardware is accepted:

  • Staging. The appliance is assembled, configured, and load-tested on a controlled staging network using synthetic or explicitly approved data before it ever reaches your site.
  • Egress verification. A disconnected rehearsal shows the approved AI workflow remains available with no public internet path, evidenced by packet capture, firewall review, and a documented destination inventory.
  • Grounded retrieval. Answers based on approved documents return traceable source citations, checked against a representative question set.
  • Backup and restore. The service and approved knowledge base can be rebuilt from the defined backup set, proven with a destroy, restore, and re-query exercise with timestamps and results.
  • Model quality. The selected model produces coherent, useful output for the agreed priority workflows against a customer-approved evaluation set.

Every engagement should define what must work and what evidence will demonstrate it before handoff, and close with a signed acceptance package plus operator training and recovery runbooks so ownership is clear after go-live.

Sizing

Sizing basics.

Four inputs drive the hardware more than anything else. Get rough answers to these before you shop for a specific box.

Users and concurrency. Named users set the overall scale, but the number of generations running at the same time drives the accelerator memory and throughput you need. A five-to-twenty-five-user pilot with a handful of active generations is a very different build from a department with dozens of concurrent sessions.

Model class. Faster small-to-mid local models suit focused teams; mid-to-large models raise quality and throughput for shared use; enterprise sites may route a portfolio of models—including specialized code, vision, OCR, and embedding models—per workflow. Larger and higher-quality models demand more memory.

Corpus size. The retrieval estate ranges from thousands of indexed chunks for a workgroup knowledge base up to millions of chunks for a multi-domain estate with governed ingestion and permissions. Bigger corpora change both storage and retrieval design.

Treat any published tier range as a planning baseline, not a guaranteed limit. Final capacity depends on model and quantization, context length, retrieval design, data volume, concurrent use, latency targets, licensing, facilities, and acceptance testing. The fastest way to a starting point is the private AI deployment configurator, which turns seven quick choices into a practical tier.

Compliance

Compliance reality.

An appliance can fit inside a compliance program. It cannot hand you compliance. Be skeptical of any vendor that says otherwise.

SkilakSpool is designed to operate inside your existing CMMC, HIPAA, ITAR, or other security program—its private-by-default network posture, governed identity, approved-egress model, and customer-controlled data lifecycle are meant to slot into controls you already run. That is the honest ceiling of what a product can do.

No product by itself creates compliance. Certification, accountability, and the evidence your assessors require remain with your security program. A system that keeps sensitive data on infrastructure you own removes one large obstacle to using AI on regulated work, but the policies, the people, and the audit trail are still yours to own. When a claim sounds like “buy this and you are compliant,” treat it as a warning sign.

If that framing fits how your organization thinks about risk, request a private AI briefing and bring one real workload. That is the fastest way to test whether an on-premises deployment is justified for you.

Ready to scope it?

Bring one workload to the table.

Tell us who uses it, what it can read, and what must stay inside. We'll map the users, boundary, and proof needed for a practical first deployment.