Pro vs Flash: Choosing Your Sovereign Deployment Model

Same Sovereignty, Two Operating Models

Every 100xprompt deployment shares one promise: your code, data, and context never leave your perimeter. The difference between our two tiers is not privacy, it is who racks the GPUs and runs the stack. Pro puts everything on hardware you own. Flash gives you the same isolation as a fully managed, air-gapped service. Both serve leading open-source models, so the intelligence is the same either way.

100X Pro: Your Hardware, Your Control

Pro is for teams that want to own the whole stack. The models run on your GPUs, in your data center or private cloud, under your change control. You decide the hardware, the scaling strategy, and the upgrade cadence.

  • Best for: organizations with existing infrastructure, strict air-gap mandates, or sustained inference volume where owning GPUs wins on cost.
  • You manage: the GPU fleet and its capacity. Our GPU infrastructure guide covers what that takes.
  • The payoff: no per-token bill, ever. The economics are laid out in the cost case.

100X Flash: Managed and Air-Gapped

Flash is for teams that need sovereignty but do not want to operate GPUs. It runs as an isolated, air-gapped deployment, dedicated to you, with no shared tenancy and no data leaving the boundary, while we handle provisioning, scaling, and updates.

  • Best for: teams that want to move fast, have no GPU operations practice, or need to prove value before committing capital to hardware.
  • We manage: the infrastructure and the model serving; you consume it through the same 100X Code CLI.
  • The payoff: sovereign isolation with the convenience of a managed service.

How to Choose

A simple rule: if you already run infrastructure and have steady, heavy inference, Pro will cost less and give you more control. If you want sovereignty without standing up a GPU practice, start on Flash. Many teams begin on Flash to prove the workflow, then graduate to Pro as volume grows, and because both are private deployments of the same platform, moving between them is a deployment decision, not a rewrite.

Both Are Model-Agnostic

Whichever tier you pick, you are not locked to one model. Pro and Flash both serve open weights and let you swap the underlying model as better ones ship, the reason the platform is built around a harness, not a single model. Compare the tiers on the pricing page.

Keep reading → Private LLM deployment guide  ·  GPU infrastructure  ·  The cost case