Large-model private endpoints
Sovereign 192
Two 96GB-class GPU endpoints for controlled, IP-sensitive AI workloads.
Professional server GPUs, modern PCIe 5.0 platform and the memory headroom needed for larger private inference and high concurrency.
Validation-stage standard configuration.
This public proposal remains indexable while clearly qualifying its evidence state. Price, exact parts, compatibility, warranty, delivery and workload results require a current quotation and written acceptance record.
- Price status
- Indicative, ex VAT
- Commercial model
- Reviewed 25 July 2026
- Range position
- Standard progression
Buyer fit
Start with the reason to own it.
Two 96GB-class private endpoints for organisations whose accepted workload needs substantially more memory on each GPU and a tightly governed operating route.
Indicative purchase price
£97,000
ex VAT · £116,400 inc VAT
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.
A credible fit
- Large-model private inference after benchmark validation
- Separate production and test endpoints
- Controlled-network or designed air-gap deployments
Choose another route when
- A compliance certificate in a box
- An untested model or latency promise
- A business without operational ownership
Interactive system inspection
Inspect the platform behind Sovereign 192.
Move from the complete Sovereign 192 platform family into its internal layout, cooling, accelerator plane and management boundary. These views explain the engineering questions; the final ordered build and acceptance record remain decisive.
01 Platform
Start with the complete machine.
A serviceable 4U rack platform provides the physical boundary. Exact dimensions, rails, weight, power supplies and ordered components still belong in the written quotation.
02 Internal layout
See where density becomes an engineering decision.
GPU positions, processor sockets, memory, storage and cable paths compete for space and airflow. The final layout must be compatible as one system, not merely as a component list.
03 Cooling
Treat heat removal as part of the product.
High-airflow fan modules serve a tightly controlled front-to-back path. Loaded wall power, room conditions, noise and heat rejection require a measured site and reference-build record.
04 Accelerator plane
Count independent workers, not imaginary pooled memory.
A dense accelerator layout can run separate jobs or supported parallel workloads. Per-GPU memory, runtime behaviour, context and concurrency decide what the system can actually serve.
05 Operations
Make management visible before handover.
Remote health, access, sensors, logs, credentials, recovery and update ownership are part of the appliance. The final interface and permissions are recorded for the ordered platform.
Showing system view 1 of 5: Platform.
Workload route
Every use needs its own acceptance test.
Model name and aggregate VRAM do not prove a business outcome. Each proposed route below states how the capacity could be used and what must be tested before the order treats it as suitable.
01
Large-model private inference
- Proposed use
- Evaluate a model that fits within one 96GB GPU, with the second GPU available for another endpoint, staging or supported parallel execution.
- Acceptance evidence
- Test the exact model, quantisation, context, concurrency and runtime against the written quality and latency set.
02
Production and test separation
- Proposed use
- Keep one endpoint on a pinned production version while the other supports a controlled model or software change process.
- Acceptance evidence
- Verify access, version records, rollback, monitoring and service recovery before handover.
03
Controlled-network deployment
- Proposed use
- Design a restricted or offline operating route with documented model transfer, updates, backups and administrator access.
- Acceptance evidence
- Witness the agreed transfer, restart, recovery and evidence procedures in the target network design.
Specification certainty ledger
Known, proposed and still to be confirmed.
A proposal should expose missing facts. The final order replaces every confirmation row with an exact part, measured result or named customer decision.
| Item | Current value | Status | What the record must show |
|---|---|---|---|
| Form factor | Rack server | Package proposal | The site, delivery and support route is designed around this physical class. |
| Physical platform | OEM 4U PCIe 5.0 dual-socket GPU server platform | Package proposal | Platform family selected for the proposed configuration. |
| GPU route | 2 × RTX PRO 6000 Blackwell Server Edition 96GB | Package proposal | Exact manufacturer, part number, condition and serials belong in the final bill of materials. |
| Per-GPU memory | 96GB | Package proposal | The safer model-fit starting point before any supported multi-GPU test. |
| Aggregate GPU memory | 192GB across 2 GPUs | Package proposal | Not presented as one universal memory pool. |
| System memory | 256GB | Package proposal | Memory population, speed and expansion route require the final platform bill. |
| Primary storage | 4TB NVMe | Package proposal | Drive model, endurance, layout and backup destination are confirmed in the order. |
| Network | 10GbE | Package proposal | Customer switching, cabling, storage traffic and segmentation remain part of site design. |
| Power planning | Plan around 2.1kW under load; site pre-flight required | Customer decision | A qualified site review and measured reference build must replace the planning figure. |
| CPU and motherboard | To be stated in the final bill of materials | Supplier confirmation | No unverified processor, lane or motherboard claim is made on this proposal page. |
| Power supplies and leads | To be confirmed for the ordered build and UK site | Supplier confirmation | Include PSU count, rating, redundancy position, input requirements and lead specification. |
| Remote management | Platform route to be confirmed and access policy agreed | Supplier confirmation | The supplier interface shown in the gallery is a reference, not a promise of the final feature set. |
| Dimensions and weight | Exact ordered-system values required | Supplier confirmation | Placement, handling and the delivery route depend on the exact ordered-system values. |
| Hardware warranty route | Exact supplier and component warranty route required | Supplier confirmation | The final order records the whole-system, component, onsite, return and dead-on-arrival boundaries. |
96GB
Per GPU is the first sizing boundary.
Sovereign 192 has 2 GPUs and 192GB aggregate VRAM. A runtime may split a supported model or distribute independent jobs, but the page does not claim one physical 192GB memory pool. Quantisation, context, KV cache, batching and concurrency still change fit.
Software and security
The usable product is more than the chassis.
The final stack stays deliberately small. Versions, licences, access and recurring ownership are recorded so the customer is not left with an opaque collection of containers.
Proposed software baseline
- Ubuntu LTS on a recorded operating-system version
- NVIDIA driver, CUDA components and container support validated for the ordered hardware
- Docker Engine and NVIDIA Container Toolkit
- One primary model server selected from Ollama, vLLM or llama.cpp for the accepted workload
- Open WebUI or another reviewed browser interface
- Named authentication, TLS and reverse-proxy approach
- GPU, node and service monitoring with an agreed log-retention period
- Pinned versions, a software bill of materials and a model source and licence record
Security ownership
- GPU RIGS baseline
- Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
- Customer or contracted operator
- User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
- Shared before acceptance
- Model and software licence checks, retention and logging choices, recovery test, acceptance criteria and a named owner for every recurring task.
Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.
Testing and acceptance
The evidence pack is part of the machine.
No public benchmark is invented for this page. The accepted workload, exact build and disclosed test conditions decide what can be claimed after the reference system exists.
- 01
Record the final bill of materials, serial numbers and firmware versions.
- 02
Run at least 24 hours of GPU, CPU, memory and storage stress testing.
- 03
Capture temperature, fan, error, health and wall-power evidence under the agreed load.
- 04
Test cold boot, restart and the available remote-management route.
- 05
Check drive health, network throughput and the container and GPU runtime.
- 06
Run model and workload smoke tests against the written acceptance set.
- 07
Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.
- 08
Test an agreed fault or recovery route and provide the resulting handover record.
What is not claimed today
No fixed users, tokens per second, latency, model size, accuracy, availability, savings or marketplace contribution is stated without the missing configuration and test conditions.
Read the evidence methodSite, power and cooling
The room is part of the specification.
Planning power is not a measured promise. The final pre-flight records the circuit, voltage, loaded wall power, heat route, noise tolerance, rack, network, UPS decision and operating owner.
Read the site guide2.1kW
Current loaded-system planning basis
- A controlled server room or colocation position for a 4U PCIe 5.0 GPU server
- Electrical planning around a 2.1kW loaded system, subject to exact measured evidence
- Cooling, airflow, rack and UPS decisions recorded before procurement
- 10GbE baseline network planning with faster storage or fabric quoted where the workload requires it
- A named technical and security owner for identity, updates, backup, logs and incident response
Delivery, support and warranty
A written boundary before money moves.
The proposed route covers discovery, configuration, evidence and remote handover. It does not quietly absorb facilities work, migration or an always-on managed service.
Included in the proposed baseline
- Documented workload and site-fit review
- Confirmed bill of materials before procurement
- Configuration, burn-in and agreed smoke-test evidence
- Asset schedule, admin notes and user quick-start material
- Collection or the quoted kerbside or pallet-delivery route
- Remote onboarding and 30-day configuration-defect support
Separate scope or customer responsibility
- Building electrical work, rack, UPS, cooling or structured cabling
- Nationwide on-site installation unless separately quoted
- Migration of customer data, every integration or every application
- Continuous managed operations, security monitoring or a 24-hour support agreement
- Third-party model, API, marketplace or software charges
- A compliance certificate, performance guarantee or income guarantee
Supplier and warranty gates
Unresolved items stay visible.
- Exact RTX PRO 6000 Blackwell Server Edition part, condition, firmware and platform support confirmed
- Whole-system and GPU warranty scope, dead-on-arrival and return routes recorded
- Dual-GPU power, cooling and PCIe topology confirmed on the final bill of materials
- Reference-build burn-in and the agreed workload test completed before an outcome claim
- Current delivered price, lead time and UK delivery terms confirmed in the quotation
Evidence and change record
Proposal inputs remain dated.
- Package model
- UK proposal reviewed 25 July 2026
- Price state
- Calculated from dated public supplier inputs, not a live quotation
- Performance state
- No public benchmark until the exact reference build is tested
- OEM platform record
- Exact platform identity and bill of materials provided in the written quotation.
Commercial reality
Tax and spare capacity are supporting questions.
Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.
Finance, VAT and capital allowances
- The displayed price is an indicative proposal based on dated public supplier inputs. The final bill of materials and supplier quote control the order.
- Prices exclude VAT. VAT recovery depends on the buyer, its taxable activities and the normal evidence rules.
- Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
- A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
- There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Optional idle capacity
- Marketplace mode is off by default and is excluded from the purchase case.
- A separate environment, no customer data mounts, network controls and a local kill switch would be required.
- The customer, insurer, supplier warranty and marketplace terms must permit the proposed use.
- Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
- Any pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time before continuation.
Limits and alternatives
A good specification leaves room for “no”.
The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to making this package carry a workload it has not proved.
Package boundaries
- Air-gapped is an implemented operating state, not a hardware label.
- Large-model fit must be tested at the actual context, quantisation and concurrency.
- Identity, backup, patching, logging and governance remain shared responsibilities.
- Each GPU has 96GB of local VRAM. The 192GB figure is aggregate capacity.
- Air-gapped operation is a designed operating state, not an automatic chassis feature.
- Local hosting does not certify security, accuracy, confidentiality or UK GDPR compliance.
- Model fit and speed are not claimed until the exact software and workload are tested.
A better fit when several 32GB workers and mixed queues matter more than 96GB on each GPU.
Consider this route Scale 384Consider when four independent 96GB workers and greater concurrent service separation are justified.
Consider this route Hosted or hybrid AIRetain where frontier-model quality, elastic demand or the absence of an internal operator outweighs local control.
Consider this routeQuestions answered
Sovereign 192 questions that affect the order
These answers preserve the validation boundary. The written quotation and acceptance plan replace proposal assumptions.
Does Sovereign 192 provide 192GB to one model?
It provides two GPUs with 96GB each. Supported runtimes may use both, but performance and memory behaviour depend on the model and execution method. One 96GB endpoint is the safer starting assumption.
Is it automatically air-gapped?
No. An air gap needs a designed network, controlled transfer route, update process, administrator procedure, backup and recovery plan. Those controls are agreed separately.
Does local deployment make the system UK GDPR compliant?
No. Local infrastructure can change the data route, but lawful basis, access, retention, accuracy, security, processor roles and data-subject rights still need organisational controls.
Are performance figures available?
No public benchmark is claimed on this validation-stage page. The order must name the model, quantisation, context, concurrency, runtime and acceptance test before a performance result is meaningful.
Prepare the next decision
Specify the work before the parts.
Record the workload, data boundary, users, site and acceptance test. No confidential documents or credentials are needed for the first brief.
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.