Private inference at organisational scale
Scale 384
Four independent 96GB workers for demanding concurrency and workload isolation.
A flagship four-GPU server for AI consultancies, engineering teams and organisations running several large private services.
Validation-stage standard configuration.
This public proposal remains indexable while clearly qualifying its evidence state. Price, exact parts, compatibility, warranty, delivery and workload results require a current quotation and written acceptance record.
- Price status
- Indicative, ex VAT
- Commercial model
- Reviewed 25 July 2026
- Range position
- Standard progression
Buyer fit
Start with the reason to own it.
Four independent 96GB workers for organisations that can justify several large private services, workload isolation and the facility and operating ownership that follow.
Indicative purchase price
£172,000
ex VAT · £206,400 inc VAT
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.
A credible fit
- Several large concurrent private AI services
- AI consultancies and private inference operators
- Organisations with proper rack, power and cooling
Choose another route when
- A general office environment
- An H100/H200 training-cluster requirement
- A buyer expecting NVLink or universal memory pooling
Interactive system inspection
Inspect the platform behind Scale 384.
Move from the complete Scale 384 platform family into its internal layout, cooling, accelerator plane and management boundary. These views explain the engineering questions; the final ordered build and acceptance record remain decisive.
01 Platform
Start with the complete machine.
A serviceable 4U rack platform provides the physical boundary. Exact dimensions, rails, weight, power supplies and ordered components still belong in the written quotation.
02 Internal layout
See where density becomes an engineering decision.
GPU positions, processor sockets, memory, storage and cable paths compete for space and airflow. The final layout must be compatible as one system, not merely as a component list.
03 Cooling
Treat heat removal as part of the product.
High-airflow fan modules serve a tightly controlled front-to-back path. Loaded wall power, room conditions, noise and heat rejection require a measured site and reference-build record.
04 Accelerator plane
Count independent workers, not imaginary pooled memory.
A dense accelerator layout can run separate jobs or supported parallel workloads. Per-GPU memory, runtime behaviour, context and concurrency decide what the system can actually serve.
05 Operations
Make management visible before handover.
Remote health, access, sensors, logs, credentials, recovery and update ownership are part of the appliance. The final interface and permissions are recorded for the ordered platform.
Showing system view 1 of 5: Platform.
Workload route
Every use needs its own acceptance test.
Model name and aggregate VRAM do not prove a business outcome. Each proposed route below states how the capacity could be used and what must be tested before the order treats it as suitable.
01
Several large private endpoints
- Proposed use
- Assign separate 96GB GPUs to teams, customers, services or model versions with a defined isolation and scheduling policy.
- Acceptance evidence
- Test representative concurrent demand, service priority, failure isolation and recovery.
02
Supported parallel inference
- Proposed use
- Evaluate tensor or pipeline parallel execution only where the selected model, runtime and PCIe topology support it.
- Acceptance evidence
- Record per-GPU memory, communication overhead, throughput and latency under disclosed conditions.
03
Consultancy or operator capacity
- Proposed use
- Provide customer-dedicated or project-dedicated endpoints under separately defined access, retention, support and commercial terms.
- Acceptance evidence
- Prove tenant boundaries, monitoring, support ownership, capacity policy and data removal.
Specification certainty ledger
Known, proposed and still to be confirmed.
A proposal should expose missing facts. The final order replaces every confirmation row with an exact part, measured result or named customer decision.
| Item | Current value | Status | What the record must show |
|---|---|---|---|
| Form factor | Rack server | Package proposal | The site, delivery and support route is designed around this physical class. |
| Physical platform | OEM 4U PCIe 5.0 dual-socket GPU server platform | Package proposal | Platform family selected for the proposed configuration. |
| GPU route | 4 × RTX PRO 6000 Blackwell Server Edition 96GB | Package proposal | Exact manufacturer, part number, condition and serials belong in the final bill of materials. |
| Per-GPU memory | 96GB | Package proposal | The safer model-fit starting point before any supported multi-GPU test. |
| Aggregate GPU memory | 384GB across 4 GPUs | Package proposal | Not presented as one universal memory pool. |
| System memory | 512GB | Package proposal | Memory population, speed and expansion route require the final platform bill. |
| Primary storage | 8TB NVMe | Package proposal | Drive model, endurance, layout and backup destination are confirmed in the order. |
| Network | 10GbE baseline; faster fabric quoted where needed | Package proposal | Customer switching, cabling, storage traffic and segmentation remain part of site design. |
| Power planning | Plan around 3.3kW under load; technical facility required | Customer decision | A qualified site review and measured reference build must replace the planning figure. |
| CPU and motherboard | To be stated in the final bill of materials | Supplier confirmation | No unverified processor, lane or motherboard claim is made on this proposal page. |
| Power supplies and leads | To be confirmed for the ordered build and UK site | Supplier confirmation | Include PSU count, rating, redundancy position, input requirements and lead specification. |
| Remote management | Platform route to be confirmed and access policy agreed | Supplier confirmation | The supplier interface shown in the gallery is a reference, not a promise of the final feature set. |
| Dimensions and weight | Exact ordered-system values required | Supplier confirmation | Placement, handling and the delivery route depend on the exact ordered-system values. |
| Hardware warranty route | Exact supplier and component warranty route required | Supplier confirmation | The final order records the whole-system, component, onsite, return and dead-on-arrival boundaries. |
96GB
Per GPU is the first sizing boundary.
Scale 384 has 4 GPUs and 384GB aggregate VRAM. A runtime may split a supported model or distribute independent jobs, but the page does not claim one physical 384GB memory pool. Quantisation, context, KV cache, batching and concurrency still change fit.
Software and security
The usable product is more than the chassis.
The final stack stays deliberately small. Versions, licences, access and recurring ownership are recorded so the customer is not left with an opaque collection of containers.
Proposed software baseline
- Ubuntu LTS on a recorded operating-system version
- NVIDIA driver, CUDA components and container support validated for the ordered hardware
- Docker Engine and NVIDIA Container Toolkit
- One primary model server selected from Ollama, vLLM or llama.cpp for the accepted workload
- Open WebUI or another reviewed browser interface
- Named authentication, TLS and reverse-proxy approach
- GPU, node and service monitoring with an agreed log-retention period
- Pinned versions, a software bill of materials and a model source and licence record
Security ownership
- GPU RIGS baseline
- Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
- Customer or contracted operator
- User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
- Shared before acceptance
- Model and software licence checks, retention and logging choices, recovery test, acceptance criteria and a named owner for every recurring task.
Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.
Testing and acceptance
The evidence pack is part of the machine.
No public benchmark is invented for this page. The accepted workload, exact build and disclosed test conditions decide what can be claimed after the reference system exists.
- 01
Record the final bill of materials, serial numbers and firmware versions.
- 02
Run at least 24 hours of GPU, CPU, memory and storage stress testing.
- 03
Capture temperature, fan, error, health and wall-power evidence under the agreed load.
- 04
Test cold boot, restart and the available remote-management route.
- 05
Check drive health, network throughput and the container and GPU runtime.
- 06
Run model and workload smoke tests against the written acceptance set.
- 07
Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.
- 08
Test an agreed fault or recovery route and provide the resulting handover record.
What is not claimed today
No fixed users, tokens per second, latency, model size, accuracy, availability, savings or marketplace contribution is stated without the missing configuration and test conditions.
Read the evidence methodSite, power and cooling
The room is part of the specification.
Planning power is not a measured promise. The final pre-flight records the circuit, voltage, loaded wall power, heat route, noise tolerance, rack, network, UPS decision and operating owner.
Read the site guide3.3kW
Current loaded-system planning basis
- A technical facility or colocation position suitable for a dense 4U four-GPU server
- Electrical planning around a 3.3kW loaded system, subject to exact measured evidence
- Cooling, rack, PDU, UPS and service-access decisions completed before procurement
- Network and storage architecture checked for the buyer's concurrency, data and backup pattern
- Named owners for scheduling, identity, monitoring, updates, incidents, capacity and recovery
Delivery, support and warranty
A written boundary before money moves.
The proposed route covers discovery, configuration, evidence and remote handover. It does not quietly absorb facilities work, migration or an always-on managed service.
Included in the proposed baseline
- Documented workload and site-fit review
- Confirmed bill of materials before procurement
- Configuration, burn-in and agreed smoke-test evidence
- Asset schedule, admin notes and user quick-start material
- Collection or the quoted kerbside or pallet-delivery route
- Remote onboarding and 30-day configuration-defect support
Separate scope or customer responsibility
- Building electrical work, rack, UPS, cooling or structured cabling
- Nationwide on-site installation unless separately quoted
- Migration of customer data, every integration or every application
- Continuous managed operations, security monitoring or a 24-hour support agreement
- Third-party model, API, marketplace or software charges
- A compliance certificate, performance guarantee or income guarantee
Supplier and warranty gates
Unresolved items stay visible.
- Exact four-GPU compatibility, PCIe topology, firmware, power and thermal support confirmed in writing
- Whole-system and GPU warranty scope, dead-on-arrival and return routes recorded
- Higher-speed networking and storage options quoted where the accepted workload requires them
- Reference-build burn-in and concurrent workload evidence completed before an outcome claim
- Current delivered price, lead time and UK delivery terms confirmed in the quotation
Evidence and change record
Proposal inputs remain dated.
- Package model
- UK proposal reviewed 25 July 2026
- Price state
- Calculated from dated public supplier inputs, not a live quotation
- Performance state
- No public benchmark until the exact reference build is tested
- OEM platform record
- Exact platform identity and bill of materials provided in the written quotation.
Commercial reality
Tax and spare capacity are supporting questions.
Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.
Finance, VAT and capital allowances
- The displayed price is an indicative proposal based on dated public supplier inputs. The final bill of materials and supplier quote control the order.
- Prices exclude VAT. VAT recovery depends on the buyer, its taxable activities and the normal evidence rules.
- Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
- A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
- There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Optional idle capacity
- Marketplace mode is off by default and is excluded from the purchase case.
- A separate environment, no customer data mounts, network controls and a local kill switch would be required.
- The customer, insurer, supplier warranty and marketplace terms must permit the proposed use.
- Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
- Any pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time before continuation.
Limits and alternatives
A good specification leaves room for “no”.
The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to making this package carry a workload it has not proved.
Package boundaries
- This is not presented as an H100/H200 training cluster.
- Higher-speed network and shared storage may be required and are not in the baseline.
- Procurement economics and exact workload evidence require enterprise-buyer validation.
- Each GPU has 96GB of local VRAM. The 384GB figure is aggregate capacity.
- The system is not presented as an H100 or H200 training cluster and no NVLink claim is made.
- 10GbE is a baseline, not an assurance that every data-heavy workload has enough network or storage throughput.
- No multi-tenant, service-level, availability or performance promise exists without a separate operating design and contract.
A more proportionate route when two 96GB endpoints provide enough memory and service separation.
Consider this route Custom specificationUse when storage, network fabric, redundancy, colocation or service design differs from the proposed baseline.
Consider this route Elastic hosted capacityMay be preferable for bursty work, uncertain utilisation or teams that cannot own a dense server platform.
Consider this routeQuestions answered
Scale 384 questions that affect the order
These answers preserve the validation boundary. The written quotation and acceptance plan replace proposal assumptions.
Is 384GB available as one GPU memory pool?
No. The proposal uses four 96GB GPUs. Some runtimes can split supported models or jobs across GPUs, but the page does not claim one universal 384GB pool.
Is Scale 384 an AI training cluster?
It is proposed as a dense private inference and independent-worker platform. It is not presented as an H100 or H200 training cluster, and no NVLink capability is claimed.
Does the price include a data-centre installation?
No. Rack, PDU, electrical work, UPS, cooling, structured cabling, colocation work and on-site installation need a separate scope where required.
Can spare capacity earn income?
It can only be assessed as an opt-in pilot after security, warranty, insurer and availability review. Marketplace acceptance, demand, rates and income are not guaranteed and are excluded from the purchase case.
Prepare the next decision
Specify the work before the parts.
Record the workload, data boundary, users, site and acceptance test. No confidential documents or credentials are needed for the first brief.
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.