Multi-workload company capacity
Business 128
Four GPU workers for teams that need throughput, separation and room to grow.
A dense multi-GPU platform for concurrent private assistants, mixed AI services, render queues and supported model sharding.
Validation-stage standard configuration.
This public proposal remains indexable while clearly qualifying its evidence state. Price, exact parts, compatibility, warranty, delivery and workload results require a current quotation and written acceptance record.
- Price status
- Indicative, ex VAT
- Commercial model
- Reviewed 25 July 2026
- Range position
- Standard progression
Buyer fit
Start with the reason to own it.
A multi-worker company platform for organisations that need several private services or queues at the same time and have a suitable technical facility.
Indicative purchase price
£70,500
ex VAT · £84,600 inc VAT
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.
A credible fit
- Several independent model endpoints
- Company RAG plus creative or batch workloads
- Teams with a proper server room or colocation plan
Choose another route when
- A quiet under-desk appliance
- A single 128GB workload without runtime validation
- A site without electrical and cooling capacity
Interactive system inspection
Inspect the platform behind Business 128.
Move from the complete Business 128 platform family into its internal layout, cooling, accelerator plane and management boundary. These views explain the engineering questions; the final ordered build and acceptance record remain decisive.
01 Platform
Start with the complete machine.
A serviceable 4U rack platform provides the physical boundary. Exact dimensions, rails, weight, power supplies and ordered components still belong in the written quotation.
02 Internal layout
See where density becomes an engineering decision.
GPU positions, processor sockets, memory, storage and cable paths compete for space and airflow. The final layout must be compatible as one system, not merely as a component list.
03 Cooling
Treat heat removal as part of the product.
High-airflow fan modules serve a tightly controlled front-to-back path. Loaded wall power, room conditions, noise and heat rejection require a measured site and reference-build record.
04 Accelerator plane
Count independent workers, not imaginary pooled memory.
A dense accelerator layout can run separate jobs or supported parallel workloads. Per-GPU memory, runtime behaviour, context and concurrency decide what the system can actually serve.
05 Operations
Make management visible before handover.
Remote health, access, sensors, logs, credentials, recovery and update ownership are part of the appliance. The final interface and permissions are recorded for the ordered platform.
Showing system view 1 of 5: Platform.
Workload route
Every use needs its own acceptance test.
Model name and aggregate VRAM do not prove a business outcome. Each proposed route below states how the capacity could be used and what must be tested before the order treats it as suitable.
01
Concurrent private assistants
- Proposed use
- Assign independent GPUs to separate teams, models or service queues instead of assuming every request shares one memory pool.
- Acceptance evidence
- Measure the named model, context and simultaneous request pattern against the team's response-time target.
02
RAG plus creative or batch work
- Proposed use
- Separate a document assistant from image, transcription, embedding or render queues so one service does not consume every GPU.
- Acceptance evidence
- Test each queue alone and under the agreed mixed workload, then record wall power and service behaviour.
03
Supported model sharding
- Proposed use
- Evaluate a multi-GPU runtime only where the model and serving engine support the intended execution pattern.
- Acceptance evidence
- Record per-GPU memory, throughput and latency at the disclosed quantisation, context and concurrency.
Specification certainty ledger
Known, proposed and still to be confirmed.
A proposal should expose missing facts. The final order replaces every confirmation row with an exact part, measured result or named customer decision.
| Item | Current value | Status | What the record must show |
|---|---|---|---|
| Form factor | Rack server | Package proposal | The site, delivery and support route is designed around this physical class. |
| Physical platform | OEM 4U dual-socket multi-GPU server platform | Package proposal | Platform family selected for the proposed configuration. |
| GPU route | 4 × proposed passive RTX 5090-class 32GB GPUs | Package proposal | Exact manufacturer, part number, condition and serials belong in the final bill of materials. |
| Per-GPU memory | 32GB | Package proposal | The safer model-fit starting point before any supported multi-GPU test. |
| Aggregate GPU memory | 128GB across 4 GPUs | Package proposal | Not presented as one universal memory pool. |
| System memory | 256GB | Package proposal | Memory population, speed and expansion route require the final platform bill. |
| Primary storage | 4TB NVMe | Package proposal | Drive model, endurance, layout and backup destination are confirmed in the order. |
| Network | 10GbE | Package proposal | Customer switching, cabling, storage traffic and segmentation remain part of site design. |
| Power planning | Plan around 3.0kW under load; site pre-flight required | Customer decision | A qualified site review and measured reference build must replace the planning figure. |
| CPU and motherboard | To be stated in the final bill of materials | Supplier confirmation | No unverified processor, lane or motherboard claim is made on this proposal page. |
| Power supplies and leads | To be confirmed for the ordered build and UK site | Supplier confirmation | Include PSU count, rating, redundancy position, input requirements and lead specification. |
| Remote management | Platform route to be confirmed and access policy agreed | Supplier confirmation | The supplier interface shown in the gallery is a reference, not a promise of the final feature set. |
| Dimensions and weight | Exact ordered-system values required | Supplier confirmation | Placement, handling and the delivery route depend on the exact ordered-system values. |
| Hardware warranty route | Exact supplier and component warranty route required | Supplier confirmation | The final order records the whole-system, component, onsite, return and dead-on-arrival boundaries. |
32GB
Per GPU is the first sizing boundary.
Business 128 has 4 GPUs and 128GB aggregate VRAM. A runtime may split a supported model or distribute independent jobs, but the page does not claim one physical 128GB memory pool. Quantisation, context, KV cache, batching and concurrency still change fit.
Software and security
The usable product is more than the chassis.
The final stack stays deliberately small. Versions, licences, access and recurring ownership are recorded so the customer is not left with an opaque collection of containers.
Proposed software baseline
- Ubuntu LTS on a recorded operating-system version
- NVIDIA driver, CUDA components and container support validated for the ordered hardware
- Docker Engine and NVIDIA Container Toolkit
- One primary model server selected from Ollama, vLLM or llama.cpp for the accepted workload
- Open WebUI or another reviewed browser interface
- Named authentication, TLS and reverse-proxy approach
- GPU, node and service monitoring with an agreed log-retention period
- Pinned versions, a software bill of materials and a model source and licence record
Security ownership
- GPU RIGS baseline
- Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
- Customer or contracted operator
- User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
- Shared before acceptance
- Model and software licence checks, retention and logging choices, recovery test, acceptance criteria and a named owner for every recurring task.
Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.
Testing and acceptance
The evidence pack is part of the machine.
No public benchmark is invented for this page. The accepted workload, exact build and disclosed test conditions decide what can be claimed after the reference system exists.
- 01
Record the final bill of materials, serial numbers and firmware versions.
- 02
Run at least 24 hours of GPU, CPU, memory and storage stress testing.
- 03
Capture temperature, fan, error, health and wall-power evidence under the agreed load.
- 04
Test cold boot, restart and the available remote-management route.
- 05
Check drive health, network throughput and the container and GPU runtime.
- 06
Run model and workload smoke tests against the written acceptance set.
- 07
Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.
- 08
Test an agreed fault or recovery route and provide the resulting handover record.
What is not claimed today
No fixed users, tokens per second, latency, model size, accuracy, availability, savings or marketplace contribution is stated without the missing configuration and test conditions.
Read the evidence methodSite, power and cooling
The room is part of the specification.
Planning power is not a measured promise. The final pre-flight records the circuit, voltage, loaded wall power, heat route, noise tolerance, rack, network, UPS decision and operating owner.
Read the site guide3.0kW
Current loaded-system planning basis
- A server room or colocation position suitable for a 4U high-density GPU server
- Electrical planning around a 3.0kW loaded system, subject to exact measured evidence
- Cooling and ventilation sized for the room and duty cycle
- An acoustic plan that does not assume normal office noise levels
- 10GbE network planning, remote-management controls, rack and UPS decisions
Delivery, support and warranty
A written boundary before money moves.
The proposed route covers discovery, configuration, evidence and remote handover. It does not quietly absorb facilities work, migration or an always-on managed service.
Included in the proposed baseline
- Documented workload and site-fit review
- Confirmed bill of materials before procurement
- Configuration, burn-in and agreed smoke-test evidence
- Asset schedule, admin notes and user quick-start material
- Collection or the quoted kerbside or pallet-delivery route
- Remote onboarding and 30-day configuration-defect support
Separate scope or customer responsibility
- Building electrical work, rack, UPS, cooling or structured cabling
- Nationwide on-site installation unless separately quoted
- Migration of customer data, every integration or every application
- Continuous managed operations, security monitoring or a 24-hour support agreement
- Third-party model, API, marketplace or software charges
- A compliance certificate, performance guarantee or income guarantee
Supplier and warranty gates
Unresolved items stay visible.
- Written confirmation that the proposed passive RTX 5090-class cards, risers, firmware, power leads and cooling are supported together
- Exact GPU manufacturer, condition and warranty responsibility recorded
- Whole-system, component, dead-on-arrival and return responsibilities recorded before order
- Measured reference-build thermals and power captured before a customer acceptance claim
- Current delivered price, lead time and UK delivery terms confirmed in the quotation
Evidence and change record
Proposal inputs remain dated.
- Package model
- UK proposal reviewed 25 July 2026
- Price state
- Calculated from dated public supplier inputs, not a live quotation
- Performance state
- No public benchmark until the exact reference build is tested
- OEM platform record
- Exact platform identity and bill of materials provided in the written quotation.
Commercial reality
Tax and spare capacity are supporting questions.
Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.
Finance, VAT and capital allowances
- The displayed price is an indicative proposal based on dated public supplier inputs. The final bill of materials and supplier quote control the order.
- Prices exclude VAT. VAT recovery depends on the buyer, its taxable activities and the normal evidence rules.
- Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
- A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
- There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Optional idle capacity
- Marketplace mode is off by default and is excluded from the purchase case.
- A separate environment, no customer data mounts, network controls and a local kill switch would be required.
- The customer, insurer, supplier warranty and marketplace terms must permit the proposed use.
- Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
- Any pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time before continuation.
Limits and alternatives
A good specification leaves room for “no”.
The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to making this package carry a workload it has not proved.
Package boundaries
- 128GB is aggregate VRAM across four GPUs.
- Passive GPU, firmware, thermal and power compatibility require written supplier confirmation.
- Actual users and throughput depend on model, quantisation, context and concurrency.
- Each GPU has 32GB of local VRAM. The 128GB figure is aggregate capacity.
- The page does not claim that any 128GB model behaves as if it has one physical 128GB GPU.
- No user-count or throughput promise is made without model, context and concurrency evidence.
- A roughly 3kW loaded platform is unlikely to suit a normal office without a designed equipment environment.
A lower-cost proof route when two independent 24GB workers are enough and the refurbished-GPU warranty is acceptable.
Consider this route Sovereign 192Consider when one workload needs 96GB on a single GPU or professional server-GPU positioning matters.
Consider this route Private-first hybridKeep accepted routine work local while retaining an approved hosted route for frontier or exceptional tasks.
Consider this routeQuestions answered
Business 128 questions that affect the order
These answers preserve the validation boundary. The written quotation and acceptance plan replace proposal assumptions.
Is 128GB available to one model?
The server has four 32GB GPUs. Some models and runtimes can use more than one GPU, but the execution pattern, memory use and speed must be tested. The page does not present 128GB as one universal memory pool.
How many people can use Business 128?
There is no fixed user number. Active requests, model size, context, response length, batching and latency target determine capacity. The acceptance test uses the buyer's likely simultaneous demand.
Will it work in an office?
The planning case is around 3.0kW under load and a high-airflow 4U chassis. Electrical, cooling and acoustic suitability must be confirmed before an office-adjacent installation is considered.
Is the RTX 5090 AI configuration confirmed?
Not yet. Passive GPU, riser, firmware, thermal and power compatibility require written supplier confirmation and a tested reference build.
Prepare the next decision
Specify the work before the parts.
Record the workload, data boundary, users, site and acceptance test. No confidential documents or credentials are needed for the first brief.
Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.