Multi-workload company capacity

Business 128

Four GPU workers for teams that need throughput, separation and room to grow.

A dense multi-GPU platform for concurrent private assistants, mixed AI services, render queues and supported model sharding.

Validation-stage standard configuration.

This public proposal remains indexable while clearly qualifying its evidence state. Price, exact parts, compatibility, warranty, delivery and workload results require a current quotation and written acceptance record.

Price status
Indicative, ex VAT
Commercial model
Reviewed 25 July 2026
Range position
Standard progression
Supplier platform / proposed build
Rear view of a 4U OEM multi-GPU server chassis
Inspect the platform
128GB aggregate VRAM* OEM 4U dual-socket multi-GPU server platform proposed physical platform GPU route 4 × proposed passive RTX 5090-class 32GB GPUs

Buyer fit

Start with the reason to own it.

A multi-worker company platform for organisations that need several private services or queues at the same time and have a suitable technical facility.

Indicative purchase price

£70,500

ex VAT · £84,600 inc VAT

Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.

A credible fit

  • Several independent model endpoints
  • Company RAG plus creative or batch workloads
  • Teams with a proper server room or colocation plan

Choose another route when

  • A quiet under-desk appliance
  • A single 128GB workload without runtime validation
  • A site without electrical and cooling capacity

Interactive system inspection

Inspect the platform behind Business 128.

Move from the complete Business 128 platform family into its internal layout, cooling, accelerator plane and management boundary. These views explain the engineering questions; the final ordered build and acceptance record remain decisive.

Three-quarter view of a 4U multi-GPU rack server platform
01 4U / 19-inch rack

01 Platform

Start with the complete machine.

A serviceable 4U rack platform provides the physical boundary. Exact dimensions, rails, weight, power supplies and ordered components still belong in the written quotation.

Inspection focus 4U / 19-inch rack

02 Internal layout

See where density becomes an engineering decision.

GPU positions, processor sockets, memory, storage and cable paths compete for space and airflow. The final layout must be compatible as one system, not merely as a component list.

Inspection focus GPU / CPU / memory

03 Cooling

Treat heat removal as part of the product.

High-airflow fan modules serve a tightly controlled front-to-back path. Loaded wall power, room conditions, noise and heat rejection require a measured site and reference-build record.

Inspection focus Airflow / service access

04 Accelerator plane

Count independent workers, not imaginary pooled memory.

A dense accelerator layout can run separate jobs or supported parallel workloads. Per-GPU memory, runtime behaviour, context and concurrency decide what the system can actually serve.

Inspection focus Per-GPU fit first

05 Operations

Make management visible before handover.

Remote health, access, sensors, logs, credentials, recovery and update ownership are part of the appliance. The final interface and permissions are recorded for the ordered platform.

Inspection focus Observe / recover / own

Showing system view 1 of 5: Platform.

Workload route

Every use needs its own acceptance test.

Model name and aggregate VRAM do not prove a business outcome. Each proposed route below states how the capacity could be used and what must be tested before the order treats it as suitable.

01

Concurrent private assistants

Proposed use
Assign independent GPUs to separate teams, models or service queues instead of assuming every request shares one memory pool.
Acceptance evidence
Measure the named model, context and simultaneous request pattern against the team's response-time target.

02

RAG plus creative or batch work

Proposed use
Separate a document assistant from image, transcription, embedding or render queues so one service does not consume every GPU.
Acceptance evidence
Test each queue alone and under the agreed mixed workload, then record wall power and service behaviour.

03

Supported model sharding

Proposed use
Evaluate a multi-GPU runtime only where the model and serving engine support the intended execution pattern.
Acceptance evidence
Record per-GPU memory, throughput and latency at the disclosed quantisation, context and concurrency.

Specification certainty ledger

Known, proposed and still to be confirmed.

A proposal should expose missing facts. The final order replaces every confirmation row with an exact part, measured result or named customer decision.

Package proposal Supplier confirmation Customer decision
Proposed Business 128 specification and evidence status
Item Current value Status What the record must show
Form factor Rack server Package proposal The site, delivery and support route is designed around this physical class.
Physical platform OEM 4U dual-socket multi-GPU server platform Package proposal Platform family selected for the proposed configuration.
GPU route 4 × proposed passive RTX 5090-class 32GB GPUs Package proposal Exact manufacturer, part number, condition and serials belong in the final bill of materials.
Per-GPU memory 32GB Package proposal The safer model-fit starting point before any supported multi-GPU test.
Aggregate GPU memory 128GB across 4 GPUs Package proposal Not presented as one universal memory pool.
System memory 256GB Package proposal Memory population, speed and expansion route require the final platform bill.
Primary storage 4TB NVMe Package proposal Drive model, endurance, layout and backup destination are confirmed in the order.
Network 10GbE Package proposal Customer switching, cabling, storage traffic and segmentation remain part of site design.
Power planning Plan around 3.0kW under load; site pre-flight required Customer decision A qualified site review and measured reference build must replace the planning figure.
CPU and motherboard To be stated in the final bill of materials Supplier confirmation No unverified processor, lane or motherboard claim is made on this proposal page.
Power supplies and leads To be confirmed for the ordered build and UK site Supplier confirmation Include PSU count, rating, redundancy position, input requirements and lead specification.
Remote management Platform route to be confirmed and access policy agreed Supplier confirmation The supplier interface shown in the gallery is a reference, not a promise of the final feature set.
Dimensions and weight Exact ordered-system values required Supplier confirmation Placement, handling and the delivery route depend on the exact ordered-system values.
Hardware warranty route Exact supplier and component warranty route required Supplier confirmation The final order records the whole-system, component, onsite, return and dead-on-arrival boundaries.

32GB

Per GPU is the first sizing boundary.

Business 128 has 4 GPUs and 128GB aggregate VRAM. A runtime may split a supported model or distribute independent jobs, but the page does not claim one physical 128GB memory pool. Quantisation, context, KV cache, batching and concurrency still change fit.

Software and security

The usable product is more than the chassis.

The final stack stays deliberately small. Versions, licences, access and recurring ownership are recorded so the customer is not left with an opaque collection of containers.

Proposed software baseline

  1. Ubuntu LTS on a recorded operating-system version
  2. NVIDIA driver, CUDA components and container support validated for the ordered hardware
  3. Docker Engine and NVIDIA Container Toolkit
  4. One primary model server selected from Ollama, vLLM or llama.cpp for the accepted workload
  5. Open WebUI or another reviewed browser interface
  6. Named authentication, TLS and reverse-proxy approach
  7. GPU, node and service monitoring with an agreed log-retention period
  8. Pinned versions, a software bill of materials and a model source and licence record

Security ownership

GPU RIGS baseline
Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
Customer or contracted operator
User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
Shared before acceptance
Model and software licence checks, retention and logging choices, recovery test, acceptance criteria and a named owner for every recurring task.

Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.

Testing and acceptance

The evidence pack is part of the machine.

No public benchmark is invented for this page. The accepted workload, exact build and disclosed test conditions decide what can be claimed after the reference system exists.

  1. 01

    Record the final bill of materials, serial numbers and firmware versions.

  2. 02

    Run at least 24 hours of GPU, CPU, memory and storage stress testing.

  3. 03

    Capture temperature, fan, error, health and wall-power evidence under the agreed load.

  4. 04

    Test cold boot, restart and the available remote-management route.

  5. 05

    Check drive health, network throughput and the container and GPU runtime.

  6. 06

    Run model and workload smoke tests against the written acceptance set.

  7. 07

    Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.

  8. 08

    Test an agreed fault or recovery route and provide the resulting handover record.

What is not claimed today

No fixed users, tokens per second, latency, model size, accuracy, availability, savings or marketplace contribution is stated without the missing configuration and test conditions.

Read the evidence method

Site, power and cooling

The room is part of the specification.

Planning power is not a measured promise. The final pre-flight records the circuit, voltage, loaded wall power, heat route, noise tolerance, rack, network, UPS decision and operating owner.

Read the site guide

3.0kW

Current loaded-system planning basis

  • A server room or colocation position suitable for a 4U high-density GPU server
  • Electrical planning around a 3.0kW loaded system, subject to exact measured evidence
  • Cooling and ventilation sized for the room and duty cycle
  • An acoustic plan that does not assume normal office noise levels
  • 10GbE network planning, remote-management controls, rack and UPS decisions

Delivery, support and warranty

A written boundary before money moves.

The proposed route covers discovery, configuration, evidence and remote handover. It does not quietly absorb facilities work, migration or an always-on managed service.

Included in the proposed baseline

  • Documented workload and site-fit review
  • Confirmed bill of materials before procurement
  • Configuration, burn-in and agreed smoke-test evidence
  • Asset schedule, admin notes and user quick-start material
  • Collection or the quoted kerbside or pallet-delivery route
  • Remote onboarding and 30-day configuration-defect support

Separate scope or customer responsibility

  • Building electrical work, rack, UPS, cooling or structured cabling
  • Nationwide on-site installation unless separately quoted
  • Migration of customer data, every integration or every application
  • Continuous managed operations, security monitoring or a 24-hour support agreement
  • Third-party model, API, marketplace or software charges
  • A compliance certificate, performance guarantee or income guarantee

Supplier and warranty gates

Unresolved items stay visible.

  1. Written confirmation that the proposed passive RTX 5090-class cards, risers, firmware, power leads and cooling are supported together
  2. Exact GPU manufacturer, condition and warranty responsibility recorded
  3. Whole-system, component, dead-on-arrival and return responsibilities recorded before order
  4. Measured reference-build thermals and power captured before a customer acceptance claim
  5. Current delivered price, lead time and UK delivery terms confirmed in the quotation

Evidence and change record

Proposal inputs remain dated.

Package model
UK proposal reviewed 25 July 2026
Price state
Calculated from dated public supplier inputs, not a live quotation
Performance state
No public benchmark until the exact reference build is tested
OEM platform record
Exact platform identity and bill of materials provided in the written quotation.

Commercial reality

Tax and spare capacity are supporting questions.

Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.

Finance, VAT and capital allowances

  • The displayed price is an indicative proposal based on dated public supplier inputs. The final bill of materials and supplier quote control the order.
  • Prices exclude VAT. VAT recovery depends on the buyer, its taxable activities and the normal evidence rules.
  • Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
  • A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
  • There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Read the guarded UK buyer notes

Optional idle capacity

  • Marketplace mode is off by default and is excluded from the purchase case.
  • A separate environment, no customer data mounts, network controls and a local kill switch would be required.
  • The customer, insurer, supplier warranty and marketplace terms must permit the proposed use.
  • Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
  • Any pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time before continuation.
Review the marketplace decision gate

Limits and alternatives

A good specification leaves room for “no”.

The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to making this package carry a workload it has not proved.

Package boundaries

  • 128GB is aggregate VRAM across four GPUs.
  • Passive GPU, firmware, thermal and power compatibility require written supplier confirmation.
  • Actual users and throughput depend on model, quantisation, context and concurrency.
  • Each GPU has 32GB of local VRAM. The 128GB figure is aggregate capacity.
  • The page does not claim that any 128GB model behaves as if it has one physical 128GB GPU.
  • No user-count or throughput promise is made without model, context and concurrency evidence.
  • A roughly 3kW loaded platform is unlikely to suit a normal office without a designed equipment environment.

Questions answered

Business 128 questions that affect the order

These answers preserve the validation boundary. The written quotation and acceptance plan replace proposal assumptions.

Is 128GB available to one model?

The server has four 32GB GPUs. Some models and runtimes can use more than one GPU, but the execution pattern, memory use and speed must be tested. The page does not present 128GB as one universal memory pool.

How many people can use Business 128?

There is no fixed user number. Active requests, model size, context, response length, batching and latency target determine capacity. The acceptance test uses the buyer's likely simultaneous demand.

Will it work in an office?

The planning case is around 3.0kW under load and a high-airflow 4U chassis. Electrical, cooling and acoustic suitability must be confirmed before an office-adjacent installation is considered.

Is the RTX 5090 AI configuration confirmed?

Not yet. Passive GPU, riser, firmware, thermal and power compatibility require written supplier confirmation and a tested reference build.

Prepare the next decision

Specify the work before the parts.

Record the workload, data boundary, users, site and acceptance test. No confidential documents or credentials are needed for the first brief.

Proposed indicative price, excluding VAT. Final price, specification, availability, delivery, warranty and model fit require a current supplier quote and written order confirmation.