Route comparison
Choose the operating shape before the chassis
Start from constraints, not a product badge. The right route can change when privacy, demand or site readiness changes.
- Variable external capacity Useful for light or changing demand where the data route is acceptable
- Local team system Lower-density option where office fit and direct use matter
- Shared managed capacity Appropriate when density, remote operation and a suitable site are justified
- Explicit split Local route by default with controlled exceptions for agreed tasks
Compare each route using the same workload, data boundary, service level and complete cost period.
Buying a private AI server is an infrastructure decision, not an online shopping exercise. The visible parts - GPU name, VRAM and purchase price - are only the beginning. A useful system must run an accepted workload, fit the customer’s site, have a supportable software route and leave somebody clearly responsible after handover.
This guide sets out the questions a UK organisation should answer before asking for a quotation.
Step 1: Write down the job
Start with a narrow statement: who will use the system, what will they ask it to do, which information may it process and what result would be good enough?
“We need private AI” is not a workload. These are:
- twenty solicitors searching an approved precedent library, with matter permissions and source citations;
- six developers asking repository-aware coding questions during the working day;
- a render queue processing independent frames overnight;
- one internal application sending predictable requests to an OpenAI-compatible endpoint.
For language models, record the likely model family, quantisation, context length, response length, simultaneous demand and acceptable latency. For retrieval, record corpus size, update frequency, file types and permission model. For batch work, record job size, queue pattern and completion window.
Step 2: Let the quality test reject the idea
Do not assume a local model is equivalent to the hosted model people use today. Build a representative, permission-safe evaluation set. Score factual support, citations, refusals, latency and reviewer effort.
The local route only passes when the outcome is good enough for the agreed use. Some organisations will keep a hosted frontier model for difficult or low-risk exceptions. That is a valid hybrid architecture, not a failed private deployment.
Step 3: Understand per-GPU memory
VRAM is the first constraint, but the shape matters.
Two 24GB GPUs contain 48GB in aggregate. They do not automatically behave like one 48GB GPU. A model must fit on a single card, use a runtime that can divide it effectively, or be served as independent workers. Context, KV cache, batch size and concurrent requests consume additional memory.
Ask a supplier to state:
- the exact GPU count and VRAM per GPU;
- the model, quantisation, runtime and context used in any benchmark;
- whether the workload is single-GPU, sharded or independent-worker;
- throughput and latency at a stated concurrency;
- the maximum tested temperature and measured wall power.
“Runs large models” or “supports 100 users” is not enough.
Step 4: Choose the physical form
A tower workstation can be the right answer for one or two GPUs, a small technical team and an office setting. It is usually easier to house and quieter than a dense rack server.
A 4U rack server earns its place when the organisation needs density, several independent workers, remote management, redundant power or a proper technical facility. It also brings meaningful airflow, heat and acoustic requirements.
Before purchase, confirm:
- rack space and depth;
- circuit capacity and connector type;
- room cooling and ventilation;
- acceptable noise at the intended location;
- network route, VLAN, DNS, NTP and remote administration;
- UPS and controlled shutdown policy;
- delivery access, weight and installation responsibility.
If the building cannot support the machine, consider a workstation, colocation or cloud service.
Step 5: Inspect the software and licence chain
A business-ready server needs a documented operating system, driver, container runtime, model server, user interface, monitoring profile and update route.
Open-source-first can reduce lock-in and make the stack inspectable. It does not remove maintenance. Open model weights can also carry licence conditions that vary by version and use.
The handover should identify:
- pinned software versions;
- model source, version and licence;
- administrative credentials and recovery materials;
- start, stop, backup and restore procedures;
- logging and monitoring locations;
- the update owner and rollback route;
- components that are included, excluded or community-supported.
Step 6: Treat privacy as shared responsibility
Keeping a workload on customer-controlled infrastructure can reduce external disclosure. It does not by itself provide security or UK GDPR compliance.
The organisation still needs a lawful basis, transparency, data minimisation, retention rules, rights handling, access reviews, incident response, backups and accuracy controls. The ICO’s AI and data-protection guidance is a useful starting point.
Ask for a responsibility matrix. The supplier may configure an initial baseline, but the customer normally owns identity, user approvals, network operation, backup copies, business continuity, ongoing patches and use policy after acceptance.
Step 7: Demand dated supplier evidence
AI hardware moves quickly. A quotation should identify the exact bill of materials, whether any component is new or refurbished, warranty length, warranty provider, return route, lead time, substitution policy and quotation validity.
For a configured appliance, acceptance evidence should include:
- asset and serial schedule;
- firmware and software versions;
- GPU health and error evidence;
- memory, storage and network checks;
- at least 24 hours of representative burn-in;
- model/runtime smoke tests under disclosed conditions;
- power and temperature observations;
- an exceptions log.
Do not let a generic platform brochure substitute for the exact machine.
Step 8: Compare total cost, not a headline
Owned capacity includes hardware, finance, electricity, cooling, support, rack or colocation, insurance, internal operation and downtime. Hosted capacity includes seats, tokens, implementation, network dependency and future price uncertainty.
Set residual value to zero in the base case unless there is reliable evidence. Exclude speculative marketplace income. Run downside, central and upside cases and show the assumptions beside the result.
For a light-use team, hosted subscriptions can remain much cheaper. Ownership is more credible where sustained use, data route, offline capability or predictable shared capacity has value beyond pure cash break-even.
Step 9: Keep tax language guarded
Qualifying equipment may be eligible for Annual Investment Allowance or full expensing, depending on the purchaser and current rules. A VAT-registered organisation may be able to recover input VAT subject to normal conditions.
That is not a generic capital-gains advantage, an automatic R&D claim or personal tax advice. Ask the organisation’s accountant to confirm treatment before relying on it.
Step 10: Use a written buying gate
Before paying a deposit, require:
- an accepted workload and quality test;
- an exact bill of materials and time-limited supplier quote;
- a written warranty and support route;
- a completed site pre-flight;
- benchmark evidence on the exact reference build;
- a total-cost model using the customer’s assumptions;
- a security and operations responsibility matrix;
- order terms covering substitutions, acceptance and returns.
The strongest buying decision is sometimes to proceed with a smaller system, use cloud for longer or wait. A credible supplier should be able to say that plainly.
Technical context
See the physical and operating boundary
Use these views to connect the guide to the machine, its airflow and its operating environment. Captions state the limits of what each image shows.
Deposit gate
Eight records before money moves
The buying decision becomes traceable when each risk has an owner and dated evidence.
- Workload and site Accepted quality test, benchmark conditions and signed pre-flight
- Exact machine Bill of materials, component status and substitution rules
- Support route Warranty provider, response path, updates and responsibilities
- Commercial record Full cost, validity, acceptance, returns and delivery
If the evidence is not specific to the ordered machine, keep the gate open.
Primary sources
- ICO: guidance on AI and data protection
- NCSC: secure system administration
- NVIDIA: vGPU and GPU documentation
Sources are checked at the review date. Platform terms, prices and public guidance can change; verify them at the point of decision.