First-pass technical fit
Private AI server sizing tool.
Choose a memory band, active demand and location. The result is an evaluation route - not a benchmark or promise that a model will fit.
Use the exact model, quantisation, context and runtime to verify this.
Use active requests, not named staff.
Next evaluation route
-
- Still requiredExact workload test
- Still requiredPower and site pre-flight
- Still requiredSupplier and warranty check
Calculation record
See what sits behind the result.
The controls are deliberately editable. Record the values used, the date and the source before relying on an output in a buying decision.
Method
- The first branch is the likely minimum memory required on one GPU.
- Active generations alter queue and cache demand; named users do not.
- Location changes the viable acoustic, power, cooling and service boundary.
Worked reading
- A 24GB target with modest concurrent demand may justify a compact evaluation route.
- A model needing more than 48GB on one GPU points to a 96GB-class test, not two unrelated 24GB cards.
- Any office result still needs a noise and facilities check.
Do not infer
- The tool does not prove that a model, context or runtime will fit.
- Aggregate VRAM is not automatically one shared memory pool.
- Only a representative benchmark can establish useful latency, throughput and quality.
Why the tool stops at an evaluation route
Model weights, context, KV cache, batching, runtime overhead and inter-GPU behaviour change the result. User count alone cannot produce a safe server recommendation. The route tells you what to test next.
Management access belongs in the specification.
A useful handover records software versions, credentials, recovery, monitoring and update ownership alongside the physical server.