Workload notebook

Start with a bounded workload, not a server model.

The same GPU count can serve three very different jobs. Define the input, demand pattern and acceptance test first, then decide whether owned capacity is proportionate.

Before hardware

Five questions that change the design

If these answers are missing, a larger specification only makes the uncertainty more expensive.

  1. What data enters?

    Name the document, repository, media or batch source. Record its owner, sensitivity, size and update pattern.

  2. What must the output pass?

    Build a representative acceptance set before discussing a package. A plausible demo is not a pass condition.

  3. How is demand shaped?

    Concurrent users, context, queue length, latency and working hours drive different capacity decisions.

  4. Who operates it?

    Name the owner for access, updates, logs, backup, incidents and the hosted fallback.

  5. Where will it run?

    Check rack space, circuits, heat, noise, network and physical access before treating the design as viable.

Workload routes

Choose by the work your team needs to finish

These are deliberately bounded patterns. A new page is added only when there is enough distinct workflow, sizing and test evidence to make it useful.

Workload 01

private RAG server

A private document assistant that shows its sources and its limits.

Build a private company document assistant with source citations, permission-aware retrieval, evaluation and a customer-controlled server.

Bring to the sizing session
Approved documents, permission groups, question set and update frequency
Acceptance evidence
Retrieval, citation support, refusal and permission boundaries
Hosted route
Choose hosted when connector maturity or managed model quality matters more than a local route.
Examine this workload
Workload 02

local AI coding server

Local code assistance measured against your repositories.

Private AI coding servers for repository-aware assistance, controlled source-code access, model evaluation and multi-developer use.

Bring to the sizing session
Repository scope, developer concurrency, context length and secret boundaries
Acceptance evidence
Correctness, security, review effort and useful response time
Hosted route
Keep hosted coding tools where their capability and commercial terms fit the repositories.
Examine this workload
Workload 03

GPU rendering server

Dense GPU workers for jobs that do not need one shared memory pool.

Physical GPU servers for rendering, image, transcription, embeddings and independent batch workers - with power and marketplace cautions.

Bring to the sizing session
Job queue, software licences, completion target, wall power and working hours
Acceptance evidence
Completed useful work, queue time, stability and measured energy
Hosted route
Use rented capacity where work is irregular or deadlines need rapid scale.
Examine this workload

Evidence sequence

A package follows the acceptance set

This order prevents a polished hardware specification from becoming a substitute for workload proof.

  1. 01

    Collect representative inputs

    Use authorised examples that reflect difficult and ordinary work.

  2. 02

    Agree pass and refusal conditions

    Name the quality, latency, security and human-review requirements.

  3. 03

    Test a reference configuration

    Record model, runtime, context, concurrency, output and wall power.

  4. 04

    Confirm the operating site

    Validate facilities, support ownership, network and fallback route.

Workload decisions become physical

GPU layout, cooling, remote management and the installation setting all affect whether a reference workload can become a dependable business service.

Three-quarter supplier render of a 4U OEM multi-GPU rack server
Three-quarter supplier render of a 4U OEM multi-GPU rack server
OEM platform reference render. It is not evidence of a completed customer build or final specification. OEM supplier reference image. Written reuse permission pending.
OEM platform reference render. It is not evidence of a completed customer build or final specification. OEM supplier reference image. Written reuse permission pending.
Open 4U OEM GPU server chassis showing passive GPUs, cooling fans, processors and memory slots
Open 4U OEM GPU server chassis showing passive GPUs, cooling fans, processors and memory slots
OEM supplier render showing one possible internal layout. Components vary with the ordered build. OEM supplier reference image. Written reuse permission pending.
OEM supplier render showing one possible internal layout. Components vary with the ordered build. OEM supplier reference image. Written reuse permission pending.
Diagram combining model weights, context, cache and active requests into a memory headroom check
Diagram combining model weights, context, cache and active requests into a memory headroom check
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. Original explanatory plate. It sets out a decision method, not a measured result.
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. Original explanatory plate. It sets out a decision method, not a measured result.
Exploded supplier render of passive GPUs arranged above an open 4U rack chassis
Exploded supplier render of passive GPUs arranged above an open 4U rack chassis
Supplier layout render used to explain GPU density and airflow. It does not represent a confirmed package configuration. OEM supplier reference image. Written reuse permission pending.
Supplier layout render used to explain GPU density and airflow. It does not represent a confirmed package configuration. OEM supplier reference image. Written reuse permission pending.

Smallest useful next step

Bring one workload, one owner and one pass condition

Do not upload confidential material through the public site. A requirements brief can describe the workload shape first, with a controlled evidence route agreed separately.