Local control, practical limits

Run useful AI locally - without pretending local is always better.

Local AI places the model and approved data on customer-controlled hardware. It can improve control, predictability and offline capability, but the customer also owns operations, energy and lifecycle.

Decision in one minute

Choose local when the workload and control case are strong. Use hybrid when a local default and an approved hosted exception give the best balance.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. Original explanatory plate. It sets out a decision method, not a measured result.
A private deployment starts with the permitted data path, access policy and logging boundary.

Buyer field note

Let the work and the room reject the wrong machine.

01 / Fit
Start with the accepted task, not a chassis or GPU headline.
02 / Site
Check power, cooling, noise, network and maintenance access.
03 / Alternative
Keep cloud, colocation or a smaller system in the decision.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. Original explanatory plate. It sets out a decision method, not a measured result.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. Original explanatory plate. It sets out a decision method, not a measured result.

What stays local

The agreed model, prompts, retrieval index, documents, logs and output can remain within the customer’s defined environment.

External connectors, updates, telemetry and fallback models must be individually identified; “local” should not be a vague blanket claim.

Capacity budget

Memory demand has more than one layer

Model
Weights and quantisation.
Context
Working input and cache.
Concurrency
Simultaneous active work.
Headroom
Operations and change allowance.
The bands are a checklist, not a scale. Exact memory belongs to a named model and test. Decision aid for: What stays local

Connect the category to the physical and operating system.

Platform views and technical plates show what must be tested beyond a product label or aggregate specification.

GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image. Written reuse permission pending.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image. Written reuse permission pending.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A credible handover records the supplied assets, checks, operating evidence, instructions and unresolved items. Original explanatory plate. It sets out a decision method, not a measured result.
A credible handover records the supplied assets, checks, operating evidence, instructions and unresolved items. Original explanatory plate. It sets out a decision method, not a measured result.

What the business takes on

Owned hardware needs power, cooling, monitoring, updates, backup, support and a replacement path. Someone must own identity and daily availability.

The fit check includes those operational costs before comparing with a hosted service.

Site boundary

The room is part of the specification

  1. Electrical Circuit, protection and UPS policy.
  2. Thermal Airflow, room heat and cooling route.
  3. Network Access, bandwidth and isolation.
  4. Service Rack space, noise and maintenance access.
A technically suitable server is still the wrong purchase when the operating site cannot support it. Decision aid for: What the business takes on

Open models can be excellent within a clear brief

Private document search, extraction, summarisation, coding and structured workflows are strong candidates. Some complex reasoning and multimodal tasks may still favour hosted frontier models.

Acceptance uses the customer’s representative examples rather than generic leaderboard claims.

Procurement record

Make the route to acceptance inspectable

  1. Fit brief Workload and alternatives recorded.
  2. Reference test Conditions and results retained.
  3. Exact build Bill of materials and substitutions visible.
  4. Handover Owners, versions and exclusions signed off.
The evidence trail should explain why this platform was chosen and what remains outside the order. Decision aid for: Open models can be excellent within a clear brief

No GPU RIGS token meter

GPU RIGS does not propose a per-token charge for work run on customer-owned hardware. Electricity, support, model licences and any chosen hosted fallback remain real costs.

That distinction belongs in every TCO calculation.

Questions answered

Straight answers to common questions

What is a local AI server?

A physical server in a customer-controlled location that runs AI models for local users or applications rather than relying on a hosted AI endpoint for every request.

Is local AI private by default?

It can reduce external data sharing, but privacy depends on network, access, telemetry, connectors, logs and operating controls.

Will a private AI server always be cheaper than cloud AI?

No. SaaS is usually the better-value choice for a small team with light or irregular use. Owned capacity becomes more credible when privacy, offline operation, many shared users or sustained workloads have independent value. We show the comparison rather than forcing the server answer.

Put the claim to work

Turn this guidance into a testable requirement.

The brief asks about workload and operating conditions - not just budget.