Compare dedicated and shared GPU infrastructure for AI training and inference, including allocation, isolation, topology, consistency and commercial fit.
· Arvica Cloud · Analysis & buying guide
What “shared” can mean
Shared GPU is not one architecture. It may mean jobs scheduled at different times on a pooled fleet, time-sliced access to one physical GPU, virtual machines with assigned accelerators, or hardware partitioning such as NVIDIA Multi-Instance GPU on supported devices. NVIDIA's GPU Operator documentation notes that time-sliced replicas share compute time and do not provide the memory or fault isolation of MIG. Ask the provider to name the allocation mechanism rather than accepting “shared” as a complete specification.
What dedicated infrastructure means
Dedicated should identify a boundary: a full physical GPU, an entire host, a complete eight-GPU node, or several nodes and their cluster fabric. A dedicated GPU inside a shared host is different from a single-tenant server with dedicated local disks and network paths. Record each layer—GPU, host CPU and memory, local storage, shared storage, network and management plane—in the proposal and contract.
Performance consistency is about contention and topology
A shared service can be well engineered, but buyers need to know which resources can contend and how placement changes between runs. Dedicated infrastructure can make GPU topology, NUMA placement, NVMe layout, driver versions and NIC attachment easier to keep stable. That stability matters for long experiments and acceptance testing. It does not establish a particular benchmark result; the workload still needs a repeatable test on the offered configuration.
Tenancy and isolation need explicit evidence
NVIDIA documents that MIG can partition supported GPUs into separate instances with isolated memory-system paths and fault domains, while time slicing trades away those properties for broader sharing. Virtualisation, containers and scheduling add further boundaries. None of these labels alone proves that an arrangement meets a customer's security requirement. Review identities, privileged access, storage lifecycle, network segmentation, logging and deletion alongside the hardware allocation.
Long-running training workloads
Dedicated capacity is attractive when a training run spans days, placement and topology must remain stable, checkpoints are large, or job interruption has a high recovery cost. A full eight-GPU node also gives one defined NVLink/NVSwitch domain on the relevant HGX or DGX platform. Before reserving it, verify that the workload can use the topology efficiently; the existing Arvica eight-GPU scaling guide covers that measurement rather than duplicating it here.
Inference economics depend on utilisation
Shared or on-demand capacity often suits variable traffic, development, batch jobs and pilots because the buyer does not commit to idle hardware. Dedicated capacity can suit sustained high utilisation, predictable latency objectives or a fixed production topology. Compare cost per completed request or useful token under the same quality and latency boundary—not only a GPU-hour. Arvica's reserved-versus-on-demand analysis provides the utilisation framework, so this article does not repeat invented rates.
When a dedicated eight-GPU node makes sense
A dedicated node is a credible candidate when the application needs most or all eight GPUs, the intra-node fabric matters, stable access is worth a term commitment, and storage and networking can be specified with the node. It can also be the building block for a future cluster if the first deployment reserves a compatible fabric and operations path. Availability, exact platform, tenancy and timing still require written confirmation.
When shared or on-demand is sufficient
Use shared or on-demand infrastructure when demand is intermittent, jobs tolerate placement changes, the workload fits a smaller allocation, or a pilot is still resolving model and memory requirements. It is often the cleaner way to avoid premature reservation. Ask how resources are labelled, whether capacity can disappear between jobs, what persists after shutdown and how storage and transfer are billed.
Questions before signing a reserved contract
Define the physical and logical tenancy boundary, GPU and host specification, topology, location, storage persistence, network, software image, maintenance process, support access, start window, acceptance tests, renewal and exit process. Separate requirements from preferences. A proposal should confirm what is deployable and when; a public product family name is not a reservation, service level or inventory statement.
Frequently asked questions
Is dedicated GPU always single-tenant?
Not necessarily. The contract should identify whether the GPU, host, storage and network are dedicated and who can administer them.
Can shared GPU still provide isolation?
Yes, through different mechanisms, but the properties differ. For example, NVIDIA states MIG provides memory and fault isolation that simple time slicing does not.
Download an editable worksheet to agree scope, compare configurations and decide whether to stop, adjust or scale. An evaluation method, not a customer case study.
Compare scope, allocation, storage, networking and pilot acceptance before committing to a B300 node. Includes a worked cost formula, not a supplier price.
Build a representative inference test with model, precision, context, concurrency and quality targets. Connect the measured result to a transparent GPU budget.