Enterprise AI infrastructure

Planning a 32-GPU AI Cluster in Australia

Plan a 32-GPU AI cluster in Australia using a four-node model, with workload, fabric, storage, power, residency, tenancy and acceptance requirements.

· Arvica Cloud · Analysis & buying guide

Start with the workload, not the GPU count

Thirty-two GPUs is a topology decision only after the workload is defined. Record model architecture and size, training or inference objective, precision, context and batch distribution, dataset volume, checkpoint cadence, completion target, expected utilisation and software stack. Estimate where memory, compute, communication and storage time are spent. A workload that fits one node may not benefit from four; another may need 32 GPUs primarily to meet a delivery window.

A four-node planning model

A useful planning model is four complete eight-GPU nodes. Each node has its own internal scale-up domain; a separate scale-out fabric connects the four nodes. This is an architecture model, not an inventory statement or a promise that every GPU family is available in that shape. Confirm the exact node platform, GPU SKU, host specification, local storage, NIC layout and deployment site in the proposal.

Choose the GPU against memory and software evidence

H200, B200 and B300 can each be relevant. Compare memory fit, supported precision, framework maturity, measured single-node throughput and complete deployable configuration. B300's larger memory contexts can help memory-bound training or reasoning inference; B200 may be sufficient when the model fits and its capacity window or commercial profile is better; H200 may reduce migration risk for a validated Hopper stack. Do not convert family-level marketing into a universal ranking.

Fabric and topology

The cluster needs a documented inter-node fabric: InfiniBand or an engineered Ethernet/RoCE design, NICs and rails per node, switch topology, oversubscription, routing, congestion handling and failure domains. Map GPUs to NICs and CPU roots. Validate with controlled NCCL collectives and the real workload. A nominal 400 or 800 Gb/s port does not establish effective bandwidth, and one-node NVLink results do not predict four-node scaling.

Storage and checkpoint traffic

Size storage from dataset ingestion, local cache, model artefacts, checkpoint size and frequency, restart objectives and export path. Fast local NVMe can reduce repeated reads, while shared storage gives all nodes a common dataset and checkpoint view. Specify throughput and concurrency as well as terabytes. Keep storage traffic from unexpectedly competing with collective communication, and test sustained reads and checkpoint writes under realistic load.

Power, cooling and rack planning

Use the selected complete-system specification, not GPU TDP. As one clearly labelled reference, NVIDIA lists DGX B300 at 14.5 kW and about 49,476 BTU/hour maximum heat output. Four such DGX reference systems would represent 58 kW of IT load before switches, storage and other equipment. That arithmetic must not be applied to an HGX OEM server: its power, cooling method, rack units, weight and supply design are configuration-dependent.

Data location and residency requirements

Specify compute, persistent storage, backup, logs and support-access locations. For personal information, review the organisation's Australian Privacy Principles obligations and cross-border handling with qualified advisers. A preferred Australian city is not the same as a contractual residency requirement. The proposal should name the deployable site and any agreed access or replication restrictions before commitment.

Single tenancy and the access model

Define whether all four hosts are single-tenant, whether storage and network are shared, who holds administrative privileges, how SSH and service accounts are managed, and which party patches drivers and firmware. Include monitoring, incident escalation and maintenance responsibilities. Dedicated hardware can simplify the boundary, but it does not answer every control-plane, support or data-lifecycle question.

Reserved term and deployment window

Separate target start date, latest acceptable operational date, reservation term and expansion option. Ask what constitutes ready for acceptance, how partial delivery is handled, and whether the fabric and storage arrive with the compute. Newly committed systems may have a different timeline from existing capacity. Availability, price, location, configuration and timing belong in the written proposal; this planning article does not establish any of them.

Acceptance before production use

Agree health, topology, network, storage, thermal and workload checks before the term begins or during a defined acceptance window. Preserve raw outputs and software versions. Arvica's existing GPU rental acceptance checklist covers a fuller process, so the RFQ should reference it rather than recreating an improvised checklist. Acceptance thresholds must be agreed for the actual system and workload.

What belongs in the RFQ

Include workload and software, preferred GPU and acceptable alternatives, exact GPU and node count, location and residency, fabric, storage, tenancy, access, term, start window, support expectations, acceptance tests and the likely path to 64 GPUs. Mark unknowns explicitly. Arvica can use that brief to prepare a technical capacity proposal for review; submission does not reserve infrastructure or create a payment obligation.

Frequently asked questions

How many servers are used in the planning model?

Four complete eight-GPU nodes. This is a planning model, not a statement of current inventory or a required topology for every workload.

Does a 32-GPU cluster require InfiniBand?

Not universally. An engineered Ethernet/RoCE fabric can also be appropriate. Validate the complete topology and workload rather than selecting by protocol name alone.

Related capacity and planning pages

Sources

Reviewed: 2026-10-04

  1. NVIDIA HGX AI Factory — components ↗
  2. NVIDIA NCCL documentation ↗
  3. NVIDIA DGX B300 System User Guide ↗
  4. OAIC — APP 8 cross-border disclosure ↗

Apply this to your workload

No charge before quote acceptance; the infrastructure provider is disclosed before activation.

Buying guides and pilot worksheet