Plan a B300 deployment using complete-system power, cooling, rack, NVLink, scale-out networking and storage requirements—with DGX, HGX and OEM contexts kept separate.
· Arvica Cloud · Analysis & buying guide
Start with the complete system, not GPU TDP
A B300 server includes GPUs, host CPUs, system memory, NVSwitch, NICs, NVMe, fans and power supplies. Facility planning must therefore use the selected server's rated inputs and thermal specification. GPU TDP cannot be multiplied into a safe rack design, and a family name does not identify the number of power feeds, voltage, airflow, weight or cooling method.
DGX B300 is one labelled reference
NVIDIA's DGX B300 user guide lists 14.5 kW power consumption, 49,476.054 BTU/hour maximum heat output, 10 rack units and 12 AC power inlets for the PSU design. Those figures describe DGX B300. They are useful for order-of-magnitude planning, but they must not be presented as the specification of an HGX B300 partner server or another OEM implementation.
HGX and OEM systems are configuration-dependent
HGX describes an accelerated-computing platform integrated by server partners. CPU selection, host memory, storage, NICs, chassis, power supplies and cooling can vary within certified designs. Obtain the exact OEM model and revision, rack and power drawing, thermal limits, cable schedule, firmware matrix and support process. If a supplier quotes a DGX figure for an HGX server, require a system-specific document before facility approval.
Electrical planning
Confirm input voltage, feed type, connector and PDU requirements, A/B redundancy behaviour, circuit and rack limits, power-capping support, maintenance isolation and headroom for switches and storage. Review what happens after one or more feeds are lost; redundancy does not always mean full performance continues. For multiple nodes, calculate the complete rack and facility load from the selected bill of materials, not from a single server headline.
Cooling and airflow
Translate rated heat output into a site design that covers inlet conditions, airflow direction, containment, recirculation risk, fan-failure behaviour, facility redundancy and monitoring. High-density systems can exceed the assumptions of a conventional server room. Air-cooled and liquid-cooled B300-class systems exist in different platform contexts; neither cooling method should be inferred from the GPU name. Use the complete rack design and local facility engineering sign-off.
Rack density, weight and service access
Plan kW per rack together with rack units, static and rolling weight, floor loading, chassis depth, lift and installation path, cable bend radius, PDU placement and front/rear service clearance. A server can fit by rack units and still fail the power, weight or cooling boundary. Multi-node clusters also need space and power for fabric switches, storage and management equipment.
Inside the node: NVLink and NVSwitch
NVIDIA lists fifth-generation NVLink for HGX B300, with 1.8 TB/s GPU-to-GPU bandwidth and 14.4 TB/s aggregate NVLink switch bandwidth across the eight-GPU platform. This is the scale-up fabric. Confirm the actual eight-GPU topology and software visibility during acceptance. Do not confuse the internal NVLink domain with the Ethernet or InfiniBand ports used for storage, clients and other nodes.
Between nodes: specify the complete fabric
For a single node, external networking carries data, storage, management and user traffic. For 32 GPUs or more, it can also carry synchronized collectives. Specify InfiniBand or RoCE/Ethernet, NIC model and count, link speed, rails, switch topology, oversubscription, routing, congestion control, optics and separation from storage and management. OEM NIC layouts vary, so a B300 label cannot answer these questions.
Storage and data paths
Define boot devices, local dataset cache, checkpoint space, model registry, shared file or object storage, persistence after compute stops, backup, ingestion and export. Size throughput and concurrency, not only capacity. For multi-node training, validate that checkpoint writes and dataset reads do not starve the compute fabric or leave GPUs waiting for data.
Design the first node for the future cluster
If 32 or 64 GPUs are plausible, decide early whether the first node's NICs, optics, rack, power, cooling, storage and scheduler can join the future design. Reserve switch ports and facility headroom only where the growth plan is credible. A compatible first node reduces rework, but the cluster should still be justified by measured single-node utilisation and scaling evidence.
What the capacity proposal should contain
Require the exact system, GPU SKU and memory, CPU, RAM, local storage, external storage, NICs, fabric, rack and electrical requirements, cooling method, location, tenancy, support boundary, term, deployment window and acceptance process. Arvica can scope these elements for Australia and APAC. Exact availability, configuration, site, price and timing are confirmed in the written proposal—not inferred from this guide.
Frequently asked questions
How much power does an eight-GPU B300 server use?
NVIDIA lists 14.5 kW for DGX B300. HGX partner and OEM systems can differ, so use the selected complete-system specification for final design.
Does every B300 system require liquid cooling?
No. Cooling depends on the complete platform and rack design. Confirm the selected server's method and facility requirements.
Download an editable worksheet to agree scope, compare configurations and decide whether to stop, adjust or scale. An evaluation method, not a customer case study.
Compare scope, allocation, storage, networking and pilot acceptance before committing to a B300 node. Includes a worked cost formula, not a supplier price.
Build a representative inference test with model, precision, context, concurrency and quality targets. Connect the measured result to a transparent GPU budget.