A100 vs H100 vs consumer GPUs: which should you rent?
Choose a rental GPU using model memory, measured throughput and the cost of finishing your workload—not the hardware name alone.
· Arvica Cloud · Analysis & buying guide
Which GPU should I shortlist?
Start with a configuration that can hold your workload at the required precision and context length. A smaller workload may justify a consumer GPU trial; a workload that needs more device memory or multi-GPU communication deserves a data-centre configuration trial. Request the exact SKU, usable memory, GPU count, CPU, RAM, storage and interconnect. NVIDIA lists distinct H100 form factors and specifications; an H100 label alone is not a complete server quote. [1]
Is the faster card always cheaper to use?
Compare the price of completing the same accepted job. In an illustrative example, four GPU-hours at A$2 cost A$8; two GPU-hours at A$3 cost A$6. These are invented rates, not offers. Repeat the same workload, precision and quality checks on both configurations, then add setup time and other charges. For inference, compare cost at your response-time target and expected concurrency, not only maximum tokens per second.
Can several smaller cards replace one larger card?
Only if the model and serving or training framework support the required partitioning. Treat memory fit and speed as separate acceptance checks. Communication, CPU resources and the exact topology can change the outcome. Ask for a representative multi-GPU test before paying for a long commitment; adding cards is not a guarantee of proportional speed.
What should I include in a quote request?
Send the model name and revision, training or inference task, precision, peak concurrency or batch size, required location, GPU count if known, start date and rental term. Ask Arvica to compare feasible configurations against those requirements. Confirm availability, commercial terms and trial acceptance in the written quote.
Sources
Reviewed: 2026-09-12
Apply this to your workload
Further reading
LoRA fine-tuning vs inference: do you need the same GPU? →
GPU rental total cost: what is included beyond the hourly rate? →
Is private company data safe on rented cloud GPUs? →
How much GPU memory does a 70B model really need? →
Your GPU is busy waiting: investigate the pipeline before renting more →
Eight GPUs are not automatically eight times faster →
Before the long reservation: a GPU server acceptance checklist →
AI inference economics: why the GPU-hour is only the starting point →
B300 and Rubin: plan the workload before chasing the roadmap →
AI data-centre power demand: what compute buyers should ask →
Reserved vs on-demand GPUs: calculate the utilisation break-even →