The usual story of a failed edge AI deployment is not a bad model. It is a server that reboots under load, a breaker that trips at 3 am, or a GPU that quietly halves its clock speed because the cupboard it lives in is at 40 °C. Power is the first bottleneck you hit when you take inference out of the cloud. It is also the easiest to get right if you do the arithmetic before you buy.
This guide covers sizing the power supply (PSU) inside a GPU inference box, the circuit it plugs into, and the UPS in front of it, with figures taken from manufacturer spec sheets and Intel’s ATX design guide. It is general guidance. Any work on fixed wiring in Australia must be done by a licensed electrician to AS/NZS 3000.
Step 1: understand what the spec sheet power numbers mean
GPU vendors publish several different numbers, and mixing them up is the most common sizing mistake.
- Total board power (TBP) or total graphics power (TGP) is the sustained power the card is designed to draw at its default limit. NVIDIA’s datasheet lists the RTX PRO 4000 Blackwell at 140 W total board power (its product page quotes 145 W maximum). The GeForce RTX 5090 is listed at 575 W total graphics power.
- Required or recommended system power is the vendor’s suggestion for the whole PC’s PSU rating. For the RTX 5090, NVIDIA states 1000 W, “based on a PC configured with a Ryzen 9 9950X processor”, and recommends a PCIe CEM 5.1 compliant PSU.
- Configurable power limits. Data-centre and workstation cards often allow the limit to be set lower. The RTX PRO 6000 Blackwell Server Edition is listed at “up to 600 W (configurable)”, and Lenovo’s product guide notes it can be capped at 450 W for density. The Max-Q workstation variant is specified at 300 W.
None of these numbers describes the transient behaviour of the card. Modern GPUs change load in microseconds, and the brief peaks can be far above TBP. That is what trips PSUs.
Step 2: build a sustained power budget
Start with sustained draw for every component at the power limits you will actually run:
DC load (W) = Σ GPU power limits
+ CPU sustained package power (PL1/PPT or server TDP)
+ motherboard, RAM, NICs (~50–100 W)
+ storage (~5–10 W per NVMe, ~8–10 W per HDD)
+ fans and pumps (~2–5 W each, more for server fans)
Worked example: a two-GPU inference box
- 2 × RTX PRO 6000 Blackwell Max-Q at 300 W = 600 W
- A workstation CPU at about 350 W sustained = 350 W
- Board, 256 GB RAM, 25 GbE NIC = 100 W
- 4 × NVMe = 30 W
- 8 fans = 30 W
Sustained DC load ≈ 1,110 W.
Step 3: add headroom for transients and ageing
Intel’s ATX Version 3 power supply design guide is the reference for how desktop PSUs must handle GPU transients. For PSUs rated above 450 W and fitted with a 12V-2x6 connector, it requires tolerance of these power excursions:
| Excursion (% of PSU rating) | Duration | Test duty cycle |
|---|---|---|
| 200% | 100 µs | 5% |
| 180% | 1 ms | 8% |
| 160% | 10 ms | 12.5% |
| 120% | 100 ms | 25% |
| 100% | continuous | – |
Source: Intel ATX Version 3 Multi Rail Desktop Platform Power Supply Design Guide, Rev 2.1a, Table 3-3. PSUs ≤450 W or without a 12V-2x6 connector have lower limits (150% / 145% / 135% / 110%).
What this means in practice:
- Buy ATX 3.1 (or later) PSUs for any GPU box. A pre-ATX 3.0 unit sized “correctly” on paper can trip its over-current protection on a millisecond GPU spike that an ATX 3.1 unit rides through.
- Do not run the PSU at its limit. Aim for sustained DC load at about 60–75% of the PSU rating. Efficiency curves usually peak around 50% load, fan noise is lower, components run cooler, and there is margin for capacitor ageing.
- Treat vendor “required system power” as a floor, not a target.
For our worked example, 1,110 W ÷ 0.7 ≈ 1,590 W, which points to a 1,600 W ATX 3.1 PSU. A 1,300 W unit would technically carry the sustained load but leaves little room for the next upgrade.
Connectors: 12V-2x6 and cable discipline
High-power GPUs now use the 16-pin 12V-2x6 connector (the revised version of 12VHPWR). PCI-SIG’s CEM 5.x specification calls it the PCIe CEM5 16-pin connector, rated up to 600 W per connector. Some practical rules:
- Use native PSU cables where you can. If you must use an adapter, use the one supplied by the GPU maker. NVIDIA allows the RTX 5090 to be powered from four PCIe 8-pin cables through its adapter, or one 600 W PCIe Gen 5 cable.
- Seat the connector fully and avoid tight bends close to the plug. Partially seated connectors concentrate current on fewer pins.
- One cable per 8-pin input. Do not daisy-chain pigtails on high-draw cards.
- Check the PSU’s 12V-2x6 count. Two 600 W GPUs need two native 600 W cables or a correct adapter configuration.
Step 4: convert DC load to wall power
The PSU is not 100% efficient, and the circuit, UPS and electricity bill all see AC wall power:
Wall power (W) ≈ DC load ÷ PSU efficiency at that load
At 80 PLUS Platinum or Titanium levels, efficiency at mid load is typically in the low-to-mid 90s per cent. For our example:
- 1,110 W ÷ 0.92 ≈ 1,205 W at the wall.
Measure it. A plug-in power meter, a metered PDU or the UPS’s own load reading will give you real numbers under your own workload. Inference loads vary a lot: a GPU waiting for requests draws a fraction of its limit, while long batch jobs pin it at the cap.
Step 5: check the circuit
This is where offices and home labs come unstuck. A standard Australian power point is rated at 10 A, and the nominal supply voltage is 230 V (AS 60038), so one outlet is good for roughly 2,300 W. Several points normally share a circuit protected by a single breaker, and that circuit may also feed a kettle, a heater or a photocopier.
For our 1,205 W box:
- It fits on one 10 A outlet with headroom on paper.
- Two of them on one office circuit, at about 2,400 W combined, would exceed a 10 A rating once everything else on the circuit is counted.
- Power-on inrush current, when PSU capacitors charge, can trip sensitive breakers if several machines start together after an outage. Stagger restarts in the BIOS (“restore on AC power loss” with a delay) or use a sequenced PDU.
Continuous high loads also heat plugs and sockets. If you plan to run multi-GPU systems around the clock in an office, have a licensed electrician assess the circuit and, if needed, install a dedicated circuit or a 15 A outlet. Never plug a 15 A plug into a 10 A socket using an adapter.
Step 6: size the UPS
A UPS carries two ratings: VA (apparent power) and W (real power). Modern PSUs with active power factor correction draw close to unity power factor, so the watt rating is the real limit. For example, APC’s Smart-UPS SMT1500RMI2UC is rated at 1500 VA but 1000 W, which is too small for our 1,205 W example even though “1500” looks like enough.
Sizing rules:
- UPS watt rating ≥ measured wall draw × 1.25.
- Check the runtime chart at your load. Runtime falls steeply as load rises, and a UPS that gives 30 minutes at 300 W may give five minutes at 1,000 W.
- Decide what the UPS is for. In an office it is usually to ride through brief interruptions and shut down cleanly. Configure automatic shutdown (NUT or the vendor’s agent) so a long outage does not corrupt a model store or database.
- Line-interactive units suit most office loads. Online (double-conversion) units give cleaner output and zero transfer time, at a higher cost and with more heat.
Step 7: remember that every watt becomes heat
All of the electrical power ends up as heat in the room: 1 W is about 3.41 BTU/h. Our 1,205 W box is a 4,100 BTU/h heater that never switches off. A small office comms cupboard with a door and no dedicated cooling will heat-soak within hours, and GPUs respond by thermal throttling, cutting clocks to protect themselves. NVIDIA’s own spec page lists a maximum GPU temperature of 90 °C for the RTX 5090.
Two things follow:
- Budget the cooling energy too. Air-conditioning a room to remove that heat adds to your power bill. A common planning allowance is 30–50% on top of IT load for small rooms with split systems. The cost calculator on Aussie Racks lets you set this overhead yourself.
- Airflow direction matters. Passively cooled data-centre cards such as the L40S and RTX PRO 6000 Server Edition rely on the server chassis pushing air through them. Put one in a desktop tower and it will overheat. Workstation cards with blower or active coolers are the right choice for towers.
A one-page sizing checklist
- List every component with its configured power limit, not its marketing peak.
- Sum to a sustained DC load.
- Choose an ATX 3.1 PSU so that the sustained load is about 60–75% of its rating.
- Confirm 12V-2x6 / PCIe connector counts and use native cables.
- Convert to wall power using the PSU’s efficiency at that load, then measure it.
- Check the circuit: 10 A ≈ 2,300 W per outlet at 230 V, shared across the circuit.
- Size the UPS in watts with 25% headroom, and check runtime at load.
- Plan cooling for every watt, and include its energy cost.
When the cupboard is not enough
If the arithmetic says you need a dedicated circuit, an online UPS and a new air conditioner, it is worth comparing that against colocation. In a facility, power, conditioned cooling and suppression are already built. Green Racks’ Osborne Park, WA facility provides 2N UPS (150kVA), dual A/B feeds per rack, N+1 cooling and NOVEC fire suppression, with an on-site generator coming soon.
Sources
- Intel ATX Version 3 Multi Rail Desktop Platform Power Supply Design Guide, Rev 2.1a (PDF)
- NVIDIA RTX PRO 4000 Blackwell datasheet (PDF)
- NVIDIA GeForce RTX 5090 specifications
- NVIDIA RTX PRO 6000 Blackwell Server Edition
- NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
- Lenovo Press: ThinkSystem NVIDIA RTX PRO 6000 Blackwell Server Edition product guide
- APC Smart-UPS SMT1500RMI2UC product page