Skip to content
sales@axislinktrading.com
AXISLINKTRADING LIMITED

GPU server buying guide: how to spec, source and deploy AI hardware

AI & GPU infrastructure buying · Published 2026-07-25

Buying a GPU server comes down to four decisions taken in the right order: match the GPU class to the workload, pick the form factor (PCIe or SXM), decide between new and refurbished silicon, and choose a platform with the power, cooling and connectivity that the GPUs demand. Get the order right and the rest is procurement mechanics; get it backwards (platform first, workload last) and you buy either too much machine or one that throttles.

This is the hub guide for our AI and GPU infrastructure series. It reflects how we actually spec and quote GPU systems for buyers at AXISLINK, and it links to deeper dives on pricing, export rules and specific model comparisons as those guides publish.

At a glance

  • Workload determines GPU class; everything else follows
  • SXM modules deliver full interconnect bandwidth; PCIe cards trade speed for flexibility
  • Refurbished datacenter GPUs can cut cost sharply if provenance is verified
  • An 8-GPU node is a power and cooling project, not just a server purchase
  • Advanced AI accelerators face export controls: plan compliance before ordering
  • Lead times vary by channel; secondary market moves faster than allocation queues

Start with the workload, not the GPU

The market talks in GPU model numbers, but specification starts with what the machine must do. Rough mapping as it stands in the current generation landscape:

WorkloadWhat matters mostTypical GPU class
LLM training and fine-tuningMemory capacity, interconnect bandwidth, multi-GPU scalingFlagship datacenter accelerators (H100/H200 class, SXM)
LLM inference at scaleMemory capacity per dollar, throughputH100/A100 class or high-memory inference cards (L40S class)
Diffusion, rendering, mediaRaw compute, VRAM, driver ecosystemL40S class or workstation cards (RTX 6000 family)
Classic ML, analytics, CVCost per GPU, single-node simplicityPrevious-generation datacenter cards (A100, even V100 class)
VDI and virtual workstationsLicense model, density per cardPurpose-built virtualization GPUs

Two practical corollaries. First, previous-generation flagships remain genuinely useful: an A100 cluster still trains and serves serious models, at a fraction of current-flagship pricing on the secondary market. Second, buying a bigger GPU than the workload needs does not future-proof anything if the interconnect, storage or network becomes the bottleneck first.

PCIe or SXM: the form factor fork

The same GPU generation usually ships in two forms, and the choice shapes the whole server.

PCIe cardsSXM modules
InstallationStandard slots in many server modelsSoldered to a dedicated baseboard (HGX class)
GPU-to-GPU bandwidthLimited; bridge connectors on some pairsFull-mesh high-bandwidth interconnect
Power and coolingLower per-card draw, air-coolable in 2UHighest draw; purpose-built chassis required
FlexibilityAdd or move cards between systemsFixed 4-GPU or 8-GPU configurations
Best fitInference, mixed workloads, gradual scalingMulti-GPU training where scaling efficiency pays

The honest rule of thumb from our quoting work: single-node inference and mixed research boxes are usually PCIe builds; serious multi-GPU training clusters are SXM builds, and the premium buys real training throughput rather than a spec-sheet number.

New or refurbished: what the secondary market really offers

Datacenter GPUs enter the secondary market when hyperscalers and AI labs refresh fleets. That supply makes previous-generation accelerators dramatically cheaper than new flagships, and it is how many teams afford real capacity. What separates a good refurbished purchase from a gamble:

  • Provenance: who ran the card, in what environment, decommissioned why
  • Testing: burn-in results per card, not per batch, with memory and thermal checks
  • Condition grading stated in writing, with photographs for used stock
  • Warranty: a stated return window from the seller, since manufacturer coverage rarely transfers
  • Firmware and licensing state: locked or vendor-tied cards exist and must be disclosed

This is a provenance business. The reason buyers route refurbished purchases through an accountable trading counterparty rather than a marketplace listing is that verification, grading and recourse are the actual product; the silicon is the same either way. Our GPU and accelerator stock lists condition and grading explicitly for exactly that reason.

The platform: the server around the GPUs

GPUs do not run themselves. The platform decision is where budgets quietly break.

Chassis and density

Common shapes: 2U servers taking 2 to 4 PCIe cards; 4U/5U towers and rack chassis taking 4 to 8 PCIe cards with better airflow; and purpose-built 6U-8U HGX platforms carrying 8 SXM modules (Dell PowerEdge XE9680 class, HPE Cray class, Supermicro HGX systems). A chassis that fits the cards is not the same as a chassis that cools them at sustained load.

Power is the hidden constraint

An 8-GPU flagship node draws on the order of 10 kW under load. That is beyond many office server rooms and a meaningful fraction of a rack budget in a colo. Before ordering, confirm rack power budget and PDU capacity, cooling capacity for the heat actually produced, and whether the facility charges by provisioned power. More than once we have re-quoted a buyer from one 8-GPU node to two 4-GPU nodes purely because of facility power limits.

CPU, memory, storage, network

Balance rules that hold across generations: system memory at least matching total GPU memory (twice is comfortable for training); NVMe scratch fast enough to keep the GPUs fed; and for multi-node training, a fabric (high-speed Ethernet or InfiniBand class) that matches the ambition, because a starved interconnect idles expensive silicon.

Export controls: plan compliance before you order

Advanced AI accelerators are subject to export controls, most prominently the US Bureau of Industry and Security rules restricting shipment of top-tier datacenter GPUs to certain destinations. What this means for a buyer in practice:

  • Flagship accelerator availability differs by destination market; restricted destinations cannot lawfully receive certain SKUs
  • Sellers will ask for end-user and end-use information on controlled models; treat that as a sign of a lawful counterparty, not friction
  • Regulations change; the current rule set, not last year's summary, governs your shipment

We state this plainly in our own terms: orders may require end-user information, and we decline transactions that cannot be lawfully fulfilled. A dedicated guide in this series unpacks the rules for overseas buyers in more depth.

Sourcing channels and lead times

Three broad channels supply GPU systems. OEM allocation (ordering a configured system from Dell, HPE, Supermicro or Lenovo channels) delivers warranty-backed new systems, with lead times that stretch when demand spikes. The secondary and refurbished market moves fastest and prices previous generations attractively, with quality riding entirely on verification. Trading companies like ours sit across both: sourcing new or refurbished GPUs and platforms, verifying and consolidating them, and shipping with export documentation from Hong Kong or mainland China. Which channel wins depends on whether your constraint is budget, lead time or warranty coverage.

A worked example of the decision path

A buyer wants to fine-tune mid-size open models and serve them to a few thousand users. Decision path: workload is mixed training and inference; memory capacity matters more than peak interconnect, so 4 to 8 PCIe accelerators in the H100/A100 class beat an SXM baseboard on cost; refurbished A100 80GB cards with documented provenance halve the silicon budget; platform is a 4U chassis with redundant power and NVMe scratch; facility check confirms the rack can feed it. That is a complete, defensible specification, and every step of it is a conversation we have with buyers weekly.

FAQ

What is the difference between a GPU server and a regular server?

A GPU server is built around accelerator power and cooling: higher-wattage power supplies, airflow designed for cards or modules at sustained load, and PCIe or SXM topology to connect them. A regular server can sometimes take one or two cards; a GPU server is designed for a full complement.

Should I buy or rent GPU compute?

Renting wins for bursty or exploratory work; owning wins when utilization is sustained, data gravity or privacy matters, or cloud GPU pricing exceeds amortized hardware cost for your duty cycle. Many teams rent to prototype and buy for production. A separate guide in this series works through the arithmetic.

Are refurbished datacenter GPUs reliable?

Datacenter cards are engineered for continuous operation, and well-decommissioned fleet cards with per-card test results have long useful lives. The risk concentrates in provenance and testing, not in the silicon concept: buy graded, tested, documented stock from an accountable seller.

How much power does an 8-GPU server need?

Flagship 8-GPU SXM nodes draw on the order of 10 kW at load; PCIe builds less, but still far beyond ordinary office wiring. Confirm facility power and cooling before ordering, not after delivery.

Can I upgrade a 4-GPU server to 8 GPUs later?

On PCIe platforms, only if the chassis, power supplies and airflow were sized for it from the start. SXM baseboards are fixed configurations. If growth is likely, buy the larger platform with fewer cards and populate it over time.

Sourcing this hardware?

AXISLINK supplies business buyers worldwide from Hong Kong and mainland China, with export documentation and freight arranged. Send your requirements and receive a formal quotation, typically within 24 hours.

Request a quote