Beyond GPUs: Why Data Center Infrastructure is Being Redesigned for AI Inference
As SiliconANGLE's coverage of the Supermicro Open Storage Summit details, the AI inference stack has moved beyond GPU FLOPS to a system-level engineering problem where storage latency, network…

As SiliconANGLE's coverage of the Supermicro Open Storage Summit details, the AI inference stack has moved beyond GPU FLOPS to a system-level engineering problem where storage latency, network bandwidth, and per-watt I/O now dictate token economics. The reporting draws on interviews with IBM's Ka Wai Leung, Super Micro's William Li, and Kioxia's Anders Graham, who framed inference as a workload-class problem and quantified the generational storage gap in IOPS-per-watt. For data center operators, the redesign is no longer optional — power and physical capacity ceilings are forcing a co-design of the inference stack that spans accelerators, storage, and networking.
Workload-class differentiation
Training workloads tolerate high sequential throughput as checkpointing and gradient passes amortize cost across long-running contexts. Inference inverts the dependency. Interactive chat is bound at the tail of latency percentiles; batch inference optimizes tokens-per-second-per-watt; agentic systems consume expanding context windows that continually retrieve from proprietary corpora. Graham positioned 2026 as the year of inference, arguing that low-latency random reads and retrieval-augmented generation make storage the new bottleneck layer. Leung added a structural overlay: enterprise data remains scattered across mainframes, object stores, and operational systems, and naive ingestion into AI factories conflicts with data gravity and data sovereignty constraints. The recommended pattern is workload-aware topology selection — choosing the storage, network, and compute substrate that matches the dominant request profile — rather than a uniform stack.
The NAND generation gap
Kioxia's transition from BiCS8-based CM7 to CM9 drives produced measurable per-watt I/O gains. Graham cited a 76% improvement in random read IOPS-per-watt and more than 100% improvement in random write IOPS-per-watt on equivalent workloads. The numbers are not cosmetic: with data center operators hitting power and physical capacity ceilings, NAND generation and controller architecture become first-order cost levers rather than commodity substitutions. Supermicro's reference designs pair these storage tiers with Nvidia accelerators, indicating that the inference stack is now being co-designed at the SKU level rather than assembled opportunistically.
Capital reallocation around the inference stack
The hardware-side shift is mirrored in capital flows. VivoPower announced a separate AI Infrastructure Platform headquartered in Singapore to manage its non-Nordic 2.2GW data center portfolio, concentrated in GCC and ASEAN markets, with a planned primary listing on the London Stock Exchange and secondary listing on the Abu Dhabi Securities Exchange. The structure transfers future construction capital expenditure to the new entity while VivoPower retains de facto control, isolating funding obligations from the parent. In Europe, Velatir secured a €5M seed round to expand AI infrastructure deployment, while Airgain reported two IoT design wins tied to AI data center infrastructure — the latter two without disclosed specification detail in available reporting. The pattern is consistent with the technical analysis: the bottleneck layers — storage, networking, power delivery, edge telemetry — are attracting discrete funding vehicles rather than being absorbed into generic compute capex. For cluster architects, the practical implication is that inference procurement decisions must increasingly be evaluated at the rack level — through tokens-per-second-per-watt accounting — rather than through accelerator benchmarks alone.