ai-newspaper.
Infrastructure & Hardware

Together AI Secures $240M IBM Cloud Deal to Scale Inference Capacity

A $240 million commitment from Together AI to deploy Nvidia HGX B300 systems inside IBM Cloud data centers signals how aggressively inference providers are absorbing last-generation accelerators to…

Together AI Secures $240M IBM Cloud Deal to Scale Inference Capacity

A $240 million commitment from Together AI to deploy Nvidia HGX B300 systems inside IBM Cloud data centers signals how aggressively inference providers are absorbing last-generation accelerators to secure capacity, according to reporting from The Register. The cluster, scheduled to come online in Q1 2027, marks IBM's first large-scale deployment of B300 silicon for inference workloads — a hardware profile optimized for token-throughput economics rather than peak training FLOPs.

The hardware contract

The deal centers on Nvidia's HGX B300 platform, announced in early 2025 and positioned well below the 72-GPU rack-scale systems that dominate Nvidia's marketing material. Each B300 node draws 14–15 kW, houses eight GPUs, and uses NVLink for intra-box bandwidth alongside Spectrum-X Ethernet for inter-rack fabric — a topology compatible with conventional air-cooled data centers rather than the liquid-cooled halls required for flagship accelerators. IBM framed the partnership around delivery velocity: its announcement notes that "Together AI selected IBM with Nvidia because of their innovative product roadmaps and their ability to deliver GPU capacity at the pace required for rapid AI scaling and lowest token cost."

Why B300, and why now

The architecture choice reflects the inference economics that drive Together AI's business. The provider operates primarily as a meta-layer, renting compute from hyperscalers and neoclouds while exposing an OpenAI-compatible API endpoint for open-weight models, and has begun placing GPU hardware directly in sites spanning Maryland, Memphis, and Sweden. Its tolerance for hardware heterogeneity has already pushed it toward SambaNova's Intel-Nvidia prefill platform inside Vector Core Compute's facility; the IBM arrangement extends that flexibility to a more conventional 8-GPU building block. Compared with Nvidia's higher-tier SKUs, the B300 trades peak FLOPs and memory bandwidth for higher availability and lower per-token cost — a calculus that favors sustained inference throughput over training-grade numerics. The Q1 2027 availability window also coincides with a generation shift in the open-weight model ecosystem, where quantization-friendly architectures will set the floor for viable serving configurations.

For developers consuming Together AI's endpoints, the practical implication is continuity rather than disruption: token pricing and latency characteristics will track the underlying B300 memory bandwidth and interconnect topology, and capacity pressure may ease relative to H100-saturated neoclouds. The more interesting variable to monitor is whether IBM extends B300 inference capacity to additional tenants, or whether the deployment remains a single-customer configuration through 2027.