AI Cluster Fabric Planner
19 Signals+1

AI Cluster Fabric Planner

A planning and simulation tool for choosing GPU cluster network topology, optics, and power tradeoffs before procurement.

Added Jun 18, 2026

AI infrastructure
data center networking
procurement software
Opportunity Score
Opportunity: Medium (59%)
Evidence Strength
Vol: 30%
Urg: 72%
Spec: 72%
Market Analysis
medium
The Problem

AI infrastructure buyers are no longer bottlenecked only by GPU count; cluster performance now depends heavily on backend fabric design, optics power draw, flow control, topology, and interconnect compatibility. Hyperscalers are using divergent approaches such as InfiniBand, custom Ethernet, optical circuit switching, SRD-style transport, and scale-up links, making vendor comparison difficult. Network architects need a concrete way to estimate whether a proposed cluster fabric will support training workloads without excess latency, congestion, power waste, or tenant-isolation risk.

Potential Solution

Build a SaaS tool where AI infrastructure teams model a planned GPU cluster by entering GPU type, rack count, switch family, optics type, topology, RDMA protocol, and expected collective communication pattern. The product produces topology diagrams, estimated fabric bandwidth, transceiver power load, congestion risk, failure-domain analysis, and a vendor-neutral comparison of Ethernet, InfiniBand, and emerging optical approaches. The first version can focus on procurement-stage planning using public vendor specs, user-entered BOMs, and simple all-reduce traffic models.

Why Now?

AI clusters are scaling past tens of thousands of GPUs while networking and optics consume a growing share of power and cost. New products such as 800G/1.6T switching, co-packaged optics, rail-optimized fabrics, and open accelerator interconnects make architecture decisions more complex and expensive to reverse.

Showing 1-19 of 19 signals

Principal Infrastructure Engineer, AI Cluster Performance & Validation
nscaleSep 3, 2026

Drive cluster optimization end to end, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput. Partner with Infrastructure, Platform, SRE, and customer-facing teams to translat

embedding
Principal Solution Specialist, Core Services
coreweaveAug 19, 2026

Design the commercial framework for large-scale GPU reservation deals, including MFU modeling, cluster sizing, and network bandwidth commitments that support large enterprise closings. Partner with capacity, infrastructure, and networking teams to maintain a competitive edge on compute density, interconnect performance, and platform reliability across active and prospective customer deployments.

embedding
Telecom Conveyance Engineer, Data Center Infrastructure
metaAug 11, 2026

Establish pathway capacity planning frameworks that account for current fill ratios, future growth, and evolving cable densities driven by AI/GPU infrastructure Create design automation tools, templates, and parametric models that enable scalable pathway design across multiple facility types and regions

embedding
Google Trends: GPU cluster networking
Google TrendsAug 10, 2026

Search interest for GPU cluster networking has a recent median of 35.0, a prior baseline of 0.0, and a momentum score of 1.00.

source
Senior Network Engineer (3x Openings)
voltaAug 10, 2026

GPU fabric: RoCE v2 and InfiniBand for training traffic, congestion control tuning, rail-optimized topology, and the performance validation that proves a cluster is fit for workloads. SDN and overlay integration: the software boundary between the fabric and the platform, controller and API-driven fabric programming, and multi-tenant network provisioning.

embedding

+16 more signals