Training the largest AI models is no longer simply a question of adding more GPUs. As clusters grow, power availability is becoming a hard physical constraint, forcing infrastructure teams to consider how a single training workload can operate across multiple data centres. That creates a very different networking problem. Traditional datacenter interconnect was designed to move traffic between sites. AI training demands something more exacting: huge, synchronous flows, minimal packet loss and tightly coordinated communication between GPUs. If one part of the cluster stalls, the impact can ripple across the entire job. In this interview, The Register’s Tim Phillips speaks with Rakesh Chopra, SVP, Silicon and Systems Architecture, Cisco...
Les hele artikkelen hos kilden.
Kommentarer (0)
Ingen kommentarer ennå. Bli den første til å kommentere!