待翻譯:AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale. This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software […] The post AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production appeared first on SiliconANGLE.
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale. This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software and operations must work together from Day 0 through continuous production. If one component fails to perform, expensive capacity sits idle and revenue slips away. In a series of exclusive interviews on theCUBE, I spoke with Cisco’s Will Eatherton, senior vice president and head of networking engineering, and NVIDIA’s Gilad Shainer, SVP of networking, and Marc Hamilton, VP of solutions architecture and engineering, about how the companies are tackling that execution challenge as neoclouds, sovereign AI programs and enterprises move infrastructure into production. We explored the engineering, operational and financial requirements behind building AI factories at rack scale. That is the strategic context behind the expansion of the Cisco Secure AI Factory with NVIDIA to rack-scale systems. Cisco Systems Inc., and NVIDIA Corp. are bringing together liquid-cooled compute, AI-optimized networking, validated designs and unified operations in an effort to compress the path between ordering infrastructure and producing the first token. The full rack-scale Secure AI Factory solution will be orderable through Cisco in September. The market has entered the execution era. The winners will be determined by how quickly they turn capital expenditures into productive capacity. Time to first token becomes the new infrastructure metric Neoclouds illustrate the urgency. Many have customers lined up before the GPUs arrive, which means deployment delays translate directly into deferred revenue. Enterprises face a similar issue as the cost of consuming models through external application programming interfaces grows. Sovereign AI programs add another layer of pressure because they must balance performance, data control and national infrastructure requirements. These different markets are converging around the same question: How quickly can an organization move from a purchase order to a production-ready AI factory? “When we work with neoclouds, from the moment that they put in the PO for the GPU, they already have their end customers lined up,” said Will Eatherton, senior vice president and head of networking engineering at Cisco. “And one of the challenges is just the speed, the expectation that from the moment that is all the project planned, that the GPUs have to come back, that then has to go through all of the final deployment aspects of software and then a bring up and hand over to their income.” The important shift is economic. Traditional enterprise infrastructure was frequently viewed as a cost center to be optimized downward. An AI factory is designed to manufacture a digital product: tokens. Those tokens power applications, agents and business processes, giving infrastructure a direct connection to revenue. That changes the scoreboard. Utilization, availability, tokens per second and tokens per watt become business metrics, not simply engineering measurements. “The real cost savings in an AI factory is not about cost savings, but it is about token generation and how do you drive that revenue — so having a repeatable way to do it,” said Marc Hamilton, vice president of solutions architecture and engineering at NVIDIA. A rack of components is not an AI factory The industry cannot approach this buildout using the old data center playbook. An AI factory operates as one enormous computing system assembled from thousands, and potentially hundreds of thousands, of components. GPUs, network interface cards, switches, cables, storage systems, models and software libraries must operate in concert. That makes architecture critical. NVIDIA describes the AI factory as a five-layer cake encompassing the data center’s land, power and shell; chips; infrastructure; models; and applications. Performance depends on optimizing across every layer. “Building an a factory, it’s not connecting components and hoping for the best,” said Gilad Shainer, senior vice president of networking at NVIDIA. “Building an AI factory means that you need to build a supercomputer and a supercomputer that needs to be built quickly, needs to be built fast and needs to provide the highest numbers of tokens per second, the highest numbers of tokens per power, and so forth.” Cisco’s expanded solution brings rack-scale systems and liquid cooling into an architecture supporting HGX and MGX form factors, including NVIDIA NVL72 systems and a path toward the Vera Rubin platform. Cisco wraps those systems with networking, software, sales and support. The networking architecture combines NVIDIA Spectrum-X Ethernet with Cisco technology. Spectrum-X provides the adaptive routing, congestion control, remote direct memory access and lossless capabilities required for distributed AI computing. Cisco Silicon One supports front-end, storage and data center interconnect requirements, while NX-OS or SONiC provides a familiar operating model. This is where the partnership gets interesting. NVIDIA brings infrastructure purpose-built for AI. Cisco brings the networking reach, enterprise operating model and installed experience required to connect AI factories with the data that already runs through the business. Reference architectures reduce financial and operational risk Reference architectures can sometimes be dismissed as technical checklists. That view misses their role in the AI factory market. An NVIDIA Cloud Partner reference architecture establishes how the entire system should be constructed, tested and operated. Cisco Validated Designs and Cisco Validated Infrastructure Services adapt those requirements to Cisco networking and management technologies. The goal is repeatability: Customers should not have to reinvent the system every time they deploy a cluster. The certification also has financial implications. Infrastructure lenders want confidence that the assets they finance will deliver the utilization, performance and availability required to support the investment. Following an established reference architecture can therefore affect financing terms, as well as technical performance. “We’ve actually had some of the largest finance lenders in the world say that they will only finance at preferred rates customers that are following that NCP reference architecture,” Hamilton said. “So, this is now a huge selling point for any Cisco customer, any Cisco sales rep.” Availability is where the risk becomes visible. A multibillion-dollar AI factory running at 50% or 60% availability because of cabling, firmware or software issues destroys the underlying economics. The first token matters, but the billionth token matters more. Validated infrastructure reduces the number of variables. It gives deployment teams a tested architecture, benchmarking tools and a common escalation path across Cisco and NVIDIA. In my view, that operational discipline will become one of the most valuable layers of the AI infrastructure stack. Day 2 operations will separate the winners Getting an AI factory online is only the beginning. Models change rapidly, inference software improves and organizations continually introduce new workloads. The system must be upgraded and optimized without sacrificing availability. “If you look at the Blackwell generation of GPUs, over the lifetime, or over the lifetime of that product, we’ve driven down the inference costs by x factors,” Hamilton said. “Not 1 or 2x, but 10x, 20x, 30x by going through and doing software optimization. So, being able to continuously upgrade that AI factory once you install it is super important.” Cisco is positioning Nexus One and Cisco Cloud Control as a common management layer across routing, front-end networking, storage networks and the Spectrum-X backend. The addition of AgenticOps creates an opportunity to apply AI-assisted monitoring and lifecycle management across the infrastructure. “A lot of the industry focus is up to the point that you light up the cluster and you get your first token out,” Eatherton said. “That’s been a big focus. That’s great. But the Day 2 two aspects around monitoring and health and availability and software upgrades, and these are things that from a Cisco standpoint, we’ve put a lot of focus on here over the years.” This becomes even more important as inference moves across on-premises systems, neoclouds and the edge. Enterprises want access to external capacity without creating a completely different operating environment. Neoclouds want to support multiple customers, departments and development environments with increasingly granular security and accounting controls. A common architecture can make those boundaries less visible. An enterprise should be able to run sensitive inference locally, burst into a neocloud and extend intelligence to factories, hospitals, telecommunications networks and other edge environments. The long-term opportunity is bigger than selling racks. It is about creating a repeatable operating model for intelligence. The AI infrastructure buildout may be one of the largest capital deployment cycles in computing history, but capital alone will not decide the outcome. Execution will. The market is moving from GPU scarcity to systems engineering, where networking, software and operations determine how much useful intelligence an organization produces from every dollar and watt. Time to first token gets an AI factory into the race. Continuous optimization, availability and scale are what win it. Here’s the complete interview playlist, part of SiliconANGLE’s and theCUBE’s exclusive coverage of the “Cisco Secure AI Factory With NVIDIA Expands to Rack Scale” event: Image: SiliconANGLE/ChatGPT A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media