AI News HubLIVE
サイト内リライト3 分で読了

翻訳待ち:AI inference gets a new tier as context windows grow

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference. Agentic AI is also changing the shape of the data problem. Interactions are growing longer […] The post AI inference gets a new tier as context windows grow appeared first on SiliconANGLE.

ソースSiliconANGLE AI著者: Victoria Gayton

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference. Agentic AI is also changing the shape of the data problem. Interactions are growing longer and producing more information. At the same time, the data’s size, importance and movement through the infrastructure can all affect how quickly an application responds, according to Scott Shadley (pictured, left), director of technology planning at Solidigm Inc. “One of the beautiful things that’s happened in this agentic AI, or even just the AI era, is [that] people are starting to pay attention to storage,” he said. “What’s unique about this particular era is it’s no longer one- or two-dimensional. We have data magnitude and growth in size, importance and all of the other volumetric aspects of that. But at the end of the day, it comes down to that bit of data and how fast that bit of data moves from point A to point B.” Shadley, along with Anat Heilper (center), director of AI architecture at Vast Data Inc., and Ben Lee (right), director of solution management at Super Micro Computer Inc., spoke with theCUBE Research’s Rob Strechay during the Supermicro Open Storage Summit interview series. They discussed how the companies’ respective technologies fit together as expanding context windows and KV caches create new storage and memory demands for agentic AI. AI storage infrastructure brings KV cache closer to compute As context windows expand, graphics processing unit memory alone can’t hold everything an agentic workload needs during inference. AI storage infrastructure must therefore provide additional tiers that balance proximity, capacity and speed, with each layer handling a different part of the data load, according to Shadley. “As you think through that architecture — and you need to put that context somewhere — that context can start living in what used to be a no-no zone,” he said. “So [solid-state drives] have found a new home. One of the unique things about this 3.5 tier that we’re creating is that it could not exist until we had things like [Non-Volatile Memory Express] SSDs.” That middle tier is one part of a larger architecture. Solidigm supplies SSDs that provide fast access to cached data, Supermicro integrates them into rack-scale systems and Vast Data’s AI Operating System provides network storage and the data services needed to use, manage and protect that data alongside broader AI workloads, Heilper noted. “When we talk about [key-value] cache, which is a very significant optimization that can be done in AI inferencing … in essence, it’s the ability to replace compute with storage,” she said. “This is very significant because we all know the GPU is very, very expensive. When you have very high KV cache hit rates, we both save on compute and reduce the latency significantly.” AI infrastructure needs room to evolve Supermicro’s Context Memory eXtension, or CMX, proposal targets organizations with large AI clusters and substantial data demands; other deployments may require different combinations of memory, local SSDs and network storage. The company’s broader value lies in composing those building blocks around each customer’s workload, rather than treating one architecture as a universal answer, according to Lee. “We believe solving the problem will take the whole rack because you cannot just buy more GPUs with more [high-bandwidth memory] … it’s very expensive,” he said. “All the KV cache will naturally overflow from the GPU, HBM, to the system memory, to the local SSD and to the network storage. But there’s a new industrial definition to try and fill the gap, and they call it G3.5, which is the CMX solution. This is a very AI-native KV cache tier that can fulfill the demand.” Vast Data has been testing KV cache offload with partners across different software environments. Those experiments aim to show how the architecture behaves when cached context is integrated into production inference workloads, according to Heilper. “With Nvidia Dynamo, we’ve shown that we managed to get 20 times faster time-to-first-token, which means that the latency that you perceive as a user is significantly faster,” she said. “Also, [we’ve] seen 90% savings in GPU time.” Those results depend on an AI storage infrastructure that can match storage performance and capacity to the workload. Solidigm’s D7-PS1010 performance-oriented SSD and D5-P5336 capacity drive address different points in the hierarchy, according to Shadley. Supermicro integrates those components into systems that can be configured around customer requirements. “This is not the only definition of a hierarchy stack,” Shadley said. “It’s the current primary that everybody leverages as the gold standard, but it’s continuing to evolve, and it’s unique. Being very proactive with your customer or your supplier to better understand what they know about what you need is no longer transactional. Those value-level transaction conversations are now what are going to drive the future of us deploying these types of architectures.” Stay tuned for the complete video, part of SiliconANGLE and theCUBE’s coverage of the Supermicro Open Storage Summit interview series. (* Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.) Photo: SiliconANGLE A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media