翻訳待ち:Nvidia open-sources cuFile API, accelerating GPU read/write capability for high-speed storage
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:As artificial intelligence applications become ever hungrier for faster access to data, Nvidia Corp. today announced it is open-sourcing the application programming interface for its powerful cuFile vertical data storage stack, enabling millisecond data access. The company also announced a large-scale industry initiative with technology leaders to optimize memory and storage with Storage-Next. The initiative […] The post Nvidia open-sources cuFile API, accelerating GPU read/write capability for high-speed storage appeared first on SiliconANGLE.
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
As artificial intelligence applications become ever hungrier for faster access to data, Nvidia Corp. today announced it is open-sourcing the application programming interface for its powerful cuFile vertical data storage stack, enabling millisecond data access. The company also announced a large-scale industry initiative with technology leaders to optimize memory and storage with Storage-Next. The initiative aims to bring storage makers, controller vendors, thermal design, cooling, and orchestration providers together with standard bodies to find opportunities for superior graphics processing unit-driven storage. During Nvidia’s announcement at the Future of Memory and Storage conference, the company said the cuFile API enables secure access from storage in just milliseconds. CuFile was first officially launched into general availability in July 2021 alongside the company’s GPU access CUDA Toolkit 11.4. Prior to the launch, it acted as a core component of the company’s GPUDirect Storage software. It provides direct, high-speed data access between local distributed storage and GPU memory; this allows GPUs to bypass the central processing unit and system main memory, which greatly reduces memory access delays. Using direct memory access, or DMA, it can move data directly from storage devices such as NVMe drives into GPU memory. The primary source of AI inference happens on GPUs and specialized cards that interact primarily with GPU memory; refreshing that memory is often completed by another software layer that uses the GPU to orchestrate and make decisions. CuFile cuts out the “middle man,” so to speak, providing a direct access path and enabling GPU orchestration, meaning it can rapidly pull more data. This allows the synchronization of massive data sets at rates that can keep up with elite GPU speeds. If the bandwidth to this data is delayed or pinched off by a bottleneck, it causes something called the “GPU starvation” loop, where some GPUs in a distributed array sit idle waiting for data to be ready. Increased speeds and low millisecond delays allow massive datasets from retrieval-augmented generation and agentic AI to run faster by providing them a low-latency “superhighway” between deep storage and the GPU itself. Bringing open standards to GPU accessible storage Storage-Next includes 40 leading storage and flash memory vendors, including DataDirect Networks, Inc., Kioxia Corp. and Micron Technology Inc. Each will contribute to the next generation of AI storage technologies with Nvidia. The company will build its grounded, high-speed data access for storage off SCADA, short for scaled, accelerated data access. This solution provides the groundwork for massively parallel GPUs to pull data at high bandwidth and extremely low delay. Modern AI training requires massive data access all at once and AI inference now includes colossal mixture-of-experts models, which “think” extremely quickly, rapidly pulling data and firing tools for agentic workflows. “AI success will be defined not by how much infrastructure organizations own, but by how productively they use it,” said DDN Chief Technology Officer Sven Oehme. “Our collaboration with Nvidia is helping create a more direct, efficient connection between GPUs and data — keeping accelerated computing resources productive, speeding time to insight and enabling customers to achieve stronger business and financial returns from their AI investments.” Nvidia said SCADA allows high-speed storage layers to operate securely. Technologies such as SCADA allow direct access to storage media by the AI; this punches through safeguards designed to prevent clobber – what happens when two processes attempt to write to the same data at once, with one scribbling over the other’s work. This is a security problem, not a proper feature. The other side of the coin means that this might bypass security safeguards such as encrypted and privileged access. This opens up another security hole that malicious parties could exploit. SCADA overcomes this challenge by scaling direct access by splitting access across different jobs: the user parts of the application receive raw speed but stay outside the encrypted, secure and trusted computing base; and a separate privileged computing component configures protected access following Linux protocols to prevent clobber. Image: Nvidia A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.