Skip to content
AI News HubLIVE
In-site rewrite2 min read

Tsinghua-Affiliated Team Builds a 'Smart Computing Power Grid' for Large Models

Summary

A startup founded by Tsinghua University alumni, Shishi Technology, has developed a proprietary parallel optimization technology that integrates heterogeneous computing resources and inference optimization engines, reducing per-token cost by 40%. The company aims to build a domestic token optimization factory to lower the barrier for AI deployment.

Source量子位Author: 思邈
Tsinghua-Affiliated Team Builds a 'Smart Computing Power Grid' for Large Models
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

A team from Tsinghua University has developed an innovative solution to the AI chip utilization problem: a "smart computing power grid" that efficiently converts diverse computing resources into affordable, standardized token production. Shishi Technology, founded in 2021 by researchers from the National Supercomputing Center in Wuxi, addresses the critical bottleneck in AI deployment—the stable, low-cost generation of tokens.

The company's approach hinges on integrating high-performance computing (HPC) with artificial intelligence, enabling the effective scheduling and optimization of heterogeneous computing resources. This includes not only NVIDIA GPUs but also a wide range of domestic AI chips such as those from Huawei Ascend, Kunlun Core, and others. By creating a unified resource pool with intelligent scheduling, Shishi ensures that any available compute capacity can be dynamically allocated to meet user demands, much like an electrical grid distributes power from various sources.

The core technology lies in their 'Token Optimization Factory,' which combines hardware pooling with deep inference optimization. The team has achieved significant improvements through techniques like CUDA kernel optimization, PagedAttention, continuous batching, mixed-precision inference, FlashAttention, speculative decoding, and KV cache management. As a result, compared to baseline setups, they can boost throughput by 30-50% while slashing the cost per token by 40%.

Reliability is another key pillar. Shishi has built a multi-provider failover system with automatic fallback, ensuring 99.9% availability. This architecture, akin to redundant engines on an aircraft, allows seamless switching between primary and backup resources in milliseconds, preventing service interruptions. The company's vision is to become China's largest and most advanced domestic token optimization factory, empowering industries to leverage AI with efficient, standardized computing power that maximizes the use of domestic chips.

In the long run, Shishi Technology aims to democratize AI access by making token production as reliable and cost-effective as electricity from a grid. By focusing on standardization, localization, and efficiency, they are building a foundational infrastructure for China's AI ecosystem. This move is critical as the country seeks to reduce dependence on foreign GPUs and harness the potential of its own chip manufacturing capabilities. The company's success could significantly lower the barriers for enterprises to adopt AI, accelerating the industrial transformation toward smart manufacturing, autonomous systems, and other AI-driven applications.

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • Founded in 2021 by the core team of the National Supercomputing Center in Wuxi, with founder Yan Bowen holding a postdoctoral degree from Tsinghua.
  • A unified heterogeneous computing pool supports NVIDIA GPUs and various domestic AI chips, turning idle resources into usable compute.
  • Inference optimization boosts throughput by 30%-50% and cuts per-token cost by 40%.
  • A multi-provider disaster recovery system ensures 99.9% availability for stable token production.

Highlights and analysis are generated automatically and may contain errors. Check the original source.