AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Organizations need to own their intelligence to be successful. Training and serving that intelligence at scale, however, requires petaFLOP/s of compute and terabit/s of networking, spread across many nodes. But owning intelligence doesn’t need to mean owning that hardware. For the past 1.5 years, we’ve been battle-testing a new primitive: Modal Clusters. Today, we’re excited to announce that they are generally available through a single decorator, @modal.clustered: @app.function(gpu="B300:8") @modal.clustered(size=4, rdma=True) def train_model(): cluster = modal.Cluster.from_context() container_ips = cluster.private_ips() container_rank = cluster.container_rank() world_size = len(container_ips) main_addr = container_ips[0] print(f"{container_rank=} {world_size=} {main_addr=}") ...