Organizations need to own their intelligence to be successful. Training and serving that intelligence at scale, however, requires petaFLOP/s of compute and terabit/s of networking, spread across many nodes. But owning intelligence doesn’t need to mean owning that hardware.
For the past 1.5 years, we’ve been battle-testing a new primitive: Modal Clusters. Today, we’re excited to announce that they are generally available through a single decorator, @modal.clustered:
@app.function(gpu="B300:8") @modal.clustered(size=4, rdma=True) def train_model(): cluster = modal.Cluster.from_context() container_ips = cluster.private_ips() container_rank = cluster.container_rank()
world_size = len(container_ips) main_addr = container_ips[0]
print(f"{container_rank=} {world_size=} {main_addr=}") ...