跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Runway’s WorldPrompt and the Engineering of Real-Time Worlds

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:GWM Worlds 2 uses persistent context and timed actions to steer a world model generating video and audio in real time.

來源Latent Space作者: Richard MacManus
待翻譯:Runway’s WorldPrompt and the Engineering of Real-Time Worlds
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time. One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time. To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him. Who’s building real-time interactive world models? First, some context about world models that can generate interactive video and audio in real-time. Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December. Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table: Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.” But as our interviews with Runway show, real progress is being made. The central idea of WorldPrompt WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game. “Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.” As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out. “You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said. But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character? “Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.” Sindi added that more training plus scaling the data and models is resulting in “better following.” How a video model becomes a real-time runtime Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us. “There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.” High-level view of GWM Worlds 2 process GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz. Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively. “And after that, we work on making it real-time through distillation methods,” Kahlow added. Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.” Autoregressive causal diffusion vs traditional video models; image via Runway Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results. The challenges of real-time generation Germanidis admitted that there were issues with how it generates real-time interactive video. “The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.” Errors compound; image via Runway Sindi told us there are also challenges dealing with “infinite generations” of content. “There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.” Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.” Causality and correctness While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions. Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly. “If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.” Image via Runway Sindi told us that evaluation gets harder the more complex interactions get. “If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?” To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.” More than gaming — there are agent use cases too Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works. Another, more intriguing, use case is to use it to test agents at scale. “Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said. But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding? “So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.” Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.” Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models. “You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.” Anastasis Germanidis LinkedIn: https://www.linkedin.com/in/agermanidis/ X: https://x.com/agermanidis Timestamps 00:00:00 Introduction 00:05:17 Runway’s Origins and the Bet on Generative Video 00:12:23 The Stable Diffusion Story 00:18:44 Gen-2, Controllability, and the Weekend Hack 00:23:02 From Video Generation to World Models 00:28:03 Learning From the World, Not Just Language 00:35:04 Sora, Runway’s Existential Crisis, and Gen-3 00:39:39 Why Real-Time Video Is Inevitable 00:43:06 Interface World Models: Software Without Code 00:50:25 The Fully Neural Operating System 00:55:11 World Models for Robotics 01:02:32 Robot Policies and World Action Models 01:07:47 The Lucid Dream Test 01:11:41 Video Agents and Omni Models 01:23:12 Artists, AI, and Creative Workflows 01:27:14 Physical AI and the Future of World Models Transcript Introduction: Runway, Creative AI, and the Early Thesis Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome. Anastasis [00:00:08]: Good to be here. Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out? Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of those generative models, rethink how creative tools are made. and as we built out the research behind, our generative models, it then became clear that they were useful far beyond that as well. Anastasis’ Background: Art, Simulation, and Machine Learning Swyx [00:00:57]: And it is more obvious now with, like, the real-world stuff and the world models that we’ll talk about later. I’m just kinda curious how you go from a background in, like, Zocdoc and, computer vision into Runway. Like, take us back to that early conversations with Chris and, whoever else is on your founding team. Anastasis [00:01:14]: I was always splitting through those two worlds. One was the I had my own art practice. I was making a lot of interactive art, I think for a long time. and then on the other side, I was working in startups, and I was working as a ML engineer, as a backend engineer at different companies. I’ve always been interested in, coding and computation, and especially interested in simulation and brought it back into my early artwork as well. And at the same time, I was interested in Swyx [00:01:43]: The personal site has a few, right? Anastasis [00:01:44]: Yeah. Swyx [00:01:45]: Is there one that we should pull up? Just in case there’s something that’s like. I just like to go down memory lane. Anastasis [00:01:50]: Yeah. Swyx [00:01:50]: Okay, what is this? Anastasis [00:01:51]: So this was, a project that I made, I think back in 2015, where I built this software that would give, voice instructions to people in a gallery space. So it would coordinate interactions between people. And so it will first give you an identity, like you’re an, architect, you’re 30 years old, and, you like sports. and then it would match you with another person, and you have this completely generated interaction. language models were not quite there at the time, and so it was it was a mix of some templates and some, like, some Ma [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • GWM Worlds 2 uses persistent context and timed actions to steer a world model generating video and audio in real time.

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。