本文にスキップ
AI News HubLIVE
原典の内容 · 翻訳・分析待ち6 分で読了

翻訳待ち:Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology.

ソースLatent Space
翻訳待ち:Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology. In this special panel moderated by Brandon Anderson, they explore why AlphaFold’s breakthrough was only the beginning, why scaling compute and data alone won’t solve biology, and how the next generation of AI models could transform our understanding of proteins, cells, and human disease. We go deep on the future of AI-driven biology: finding scaling laws in biological data, the tradeoffs between scientific intuition and general-purpose architectures, why protein structure prediction is far from solved, and what it would take to build predictive models of living systems. Pushmeet reflects on the lessons behind AlphaFold, the limits of human interpretability, and whether future frontier models could understand other AI systems better than we can. Sal explains why protein language models may already contain scientific knowledge we haven’t unlocked, how biological modeling must move beyond individual proteins, and why achieving Biohub’s mission to cure all disease requires thinking in terms of 10x breakthroughs rather than incremental improvements. We discuss: The Bitter Lesson for biology: why scaling compute and data isn’t enough Why finding the right scaling law matters more than blindly increasing model size How low-quality metagenomic data can improve protein language models Why AI researchers optimize for available data instead of the most important scientific problems Lessons from DeepMind on balancing modeling, data generation, and scientific expertise Why building a virtual cell requires fundamentally different datasets AlphaFold’s handcrafted architecture and the role of scientific intuition Why good data matters more than simply having more data Inductive biases, scaling laws, and the future of specialized AI architectures Why we aren’t in a post-Transformer world, but architectures are evolving Why AlphaFold didn’t actually solve all of protein folding Protein dynamics, disorder, and the limitations of static structure prediction How cryo-EM micrographs could unlock richer biological representations Moving from models of individual proteins to whole biological systems Feynman’s famous principle and why AI can now create things we don’t understand The hidden biological knowledge inside protein language models Why trustworthiness and uncertainty calibration matter more than full interpretability Whether frontier AI models could interpret other AI systems better than humans When AI could deliver 10x–100x acceleration in drug discovery Why curing all disease requires thinking about 10x breakthroughs instead of 10% improvements Pushmeet Kohli — Google DeepMind X: https://x.com/pushmeet LinkedIn: https://www.linkedin.com/in/pushmeet-kohli-4838994/ Sal Candido — Biohub Biohub: https://biohub.org/team/salvatore-candido/ X: https://x.com/salcandido LinkedIn: https://www.linkedin.com/in/salcandido/ Brandon Anderson — Moderator LinkedIn: https://www.linkedin.com/in/brandon--anderson Timestamps 00:00:00 Introduction: The Bitter Lesson for Biological Data 00:01:00 Finding Scaling Laws and the Right Data for Biology 00:04:54 DeepMind’s Bitter Lesson: Solving Problems vs. Scaling Models 00:06:50 AlphaFold, Data Limitations, and Building the Virtual Cell 00:09:52 Handcrafted Architectures vs. Scaling Compute 00:11:53 Good Data, Inductive Bias, and Model Design 00:13:39 Beyond Transformers: The Future of AI Architectures 00:14:38 Why AlphaFold Hasn’t Solved Protein Folding 00:17:39 Protein Dynamics, Design, and Cryo-EM 00:19:09 From Individual Proteins to Whole Biological Systems 00:20:43 Feynman’s Principle: Creating Without Understanding 00:22:18 The Hidden Knowledge Inside Protein Language Models 00:23:48 AlphaFold, Trustworthiness, and Interpretability 00:25:56 Could AI Understand Other AI Models Better Than Humans? 00:27:43 When Will AI Revolutionize Drug Discovery? 00:29:47 Why Curing All Disease Requires 10x Breakthroughs Transcript Introduction: Is There a Bitter Lesson for Data? Brandon Anderson [00:00:04]: Great to be here. What an exciting morning. So many cool announcements. I think the future of bioscience is being announced right here. This is the modeling session. We’re all modelers, so of course it’s natural for us to talk about data. One of the things I like to think about when it comes to data is how to scale it properly. This has brought me to the question of the bitter lesson, but recast in the frame of data. The bitter lesson, for those of you who are not AI people, is the statement that methods that scale win eventually. If you can scale enough, it wins. So my question, starting with Sal, is: Is there a bitter lesson for data? Sal Candido [00:01:00]: For sure. You obviously need the right data in order for it to work. One misconception of scaling laws is that scaling laws are everywhere and they always exist. A lot of the work is actually finding that scaling law. A lot of what we do is trying to figure out: What’s a situation where, if you put more compute into it, if you put more data into it, you get a better result out? That’s a great situation because, once that happens, you can turn the crank. It becomes an engineering problem, which is something I like. That has to do with architecture, but it also very much has to do with data. If you don’t have data with the right information and statistics to solve the problem you want, you’re not going to get a model that has the capabilities and understanding you want. You can only really pull information from the data you have and use that to generalize beyond it. Choosing the Right Data: Availability vs. Scientific Impact Brandon Anderson [00:02:09]: When it comes to data collection, how do you think about which modalities are best? A related question: Do modelers tend to work on problems where the data is available, rather than the problems that best serve their goals or have the greatest impact on translational medicine? Sal Candido [00:02:32]: For sure. I won’t speak for all modelers, but I’m lazy, so I’m going to work with what’s available. There’s a good side and a bad side to this. The good side is that, when you’re doing conventional machine learning, you’re often looking for the most pristine, high-quality data examples you can find. But if you’re like me, you’re rooting around in the back room, looking through people’s junk to see what’s there. A concrete example is training a protein language model on metagenomic sequences, which are not the highest-quality data. In fact, much of that data, I can guarantee you, isn’t even a real, whole protein. And yet it makes the performance of the model go up for designing real proteins that work and understanding proteins that we know exist. That’s the positive side. But it can lead you to a negative place where you say, “Let’s just scale up the data that we can generate easily.” I don’t think that’s necessarily the way to do it. That’s one of the things that’s exciting to me about what we’re talking about here today with the BBI. It’s going out and asking, “What is the data that we need to solve the problem?” It’s about the right resources, but also the right community. One thing we try to do at Biohub is work in the open, work with the community, and move the whole community forward. That’s critical because, if you just have people building models, we’re going to lean toward the data that exists. If you just have people generating data, they’re going to lean toward the things that can be generated. But if you can work together as a community, every step of the way, in an open fashion, you can figure out what data you actually need. Then you can find that scaling law. The Problem Comes First: Modeling, Data, and Expertise Brandon Anderson [00:04:50]: Pushmeet, what do you think about the bitter lesson for data? Pushmeet Kohli [00:04:54]: I was at DeepMind when Rich was with us and thinking about this idea of the bitter lesson. What I took from Rich’s original lecture at DeepMind wasn’t about the specific notion of whether data is useful or how we should think about machine learning. My take was that he was talking about something more conceptual. Sometimes when we’re looking at problems, we think about solutions in a very religious way: “I’m a modeler,” or “I’m a data generation person.” I think that is the bitter lesson. If you approach a problem with that mindset, you might not succeed. The problem comes first, and you should be flexible in your solution space. You should try both things. You should understand the problem you’re trying to solve. If it requires modeling effort, do the modeling effort. If it requires collecting more data, then collect data. What happened in machine learning at the time was people saying, “We’re machine learning researchers. The dataset is there. Here’s some training data, here’s some test data, and we’ll just optimize the model.” That is broken in the sense that, if your eventual goal is to solve the problem, you have to look at both aspects of what goes into the process. And what goes into the process is not just data or modeling. It’s also expertise. With AlphaFold, for example, we looked at what was possible with existing datasets because we didn’t have the core expertise, or even the resources, to say, “Let’s augment the PDB by a significant order of magnitude.” The investment that organizations and scientists across the world had put into constructing that data was invaluable. So there, you had to focus on getting the biggest bang for your buck by investing in modeling. But in other areas, say cell genomics, we took the same approach and asked, “What can you do with cell-by-gene data?” After a lot of work, it was very clear that the data was not there yet to pursue that grand ambition of building the virtual cell. This brings people together to focus on the actual challenges of advancing science rather than religiously following advances in data generation or modeling. Brandon Anderson [00:08:23]: I really like that answer. What’s the actionable takeaway? Always define the problem first and then figure out which solution space to search over. But more broadly, how should the community think about this as we move toward the next generation of translational medicine? Pushmeet Kohli [00:08:47]: The advice I give to anyone starting in this area is to think of yourself as a multidisciplinary person. Understand the problem first. Why are you working on it? What are you trying to achieve? Then think about what’s needed, whether that’s modeling, compute scaling, or data generation. Understanding the problem is extremely important, and of course you need to build your expertise. There are constraints, too. Maybe there’s only a certain amount of data you can generate, or a certain model size you can afford to train. Understand those constraints, try to fail fast, and look at which approaches will be feasible in the long term to get you to the level of impact you’re aiming for. Scientific Inductive Bias vs. Scaling: The Craft of Building Models Brandon Anderson [00:09:52]: This leads right into my next question. When I look at the evolution of modeling, AlphaFold 2 was essentially a work of art, with a lot of carefully handcrafted features. There was very intentional thought put into every part of the solution. Some of Google’s or Alphabet’s work still stays in that space, but the general consensus seems to be moving toward more scalable, general strategies. Do we still need artisanal, craft solutions for certain problems? Or are resources better spent focusing on scale first? With fixed resources and money, you can invest in compute, talent, or data. How should we think about that trade-off? Pushmeet Kohli [00:10:48]: [truncated for AI cost control]

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。