Vault 资讯瀑布媒体2026.07.16 21:30 UTC+8

🔬 未来的实验室应像数据中心——Lila Sciences的Andy Beam与Rafa Gómez-Bombarelli

Lila Sciences 访谈:将实验室视为数据中心,AI引导自动化实验,生成超10万亿科学推理token,旨在科学超级智能。

想象一个黑暗的仓库。一排排的设备,电线、管子和电子元件外露。下一个AI数据中心?不。这是Lila Sciences对科学未来的梦想。一个充满AI引导机器人和实验室设备的黑暗仓库,全天候运转新的实验,朝着科学超级智能迈进。

他们的自动化实验室观看时几乎令人着迷。他们让漂浮的板子在类似瓦力的轨道上飞速移动,使用视觉语言模型控制Windows 95盒子,并创造了世界上最多失效保修单的收藏。在这个过程中,他们构建了庞大的科学推理token库。超过10万亿个,全部经过实验验证。

本视频制作过程中没有保修失效。

说Lila雄心勃勃是一种轻描淡写。他们的目标是将科学超级智能直接接入湿实验室。他们完全押注于“苦涩的教训”,其论点由此而来:实验室是一个无限的token生成器。大规模产生数据,协同效应会给你一个能解决任何科学问题的通用推理器。他们在全力投入。生物学、化学、药物发现和材料科学,同时进行。时间会证明它是否有效,但这是一个令人兴奋的假说。

在我们最新一期节目中,我们与Lila的Andy Beam(首席技术官)和Rafa Gómez-Bombarelli(首席科学官,物理科学)坐在一起,踏上了一段探索AI驱动科学可能性的旅程,其范围几乎与Lila的目标一样广泛。

我们有没有提到他们同时做材料科学和生物学?在同一个AI科学工厂里?同一时间,同一个实验室,同一个AI。终于有一位嘉宾能解决我们之间长期争论的问题:生物学和材料科学哪个更难?

观看以了解!

我们讨论: - 互联网已经耗尽,科学是下一个。为什么Lila认为科学方法是最后一个未开发的互联网规模数据集,以及为什么他们将RL视为以自然为验证者的数据生成机制。 - 实验室作为数据中心。仪器作为图上的节点,它们之间有一个磁悬浮的“PCI总线”传输层,编排作为slurm队列。Andy不缺乏类比。 - 为什么Lila坚持自己不是一家自动化公司。他们优化灵活性和通用性而非原始吞吐量,这意味着在自动化不划算的地方,人类保持在API线之下。 - 你的实验有一个运行时间。我们将Escalante Bio的问题抛给Andy:如果科学是token生成器,你的数据收集运行时是什么?他的回答,简而言之,是你无法让核糖体更快。为什么Lila押注于快速的逐轮迭代而不是大型嘈杂的多重筛选,以及Rafa的团队如何重新构建了一个气体吸附测量,使其运行速度快约2500倍。

🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences

Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward a scientific superintelligence.

Their automated lab is almost hypnotizing to watch. They have floating plates zipping around on Wall-E-esque tracks, used vision-language models to control Windows 95 boxes, and created the world’s largest collection of voided warranties. In the process they’ve built a massive library of scientific reasoning tokens. Over 10 trillion of them, all experimentally validated.

No warranties were voided in the making of this video

To say Lila is ambitious is an understatement. Their goal is a scientific superintelligence wired directly into the wet lab. They are all in on the bitter lesson, and the thesis follows from it: a lab is an infinite token generator. Produce data at scale, and the synergies give you a general reasoner that can tackle any scientific problem. They are committing hard. Biology, chemistry, drug discovery, and materials science, all at the same time. Time will tell if it works, but it is an exciting hypothesis.

In our latest episode we sat down with Lila’s very own Andy Beam (CTO) and Rafa Gómez-Bombarelli (CSO, physical sciences) and went on a journey through the possibilities of AI-run science, almost as wide-ranging as Lila’s goals.

Did we mention they do both materials science and biology? In the same AI science factory? Same time, same lab, same AI. Finally a guest who can settle a long-running debate we’ve had amongst ourselves: is biology or materials science harder?

Watch to find out!

We discuss:

- The internet is spent, science is next. Why Lila thinks the scientific method is the last untapped internet-scale dataset, and why they treat RL as a data generation mechanism with nature as the verifier.

- The lab as a data center. Instruments as nodes on a graph, a magnetically levitating “PCI bus” transport layer between them, orchestration as a slurm queue. Andy is not short on analogies.

- Why Lila insists it is not an automation company. They optimize for flexibility and generalizability over raw throughput, which means humans stay below the API line wherever automating does not pay.

- Your experiment has a runtime. We put Escalante Bio’s question to Andy: if science is the token generator, what is the runtime of your data collection? His answer, in short, is that you cannot make the ribosome go faster. Why Lila bets on fast round-over-round iteration rather than big noisy multiplexed screens, and how Rafa’s team rebuilt a gas sorption measurement to run roughly 2,500x faster.

- What is actually in 10 trillion scientific tokens. Not sequences. Experimentally verified reasoning traces, a kind of data that Andy argues exists on the internet in quantities that round to zero.

- Breadth as a path to depth. Small molecule chemistry priors transferring to metal organic frameworks for carbon capture, and the claim that the general model beats domain-specific models sample for sample.

- If you have the data, what do you need the model for? Sri Kosuri’s koan about the ML-for-drug-discovery business model, and Andy’s answer: the coding model got better because it also read Shakespeare and carnitas recipes.

- The serendipity they want to automate. Emily Whitehead survived the first pediatric CAR-T cure only because the doctor treating her happened to know, from pediatric arthritis, which antibody would blunt her IL-6 response. Roll that dice again and you probably lose her. Breadth is how you stop depending on luck.

- Move 37 for catalysts. Model suggestions for platinum-group-free electrocatalysts that went from boring, to what a 40-paper expert called stupid, to the best performers they have made.

- Six months to in vivo CAR-T data in non-human primates, and the zero-FTE virtual startup commercial model that fell out of it. For context on why that number is startling, AbbVie paid $2.1B for Capstan on the strength of preclinical in vivo CAR-T data.

- You cannot have scientific superintelligence if you are just a good test taker. Ken Stanley, who wrote Why Greatness Cannot Be Planned, runs open-endedness at Lila. RL at scale gives you a ruthlessly Vulcan problem solver. Machine creativity is a different thing, and it is the part nobody has solved.

- The chain of thought is an unreliable narrator. The model reasons in latent space and only emits tokens. Sometimes it skips the experiment entirely and is still right. So how much do you trust the reasoning versus the verifier?

- Reward hacking when the rollout is physical. Chains of thought that collapse into repetition, and a model that got annoyed and swore at the scientist who kept asking it to redo a plate map. What happens when a pathological loop has a wet lab inside it?

- The bittersweet lesson. Rafa’s inversion of the bitter lesson: in AI, scaling is a roadmap. In materials, scaling is a filter, because only the things that scale end up mattering.

- Not your typical Flagship company. Why a famously single-asset biotech incubator spun out a platform bet, and Andy’s line that if Lila called itself a biopharma it would have a top-three GPU cluster.

- Bottlenecks they would remove by fiat. Sim-to-real for physics-based simulation, and the fact that RL training runs at roughly 5% mean FLOP utilization.

Watch on YouTube:

查看原始发布