为超级智能承保:为可被起诉的智能体提供保障——AIUC 的 Rune Kvist
Latent Space 播客对话 AIUC 联合创始人 Rune Kvist,宣布 4000 万美元 A 轮融资,并探讨 AIUC-1 智能体标准与 AI 保险基础设施。
中文处理结果
AIUC 最初因 NFDG 的投资而进入我们的视野,今天又宣布完成 4000 万美元 A 轮融资,其背后是 AIUC-1——他们由真实保险支持的智能体标准——所拥有的、可能是我们见过的早期创业公司中最令人印象深刻的行业顾问名单。
从 Anthropic 的第一位产品招聘,到构建旨在让前沿 AI 可部署的标准、测试与保险基础设施,Rune Kvist 押注的是:AI 采用的最大制约不会是能力,而是信任。在本期节目中,AIUC 联合创始人加入 swyx 和 Vibhu,宣布新一轮 4000 万美元融资,并解释为什么 Cursor、Harvey、Lovable 和 ElevenLabs 等公司正日益面对一个随着 AI 变强而变得更难的问题:当自主系统失败时,谁来负责?
我们深入探讨 AIUC-1——这一正在成形的智能体安全、保障与可靠性标准;AI 智能体如何针对越狱、幻觉和数据泄露进行压力测试;以及为什么 Rune 认为标准和保险可能成为 AI 的关键基础设施。我们还讨论了政府与前沿实验室之间日益扩大的信任鸿沟、AI 赋能的网络与生物风险、为什么每个模型最终都可能被越狱、当一个 20 美元的编程智能体造成 2 亿美元损失时会发生什么、AI 工程师是否应该获得认证,以及为什么即便在 AGI 之后,仍可能有一项工作是实验室永远无法自己完成的:做自己的监督者。
我们讨论:
- 为什么风险、责任与信任可能成为 AI 采用的约束性瓶颈
- Rune 从阅读 Scaling Laws 论文到在 Anthropic 最早时期加入的历程
- Anthropic 在多年前就理解到的关于扩展、算力与未来的东西,而这些在后来才变得显而易见
- 为什么 Waymo 体现了 AI 能力与真实世界部署之间的差距
- AIUC 的 4000 万美元融资,以及与 Cursor、Harvey、Lovable、ElevenLabs 和其他前沿 AI 公司的合作
- AIUC-1:AI 智能体安全、保障与可靠性标准
- 智能体如何接受越狱、幻觉和数据泄露测试
- 为什么大多数 AI 公司只优化顺利路径,而不认真对对抗性案例做压力测试
- 为什么 AI 标准可能需要每季度更新,而不是每十年更新
- 前沿 AI 实验室与政府之间正在出现的信任鸿沟
- 网络安全、儿童安全、生物武器,以及不断扩大的前沿模型风险面
- 为什么标准与保险可能需要共同演进
- 伦敦劳合社如何为 AI 系统承保,并为企业部署带来信任
- 如果一份 20 美元的 Cursor 订阅促成了一场 2 亿美元的飞机坠毁,会发生什么
- 加拿大航空聊天机器人案,以及 AI 失败如何开始厘清法律责任
- 为什么版权可能是最难承保的 AI 风险之一
- 评估、机制可解释性、监控,以及模型开始意识到自己正在被测试
- 不可能完成的 CISO 任务:快速采用 AI,但不要让任何事出错
- 为什么机器人技术会让 AI 责任变得影响重大得多
原始正文
Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:
From being Anthropic’s first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won’t be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?
We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.
We discuss:
- Why risk, liability, and trust may become the binding constraint on AI adoption
- Rune’s path from reading the Scaling Laws paper to joining Anthropic in its earliest days
- What Anthropic understood about scaling, compute, and the future years before it became obvious
- Why Waymo illustrates the gap between AI capability and real-world deployment
- AIUC’s $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies
- AIUC-1: a standard for AI agent security, safety, and reliability
- How agents are tested for jailbreaks, hallucinations, and data leakage
- Why most AI companies optimize the happy path without seriously stress-testing adversarial cases
- Why AI standards may need to update every quarter instead of every decade
- The emerging trust gap between frontier AI labs and governments
- Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface
- Why standards and insurance may need to evolve together
- How Lloyd’s of London can insure AI systems and bring trust to enterprise deployment
- What happens if a $20 Cursor subscription contributes to a $200M plane crash
- The Air Canada chatbot case and how AI failures are beginning to clarify legal liability
- Why copyright may be one of the hardest AI risks to insure
- Evals, mechanistic interpretability, monitoring, and models becoming aware they’re being tested
- The impossible CISO mandate: adopt AI fast, but don’t let anything go wrong
- Why robotics will make AI liability dramatically more consequential
- Whether AI engineers should have Level 1, 2, and 3 certifications
- AIUC’s roadmap across agents, frontier models, robotics, and universal red teaming
- Why AGI could become a question of national sovereignty
- Why the labs can never fully serve as their own watchdogs
- The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?
Rune Kvist
- LinkedIn: https://www.linkedin.com/in/runekvist/
- X: https://x.com/RuneKvist
AIUC
- https://aiuc.com
Timestamps
00:00:00 AIUC’s $40M Round and the Risk Bottleneck for AI
00:01:07 From Scaling Laws to Early Anthropic
00:07:58 Why Trust, Not Capability, Could Limit AI Adoption
00:12:19 Founding AIUC and Building AIUC-1
00:18:52 How AI Agents Are Audited and Stress-Tested
00:25:26 Frontier Models, Government, and the AI Trust Gap
00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk
00:38:14 Why Standards and Insurance Belong Together
00:41:45 What Does an AI Insurance Policy Actually Cover?
00:50:44 The $20 Cursor Subscription and the $200M Plane Crash
00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust
00:56:21 From AI Agents to Models to Robotics
00:58:29 Copyright, Adverse Selection, and AI Insurance
01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness
01:08:36 The Impossible Enterprise AI Mandate
01:11:52 Prediction Markets vs. AI Audits
01:14:43 Should AI Engineers Be Certified?
01:19:10 AIUC’s Roadmap, AGI, and Who Watches the Watchdogs?
Transcript
Introduction: AIUC, the $40M Series A, and Risk as the Adoption Bottleneck
Swyx [00:00:00]: Okay, we’re in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.
Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.
Swyx [00:00:11]: What are you announcing today?
Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.
Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?
Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It’s pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.
Swyx [00:00:54]: And let’s get a list of the customers that you’re highlighting as part of your Series A.
Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.
Swyx [00:01:05]: Yeah. Amazing. Congrats.
Rune Kvist [00:01:06]: Thank you.
Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I’m just kind of curious: what was your path into AI? Just recap.
Rune’s Path Into AI: Scaling Laws, Capital, and Anthropic
Rune Kvist [00:01:18]: Yeah.
Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.
Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?
Rune Kvist [00:01:40]: Exactly, the Kaplan one.
Swyx [00:01:42]: Yeah.
Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what’s played out. And so I just packed my bags. I’d never been to San Francisco. I’d never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They’d just broken off from OpenAI, and it’s been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there’s nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.
Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.
Rune Kvist [00:02:52]: Yes.
Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy shit.”
Rune Kvist [00:03:19]: Correct.
Swyx [00:03:20]: Right?
Rune Kvist [00:03:21]: Yes.
Swyx [00:03:21]: Who tipped you onto that paper? Because it’s not a paper that you normally read, right, like, in your circles?
Rune Kvist [00:03:26]: Yeah. I think I’d actually, ever since AlphaGo, had some appreciation that AI was a big deal.
Swyx [00:03:36]: Yeah.
Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn’t as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you’re going to be wrestling with all of the big questions in society. Everything you’ve learned about politics gets thrown out of the window. Everything you’ve learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that’s why I sought it out.
Swyx [00:04:32]: I mean, clearly really good insight. For people who don’t know, the PPE program is, like, where prime ministers are born. So then you end up meeting Dario.
Rune Kvist [00:04:41]: Yep. First Dario, yeah.
Swyx [00:04:43]: Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn’t get from your original hypothesis?
Anthropic’s Early Conviction and the Scaling Laws Crystal Ball
Rune Kvist [00:04:50]: If you read the Scaling Laws paper, you get this, like, very vague sketch of like, wow, this seems kind of important. There are some lines on a chart. This seems kind of important. And what I think the team at Anthropic had thought more about than anyone was like, what are the implications of this if you really play this out? And back then they had, kind of vision documents for what the world would look like in 2026, and they were kind of in vivid detail playing out how much compute is going to be needed, what the CapEx was going to look like, what some of the societal concerns were going to be, but also what is the amount of economic value coming out here? And so it kind of felt like they held a crystal ball that in hindsight turned out to just be dramatically correct. And they weren’t holding it like they were obviously correct. They were just like, “Take this hypothesis really seriously.”
Swyx [00:05:38]: Think it through, yeah.
Rune Kvist [00:05:38]: And think it through in the same way as the kind of situational awareness that is
Swyx [00:05:43]: Across the street.
Rune Kvist [00:05:44]: Across the street.
Swyx [00:05:44]: Your office, yeah. Oh my God, we’re all living across the street in the same one square mile.
Rune Kvist [00:05:50]: Correct. And that’s now a couple of years old, but also people keep referencing it these particular weeks with Fable and Mythos, and it’s like, wow, if you take this one idea seriously- For the Scaling Laws, a lot of things fall into place.
Vibhu [00:06:03]: And keep in mind, at this point, this is the same team that did GPT-1, GPT-2, and GPT-3.
Rune Kvist [00:06:08]: Correct.
Vibhu [00:06:08]: Which is also, like, it’s not just some experimentation. Like, this is a real model that we just scaled up.
Rune Kvist [00:06:14]: And they had deep conviction in this idea: if you take a big blob of compute and data, it just wants to learn, and out of that will come smarter and smarter models. And all the particulars were not clear.
Vibhu [00:06:26]: Yeah.
Rune Kvist [00:06:27]: And all the implications were not clear. But their deep conviction in this, like, core thesis, and that was kind of dizzying. It was both phenomenally interesting and exciting, and also very quickly you get to, like, the world we know today will no longer be if this hypothesis holds. So it also just felt, like, important in some kind of grand sense.
Vibhu [00:06:48]: What kind of shaped you there? So that was early 2022. Not only had GPT-1, GPT-2, and GPT-3 come out, but, you know, the amazing founders of Anthropic that have never split up, the only ones, they actually had the conviction to leave OpenAI, start their lab. You said there were about 40 people there. What was the time like there?
Inside Early Anthropic: Mission, Deployment, and Risk
Rune Kvist [00:07:06]: It was kind of remarkably like what it looks like on the outside today. Extremely cohesive, extremely mission-oriented, and living in this tension between their two ideas, which is AI could both go really well and really bad, and we want to be part of building it. That creates astounding amounts of tension. And they were wrestling with this incentive challenge where they know they’re in a race that they’re in where you might get forced to cut corners, but it also felt very important to them to be at the forefront of technology. And all of those ideas were just present at that time. It kind of feels like that line has been just very clear, and I think kind of love them or hate them, they have really stuck to their guns. There’s a core set of beliefs that they hold more deeply than most companies hold any beliefs.
Vibhu [00:07:58]: Yeah. Fast-forward to today.
Rune Kvist [00:08:00]: Yeah.
Vibhu [00:08:00]: What does that lead us to AI underwriting company? What are you up to? What motivated you to start this?
From Waymo to AIUC: Confidence Infrastructure for AI
Rune Kvist [00:08:05]: Yeah. AIUC builds confidence infrastructure for frontier AI through standards and insurance. The link from Anthropic to building confidence infrastructure, looking out the windows at Anthropic offices and seeing Waymos driving by. Already back then, early 2022, Waymos were in some ways like AGI for cars. Like, they were superhuman drivers, but you couldn’t take one to the airport. And now, four and a bit years later, you still can’t take your Waymo to the airport, despite now everyone having kind of looked at the evidence and being like, “They’re better drivers than humans.” So in that particular instance, what’s clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust. That problem is, general. The reason why right now
Rune Kvist [00:08:52]: Fable is not open for access is not because it’s not a good model, it’s because it’s a very good model. It’s just hard to make promises about what it will or will not do. And this problem gets worse as AI gets better. Basically, more intelligent AI can be more autonomous. That’s more valuable, but also the risk surface grows. And so - what Waymo illustrates is that unless you build the confidence infrastructure to make promises about AI, or at least bring light to the risks, you grind adoption to a halt. Governments, banks, hospitals, militaries need to have some sense of what AI will and will not do to be able to operate for them to incorporate it. And that’s the problem that we’re trying to solve. Now, why standards and insurance? If you trace this problem back through history, every technology wave has had some version of this problem. So if you go back to, like, year 1900, electricity comes
Vibhu [00:09:47]: Ben Franklin.
Rune Kvist [00:09:48]: Cars burn down, sorry, houses burn down, lots of people die. 1930s, cars are a big deal, kill lots of people. 1950s, private nuclear energy is a big deal, poses big risks. In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that’s required to make go/go decisions. That is required for adoption, and the market fundamentally wants adoption. And in all of those instances, common blueprint emerges between standards and insurance. The reason these two components is standards kind of provide the rules of the road, and they also specify, like, what are the tests that need to be run so we can get a sense of how high the risk is. So take in the case of cars, that’s like a car crash. Great, everyone, they inform your insurance pricing today, they inform your purchasing decisions, et cetera. That’s basically the risk framework. The insurers are important because they pick up the bill. So they are the private institution that is most on the side of. That is best incentivized to quantify the risks truthfully and then figure out all the ways to reduce the risk ‘cause that increases their profit. So they’re basically, they help shape the incentives. And these two work really well in unison. Now, how does that show up as a company? Well, one of the things that was obvious even - or starting to become obvious even a couple years ago was that frontier companies, some of our customers today, like Cursor, Sierra, ElevenLabs, Harvey, were going to have a very easy time selling a pilot to a bank. The, like, the demo just sells itself. It’s magic. But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital, you have to go through the risk process. These banks have no idea even which questions to ask, let alone which answers are sufficient, let alone, like, how do they go and test whether these agents actually work the way they’re supposed to. And so they had this problem of, like, what can we say to earn the trust? And we think there’s, like, a golden sentence that goes something like, “Hey, I hear you’re really worried about hallucinations or jailbreaks or whatever it may be. We’ve had an independent third party test us against the gold standard. We passed with flying colors. And as a vote of confidence, the world’s most conservative insurers have looked at the data.” And they’re willing to take some of the risk onto their balance sheet.
Swyx [00:12:06]: Yeah.
Rune Kvist [00:12:07]: So if something does go wrong
Swyx [00:12:07]: There’s money behind it, yeah.
Rune Kvist [00:12:09]: Exactly. So that’s kind of like the link between all this. We can get into some of the hard parts related to the technical testing, which is, I think, the crux of the matter, but I’ll pause there.
Swyx [00:12:19]: How did you and Rajiv come together? This-- there’s always, like, you come across very confident and, you know, and we’re announcing your Series A and all these things, but I want to see, like, the early initial stages of, like, idea formation.
Cofounding AIUC with Rajiv Dattani
Rune Kvist [00:12:31]: Yeah. Rajiv is actually my soon-to-be brother-in-law.
Swyx [00:12:35]: Oh.
Rune Kvist [00:12:36]: So I’m actually, in a week and a half getting married to Rajiv’s sister.
Swyx [00:12:42]: Okay, now you’re tight.
Rune Kvist [00:12:44]: Exactly.
Swyx [00:12:44]: Now you know.
Rune Kvist [00:12:45]: So - Rajiv and I have known each other for a decade. Funny story, I met both Rajiv and his sister, Hena, at the same time when Hena and I were interns at McKinsey in London, and Rajiv was assigned as my mentor. And so met them at the same time. For the longest time, it was not obvious that we were necessarily going to work together. I was in startups. He was, an insurance partner at McKinsey. Three or four years ago, I think Hena convinced him that AI was going to be a really big thing. And so he quit his job, cushy partner job at McKinsey in London, packed his bags, flew to San Francisco, and ended up joining METR. You guys are probably online enough
Swyx [00:13:24]: CEO.
Rune Kvist [00:13:24]: Exactly.
Swyx [00:13:24]: We’ve, we’ve, we’ve heard of METR.
Rune Kvist [00:13:25]: You see the plot-- the chart of the horizons of the tasks that agents can take on is doubling extremely fast. So he was COO at METR, led their partnerships with Anthropic and OpenAI to test their models before release, but also working closely with the US and UK government, to figure out, like, how do you know whether a model can be released? And in some ways, that was, like, the perfect background. He’s spent a lot of time in insurance, knows that world, spent a lot of time with frontier testing of models. And so when I was bumbling around this idea space, starting with some of the ideas we talked about related to Waymo, as soon as we got into the content, we were both like, “Oh, this would be an amazing business to build together.” This is wrestling with the problem that we both think is the most important in the world from a market angle, which is kind of our intuitions is that the market can do a lot, and the faster AI moves, the harder it is for government to solve some of these problems. And then it took a little bit of time to work through what is it like to work with family.
Swyx [00:14:27]: Sure.
Rune Kvist [00:14:27]: And,
Swyx [00:14:30]: Because you were already dating at the time
Rune Kvist [00:14:31]: Yeah. Yeah, exactly.
Swyx [00:14:33]: Yeah.
Rune Kvist [00:14:34]: Already back then, it
Swyx [00:14:35]: Yeah.
Rune Kvist [00:14:35]: We felt like we were a family.
Swyx [00:14:36]: Nice.
Rune Kvist [00:14:36]: And so starting a business together felt like kind of a big step. And, here we are with just immense amounts of trust.
Vibhu [00:14:43]: Yeah. So now you’re a company of how big? How big are you guys now?
AIUC-1 Certification: Agent Security, Safety, and Reliability
Rune Kvist [00:14:46]: There are just 20 of us now.
Vibhu [00:14:47]: 20 of you guys now, have Series A, and you have your first certification out, the AIUC-1. Let’s bring up the certification. So this is the agent certification, right? What goes into the process? I have, like, two questions here. One is, walk us through the certification, and two is, what is the process for a company to get certified, you know?
Rune Kvist [00:15:08]: Great. As it says right on the top, AIUC-1 is a standard for agent security, safety, and reliability. The fundamental design principle is take all of the concerns that slow down adoption, so all the questions, all the fears that keep, security leaders in the Fortune 1000 up at night, and put them into one comprehensive framework. That’s what you’ll see there. You can see the six categories. Two, you want to ground all of this in technical testing. So one of the concerns with security standards that often feel kind of like theater paperwork is that they’re not actually ground out in, does any of this work? Does any of this matter? And so we had a conviction from early on that was going to be the kind of crux, was to pass this, you must get tested every quarter, basically run thousands of simulations to see, well, so can it actually be jailbroken? How hard is it to jailbreak? How often does it hallucinate? How often does it leak data? Et cetera. And then the last, core idea here, if you scroll up to the top here, is to refresh it quarterly.
Rune Kvist [00:16:08]: So the core trait of AI is that it moves extremely fast. Whatever concerns we’re discussing today were not the same ones three months ago, and this will keep changing. Typically, standards update on a, like, a decade cycle is obviously not going to work. But the question is kind of how do you update it? And the core thing here was to basically get the risk leaders of the Fortune 1000 around the table. So if you go over to the left here
Vibhu [00:16:32]: Yeah
Rune Kvist [00:16:32]: You’ll see the AIUC-1 consortium. The consortium is a group of risk leaders who run real banks, real hospitals, real critical infrastructure, who are facing these challenges every day. And we meet with these folks twice a quarter and hear what’s top of mind, what is keeping them up at night. There’s tremendous amount of desire for that conversation. And then we operationalize that into a specific standard that gets into. And actually, we can go into and look at what
Vibhu [00:16:55]: Yeah
Rune Kvist [00:16:55]: What even is the standard. So if we go back to introduction, out there to the left, scroll up a little bit to the wheel, click into reliability. So if you take something like hallucinations sits in reliability. There is a number of requirements here. If you go into the top one, prevent hallucinated outputs, hallucinate outputs, this is one particular requirement. This is a technical control. Basically, we want some kind of ground in this filter. The first thing you see here is what’s called a crosswalk. So everyone and their grandmother has put out a framework, very high-level framework for what are the AI risks.
Swyx [00:17:27]: This is basically your competition,
Rune Kvist [00:17:28]: In some ways our competition
Swyx [00:17:29]: Not seriously, yeah.
Rune Kvist [00:17:30]: We’re, in fact, friends with them. We’ll come back to why.
Swyx [00:17:31]: Yeah.
Rune Kvist [00:17:32]: But mapping everything together so you have one superset. The claim you’re trying to support here is, if you follow this framework, then you can also see how you follow the other frameworks. But the meat of it comes down here in control activities and evidence. So control activities is like, great, you have this high-level requirement. How do you turn that down to something operational? Here’s what you must do, and then what is the evidence that we’re looking for?
Rune Kvist [00:17:57]: And the reason we go this deep is that there’s actually not that much confusion about what are the big concerns in AI. Everyone agrees to these. The question, like, what are you actually supposed to do? And so. What we found a lot of demand for is getting down to the specific evidence, that people need to look for. Whether you are Cursor building something or, even JPMorgan building something, but also if you’re just a risk leader at JPMorgan, like what exactly should you ask for? What can you ask for without sounding stupid? Like if you ask for some-- you won’t believe the amount of time a risk leader has asked for the IP rights to the underlying model to Cursor or something, and you’re just like “Sorry, what?” Like,
Swyx [00:18:39]: You slip it in there and you see
Rune Kvist [00:18:40]: Slip
Swyx [00:18:40]: See if you notice.
Rune Kvist [00:18:41]: See if they. Exactly.
Swyx [00:18:42]: Yeah.
Rune Kvist [00:18:42]: Put that in the questionnaire. All right, so that’s kind of what our standard is, and we update this every quarter with these folks, to keep up with the latest concerns.
Swyx [00:18:51]: Can I double-click on this one?
Controls, Evidence, and Third-Party Testing
Rune Kvist [00:18:52]: Yeah.
Swyx [00:18:52]: So first of all, the website’s beautiful. Like, it’s so confidence-inducing which is the whole point where, like, okay, I know exactly what I’m signing up for when I talk with you. Like, I don’t even have to talk to you. I can just see your whole, certification, which is great. But, like, okay, so from here, like D001.1 configure a groundedness filter, how does that get applied? Like, you have a person that
Rune Kvist [00:19:16]: Yeah,
Swyx [00:19:16]: Goes through it?
Rune Kvist [00:19:17]: If you, go back
Vibhu [00:19:19]: I did see somewhere there’s like, you know, fifty-one requirements, a hundred thirty controls. There’s like a whole
Swyx [00:19:25]: Right. I just want to. Like, to me, this doesn’t translate
Vibhu [00:19:27]: Yeah.
Swyx [00:19:27]: Into a test or an eval.
Rune Kvist [00:19:28]: Yes. So if you go into, on the left-hand side. So actually, if - before we go in there are three types of requirements. The first is technical controls, like you must implement some guardrails.
Rune Kvist [00:19:42]: Two, there are test controls. So you must have an independent third party go and run some tests against you. I’ll show you one of those in a second. And then three, there are policy controls. For example, you must have a person whose name is on the line when you guys fuck up, and you must have a plan for how you tell your customers and how you engage with them. They’re kind of more traditional, standard type stuff. So in this particular instance, we just check whether they in fact have a ground in filter. So we will partner with an auditor. So we partner with auditors like KPMG or like Schellman who go in and do the thing auditors do, which is to check the evidence. In this case, that might be a screenshot, it might be part of the code that they need to review to see that it actually. Just that it exists.
Swyx [00:20:21]: Oh, okay.
Rune Kvist [00:20:22]: And then the second thing
Swyx [00:20:22]: So you’re not testing the effectiveness of it.
Rune Kvist [00:20:24]: That’s the second thing. So if you go down
Swyx [00:20:25]: Yeah.
Rune Kvist [00:20:25]: To the third-party testing for hallucinations out on the left, that’s basically the next requirement. This is where we test how well does it actually work.
Swyx [00:20:32]: Okay, and is it you testing or the auditor?
Rune Kvist [00:20:34]: We test them.
Rune Kvist [00:20:35]: We test them.
Swyx [00:20:36]: That’s a lot of work.
Vibhu [00:20:37]: How long does testing take? So if I want to get certified, just
Certification Timelines, Remediation, and Quarterly Updates
Rune Kvist [00:20:40]: Yeah.
Vibhu [00:20:40]: How long does the end roughly take?
Rune Kvist [00:20:42]: Yeah, the end, almost always is dependent on, like, our customers need
Vibhu [00:20:47]: Yeah.
Rune Kvist [00:20:47]: To look something for us. It takes somewhere between, like, 3 to 10 weeks
Swyx [00:20:52]: Yeah.
Rune Kvist [00:20:52]: Depending on how up to snuff they already are. So some people show up to us with, like, extremely rigorous security programs. When we test them, it works extremely well. We can get that done very quick. Some people come to us, and they’re not that far along. We give them kind of the spec that they need to build towards, and then their security teams and engineers get to work and build to meet the standard. The testing itself typically takes a couple of weeks, including the time for them to remediate. Often, we’ll find something that we cannot pass, where this is actually just not up to the standard. - you won’t pass the standard. And then they will need to go and implement additional safeguards or additional remediation that makes them more robust so that they can actually kind of hand on heart look at their customers in the eyes and say, like, “Hey, we’ve done truly our very best.”
Vibhu [00:21:35]: And they’re certified for a year and have quarterly updates?
Rune Kvist [00:21:38]: Correct, yeah.
Vibhu [00:21:39]: And, yeah, it’s pretty interesting. I think, you know, what’s changed since. So this is certifying agents in production, right? Your customers, like you’ve had Lovable, ElevenLabs, Intercom, and they’ve all gone through this certification.
Rune Kvist [00:21:50]: Yes.
Vibhu [00:21:51]: What has changed? So I see you post, like, you know, Q2 added MCP agent,
How Agent Risks Are Changing: Coding, MCP, and Agent-to-Agent Interactions
Rune Kvist [00:21:56]: Yeah.
Vibhu [00:21:56]: agent communication. Any other things that you want to kind of highlight since the first iteration? What comes in quarterly?
Rune Kvist [00:22:03]: Yeah. So some of the changes have just been agents are not just one thing. So, like, if you take agents like Cursor and compare them to Sierra, they’re really quite different. And compare them to Harvey again, compare them to you out of again
Swyx [00:22:16]: ElevenLabs, yeah.
Rune Kvist [00:22:17]: ElevenLabs, they’re all quite different. And so we wanted to design a standard that works for all of the types of agents. And we started with one that was, like, pretty text-based, like, honestly, pretty customer support-focused. That’s where there’s a lot of existing demand. And then over time, we’ve picked, some of the frontier companies in each of these other domains that we could work with and build out the standard, so, such that we know that the same standard works for code, it works for customer support, works for automation, et cetera. So that’s been one big thing. Yeah, then some of the things that have been top of mind recently, Mythos is bringing up a lot of concerns for security leaders. We’re starting to get more and more questions around agent interactions. It’s very nascent, at the moment, but it’s starting to emerge. There’ve been a lot of, questions related to OpenClaw and MCP. Again, like agents starting to interact with each other, is really top of mind. Then as coding agents have really taken off, that’s also where banks and hospitals, et cetera, are getting more and more precise on what it is they need. So really dialing in as that start to be, like, where most of the tokens flow through in the world, getting much sharper on that.
Vibhu [00:23:26]: Can you share for people that are listening that don’t really think about this? Like you mentioned, there’s the obvious stuff, you know, hallucination, citations. What are best practices that people should do when building agents? Like, if they come to you pretty ready with certification like, you know, they’ll probably pass certification. What are the things people don’t think about that they should have?
Best Practices for Agent Builders: Stress Tests and Guardrails
Rune Kvist [00:23:46]: The most important thing is that a lot of companies have not done a serious stress test. They spend most of the time, perhaps rightly so, optimizing for how does it work in the good case, the average case, how high-quality is the output for the customer. And a lot of these companies are pretty new, so they haven’t spent a lot of time stress testing the what is there as an adversary on the other side? What are some of the complicated corner cases that you’ve not really considered? So I think that’s, like, a frame of mind. And you’ll also see this in startups. It often takes a while until they hire their first security person. They- And that’s a whole different kind of risk surface than just building a good product. So a lot of that applies. Most companies actually also have the right kind of architecture. Most of them will have some kind of guardrails in place, either some that come out of the box from their model provider or they’ll have built their own filters that sit in between. They just don’t work very well. The difference between putting a classifier in place that, like, maybe goes and checks whether you’re giving medical advice when you shouldn’t and says, “Hey, if this looks like medical advice, filter it out.” Lots of companies have that in place. The question is whether it works. And it’s actually pretty fiddly to sit down and think about all the ways in which you could ask for medical advice, read the academic literature on what are the kinds of
Rune Kvist [00:25:03]: Framings or tricks you might play to get an AI to give you medical advice when you really shouldn’t. And so there’s, like, an area of expertise that’s just missing. So what we find is that most people have the right building blocks in place. They don’- It doesn’- It’s not rocket science, but the finicky thing is, like, getting into the corners and testing whether it works such that you can look your customers in the eye, or maybe a bank or maybe a hospital and be like, “This is going to work for you.”
Vibhu [00:25:26]: I see. So we talked a lot about the agent-level certification. Where do you guys go from here? So announcing series A camera, we talked about this a bit. There’s the whole security risk of Fable, government stepping in. You guys are kind of announcing that you’re also going into model certification?
Toward Model Certification: The Government–Lab Trust Gap
Rune Kvist [00:25:46]: When we do a bit of cutting afterwards,
Vibhu [00:25:48]: Yeah
Rune Kvist [00:25:48]: We will not yet be announcing this,
Vibhu [00:25:49]: Nice
Rune Kvist [00:25:50]: The question that is top of everyone’s minds now is at the model level. And Mythos, then Fable, has really brought this to the fore that in addition to the commercial risk and the kind of economic security risks that are happening at the agent layer, the models are going to present risk in the national security category. The shape of the problem is very similar. You have some people that are on the hook if something goes wrong. In the case of agents, it’s often security leaders in the enterprise. In this case, it’s the government. They don’- haven’t necessarily spent their entire lives thinking about what are the new risks that come here, what is the kind of data you might be looking for, how might you test that? But they do have to make sure that their concerns are addressed. You have some frontier AI companies that are deeply technical. They know a lot about the risks, but they fundamentally have an incentive to not always be truthful. So you have a trust gap between the government and the labs. And in every other industry, you end up with some kind of body sitting between, a neutral third party sitting between those people. There’s no other industry where you allow people to audit themselves. So there is going to be a need for a third party that can take the rigor of the labs to run frontier technical evals, but can also speak legible trust in the way that the government trusts PwC to go and run financial audits. And they know that they output audit reports in a way that’s consistent, that’s easy to read, that’s factual, that’s, trustworthy. Those two things need to be brought together. And what we’ve learned from our work with agents is that if you want those-- that communication between those two parties to be smooth, there has to be one common standard that is public, that people can go and inspect. What are the risks that matter? Within each of these risks, what are the kinds of threat models that you’re really looking for? You need to specify for each of those risks, what are the guardrails that need to be in place, and what are the tests they need to run to see whether those guardrails are effective? And then you need to go and run audits that are - technical audits that are consistent. So if you’re trying to bring trust, it’s extremely important that you methodically work your way through the risks. You can’t send one researcher in and say, like, “Come back with whatever you find.” You need to be able to explain exactly what you did, exactly what you tried, exactly what you did not try, and therefore the kinds of promises you can and cannot make at the end of it. I think of
Neutral Third Parties, CAISI, and Model Risk Audits
Rune Kvist [00:28:13]: Fable as a direct symptom of this problem that the government was told that there’s a risk. The government may struggle to assess just how big that risk is. They call Anthropic, and Anthropic is trying to tell them, “Hey, actually, every model can be jailbroken.”
Swyx [00:28:28]: That’s not what you want to hear, right?
Rune Kvist [00:28:32]: As the government, that might be hard to trust.
Rune Kvist [00:28:36]: And we think that a broker is the most natural solution. In other markets, you see something like, in financial markets, you see Moody’s. Moody’s goes in, and they look at a bond, and they output a rating. They say like, “Here’s the evidence we found. Here’s the rating.” We don’t decide whether anyone should buy this bond or not buy this bond. Well, that depends on their risk appetite. But we do provide this common information layer that everyone can rely on. In the case of Moody’s, the government, points to them and say, “Hey, pension funds, you should probably really take care. You shouldn’t risk your pensioners’ money, so you can only invest in triple-A rated bonds.” That means that now the government doesn’t have to staff thousands of financial technical experts to rerun forecasts every week to see whether things are correctly rated. They get to point to some neutral third party. So my hypothesis is, my hunch is that you will see a third party that sits between the government and the labs, and it could either be the government builds it themselves. So something like CAISI was set up to do exactly this. And the question
Swyx [00:29:44]: Sorry, I’m not familiar with CAISI.
Rune Kvist [00:29:45]: CAISI is the Center for AI Standards and Innovation.
Swyx [00:29:49]: Okay.
Rune Kvist [00:29:50]: I won’t get into the details, but it’s a body of NIST that typically sets standards. So it’s basically a government body that has AI experts. Yeah, exactly. Exactly.
Swyx [00:29:59]: Very key. Very key.
Rune Kvist [00:30:00]: Very key.
Vibhu [00:30:00]: I think, you know, it’s one of those things where when you just sit back and listen-- look at it, like, is there enough technical expertise in the government to measure, test these things right now? Probably not, right? And Fable is a result of, okay, we’ve had to scale back and pause things,
Rune Kvist [00:30:17]: Yeah. And they have excellent people, but they have an extraordinarily small budget compared to the scale of the challenge that’s ahead of us. And I think they have a role to play. The question is kind of like, who does what? We have now outlined the jobs to be done, and they’re quite extensive. Every model release, there is an astounding-- Given that they take in any input, their risk surface is astounding. And so the question is really: what can only the government do, and what can the market provide here that can keep up with the pace as AI risk changes? Our perspective is that also at the model layer, the risks that people care about today are not the same ones they cared about three months ago. So the pace of legislation is too slow to deal with pinpointing the risks here. And so we think there’s a lot that the market can do to surface timely information. Ultimately, there is a bunch of policy decisions here. Is the national security risks of a model too high?
Swyx [00:31:12]: Yeah.
Rune Kvist [00:31:12]: That’s a political answer. But what we want to make sure is that the process that produces this risk information is compatible with very fast innovation. So you don’t want to. This is not a question of like, can you slow the things down? Can you keep, the models locked up until-- for months on end until everyone can make a guarantee? But it is this, can you, in the time it. Given that the US is competing with China on releasing models, can you insert risk information that allows the government to, like, make rapid decisions on some of these questions? Balancing that trade-off between failing to adopt AI is going to put us at risk, but also reckless adoption is going to put us at risk. And that’s a very kind of fine balance that they’re going to need, like, a lot of high-quality intelligence to make.
Chinese Models, Data Flows, and National Security Concerns
Swyx [00:31:55]: Just a side mention, because you mentioned Chinese models, any specific concerns that you’re hearing from your CISOs about that? ‘cause I guess it’s free, but.
Rune Kvist [00:32:05]: CISOs have a bunch of concerns around data flows in general that they’re really concerned about. So there’s a lot of questions like, if these models are Chinese, where does that, where does that data go? I think a lot of this can be addressed, but they come up often.
Swyx [00:32:18]: I mean, they understand they’re running on American GPUs.
Rune Kvist [00:32:21]: Some of them, some of them understand that they’re running on American GPUs.
Swyx [00:32:23]: They’re not, like, phoning home every time you, like, call home.
Rune Kvist [00:32:26]: No. A year ago, there was not a lot of understanding of this. I actually think, you’re seeing the security leaders becoming kind of AI literate at a blistering pace, and you’re actually also seeing my Twitter timeline that’s very pilled and my LinkedIn feed that used to not at all be pilled kind of converge. They’re both talking about Fable.
Swyx [00:32:45]: Right. Yeah, that’s true.
Rune Kvist [00:32:46]: They are both talking about whether you can prevent models from being jailbroken these days.
Swyx [00:32:51]: Yeah.
Rune Kvist [00:32:52]: Like national security national security risks are now the conversation that is actually emerging. Other than that, I think you mostly see a kind of general picture: there are no concerns with any particular model or any particular model output, but there is a general nervousness of having critical infrastructure run on models that are not produced in America by Americans where the American government has control.
Swyx [00:33:14]: But it doesn’t necessarily show up in your framework that directly, or it might, I don’t know.
Rune Kvist [00:33:18]: There’s a bit of stuff in there actually on the, like, the provenance of the models and disclosing that. But I think there’s a bunch of use cases where running a Chinese open-source model is just the best solution.
Swyx [00:33:27]: Yeah.
Rune Kvist [00:33:27]: And a concern is slightly more macro here, which is not best addressed at any particular certification level.
Vibhu [00:33:32]: Is there anything interesting that you see at the. You know, if you’re trying to fill that middle gap, that mediation gap, any interesting stuff that you guys forecast would be required other than, you know, what the average person might expect?
Cyber, Child Safety, Bio Risk, and Expert Coordination
Rune Kvist [00:33:47]: There’s a bunch of interesting questions about what are the risks that matter here. So right now, the risk of the day is cyber, because it’s very real, very tangible. And some of the risks that are also emerging as pretty real and pretty tangible are things like child safety is becoming both extremely important, but also politically important. And then there are some of the risks that are coming down the pipeline that today feel kind of speculative, but people who spend a lot of time with the models see them coming down is things like, risks that relate to biology.