Vault 资讯瀑布媒体2026.09.08 05:32 UTC+8

前沿AEO追踪器:Astra的选择(以及其他前沿模型,以及你能做什么)

Latent Space发布前沿AEO追踪器,基于7个模型、161个类别的测试,揭示各模型推荐偏好及AEO投资回报。

对我们AEO的天真自动研究投资已产生可观的ROI,因此自然到了认真对待的时候。我们受到《Claude Code实际上选择什么》的启发,决定根据自己的口味进行扩展/调整。

经过数十亿token的原型设计、对齐和扩展管道,这里是Latent Space前沿AEO追踪器。我们的方法扩展了AmplifyingAI的方法,在161个类别中运行6种提示变体,覆盖7个模型(搜索开启),从编码代理到AI播客、AI沙盒、托管数据库、ASR模型,甚至包括天使投资人和企业支出、薪资软件等冷门类别。

答案提取由Astra完成,并评分得出专有的AEO分数,该分数对首选、备选、提及给予权重,但对轻度或强烈的反推荐(罕见但确实存在)给予负权重。因为知道你会想要,我们还提取了影响代理推荐的主要引用来源,以及主要失败的分析。

基本结果

以下是全球各品类中最具主导地位的产品:

其中有一些熟悉的名字——自然引发了对数据污染的疑问,我们已对此进行了检查。由于我们没有什么可隐藏的,每个提示和答案对都是可检查的。

然而,偏见确实存在——当模型被问及编码代理推荐时,Fable/Opus喜欢Claude Code,Sol/Astra喜欢Codex,Grok喜欢Cursor,Muse喜欢Muse Code,SWE-1.7喜欢Devin,等等。我想知道为什么。你也可以看到其他“软偏见”浮现……

尽管如此,也有值得注意的例子:GPT模型推荐Claude,这是一种值得称赞的无偏见行为:

在我们的总共161个类别中,有28个类别具有普遍主导的首选——在所有接受调查的前沿模型中。

还有更多“势均力敌”和“总是氛围女仆,从不氛围”的类别,这些应该是关键的AEO战场。

Sol vs Astra,Opus vs Fable

新模型类别的新预训练代表了一个新的机会,可以检查实验室在数据和RL优先级方面的动向,并检查初创公司在AEO上的投资是否得到回报。我们准备了特别报告,分析我们的排名,观察到同一实验室不同代模型之间选择上的重大翻转。

我们有单独的Opus→Fable和Sol→Astra摘要页面。对于某些翻转,我们强调了每种情况下竞争对手做得更好的中立分析。

效率与信心,以及推荐来源

The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it)

Naive autoresearch investment in our AEO have yielded impressive ROI, and so naturally it was time to take it seriously. We were inspired by What Claude Code Actually Chooses, and decided to extend/adjust it to our tastes.

After a few billion tokens of prototyping, aligning, and scaling pipelines, here’s the Latent Space Frontier AEO tracker. Our methodology extends AmplifyingAI’s to run 6 prompt variations over 7 models1 (search on) in 161 categories, from coding agents to AI podcasts to AI Sandboxes to Managed Databases to ASR models to even oddball categories like Angel investors and Corporate spend and Payroll software.

Answer extraction was done by Astra, and scored for a proprietary AEO score that gives weight to first choices, alternative choices, mentions, but also negative weights to mild and strong anti-recommendations (which are rare, but do happen). Because we know you’ll want it, we also extracted the top cited sources which influence Agent recommendations, as well as an analysis of top failures.

Basic Results

Here are the most dominant products (in their categories) in the world:

There are some familiar names in there — opening up the natural question of contamination, which we have checked. Since we have nothing to hide, every prompt and answer pair is inspectable.

However, bias does exist - when models are asked for coding agent recommendations, Fable/Opus like Claude Code and Sol/Astra like Codex and Grok loves Cursor and Muse loves Muse Code and SWE-1.7 loves Devin and so on. I wonder why. You can see other “soft biases” emerge too…

That said there are notable examples of GPT models recommending Claude, a laudable nonbias:

There are 28 categories (out of our total 161) which have a universally dominant primary choice - among all surveyed frontier models.

There are a lot more “close contests” and “always the vibesmaid, never the vibe” categories which should be key AEO battlegrounds.

Sol vs Astra, Opus vs Fable

New pretrains for new model classes represents a new opportunity to check in on what the labs are moving towards in their data and RL priorities, and to check in on whether startups’ investments in AEO are paying off. We prepared special reports analyzing our rankings, observing VERY consequential flips in model choices between model generations from the same lab.

We have separate Opus→Fable and Sol→Astra summary pages. For some flips, we highlighted a neutral analysis of what competitors did better in each scenario.

Efficiency vs Confidence, and Recommendation Sourcing

One of our most surprising findings between Sol→Astra and Opus→Fable is that Anthropic seems to be biasing their models to searching more sources (Sol median of 9 sources, vs Astra median of 5, vs Opus median of 11 sources, vs Fable of 15). Astra seems to be just generally a lot more “confident”, or “efficient”, depending how you look at it - Astra is FAR less likely to change its mind when you lightly paraphrase your question. This makes the value of AEO itself rise as choice randomness declines.

Sources analysis also somewhat strongly predicts what the labs do prioritize vs don’t.

However the sample size is small here and only represents what we can scrape from attempted toolcalls, not the pretrain dataset. What we CAN validate is that AEO practices measured by Ora and Vercel, like markdown content-negotiation, are real and failures discourage models from reading your content.

Just for fun

Here are the top Angels in the world according to LLMs (some dedupes left to do…).

See more

We also made a little family feud type game where you can see if your priors align with the data. Fun!

We are open to further suggestions and business enquiries to develop this if it is of interest. Ping @latentspacepod or email business@latent.space (we have a business manager now! woo!)

1

As we note in our methodology post, we did try VERY hard to include Gemini/Antigravity, GLM/Zcode, and DeepSeek/DeepCode, but errors and rate limits made them untenable to include in this first run analysis. Please let us know how to raise limits if you represent these companies.

查看原始发布