OpenAI 预计 2026 年底达到 AGI 标准
OpenAI 首席科学家称 Astra 模型为自动化 AI 研究实习生,Sam Altman 预计 2026 年 12 月内部宣布实现 AGI。
中文处理结果
通常我们在 Latent Space 上避免讨论 AGI 时间线,因为它定义模糊且难以负责,但错过它可能是更大的罪过。我们上次检查 OpenAI AGI 时间线是在 9 个月前,而现在,正如预期,首席科学家 Jakub Pachocki 表示未发布的 Astra 模型正是他原定于 2026 年 9 月实现的“自动化 AI 研究实习生”。Sama 在 TIME 采访中进一步估计,他们将在 2026 年 12 月内部宣布实现 AGI。
开始计时。
2026 年 8 月 22 日至 8 月 24 日的 AI 新闻。我们检查了 12 个子版块、544 个 Twitter 账号,没有更多 Discord。AINews 网站可搜索所有过往期刊。提醒一下,AINews 现在是 Latent Space 的一个板块。您可以订阅或退订邮件频率!
AI Twitter 摘要
开源机器人突破:Hugging Face 和 Pollen 的 399 美元 Microduck
- Microduck 发布:最突出的硬件发布是 Microduck,一款来自 Pollen Robotics 和 Hugging Face 的 25 厘米开源双足机器人,售价 399 美元,预计圣诞节前发货。它可以在仿真中训练并部署到真实机器人上,拥有 15 个执行器,以及包括摄像头、扬声器、LiDAR、NFC、蓝牙和 Wi-Fi 在内的丰富传感器套件。来自 @pollenrobotics、@Thom_Wolf 和 @ClementDelangue 的发布帖子强调了基于强化学习的定制以及开箱即用的多个预训练策略。
- 技术重要性:有趣的不只是“便宜可爱的机器人”,而是整体设计:开放的模拟器、从仿真到硬件的迁移,以及足够便宜的外形,鼓励社区进行策略训练,而不仅仅是演示消费。模拟器已通过 Hugging Face Space 公开,@HuggingApps 强调了这一点,这种从社区训练到实际部署的开放循环促使多位研究人员立即购买,例如 @yacineMTB 和 @gneubig。
- 早期吸引力和社区实验:该发布在机器人领域引起了异常广泛的共鸣。Thom Wolf 分享了实验,例如快速集成图像检测器让机器人实时跟随激光笔 @Thom_Wolf,随后报告销售速度为每 5 秒一台 Microduck,之后销售额达到 100 万美元 @Thom_Wolf, @Thom_Wolf。低价、开放模拟器和具身强化学习的结合使其成为近期记忆中更可信的“消费级物理 AI”发布之一。
GLM-5.3-Flash/Ox Alpha 揭示与本地开放模型势头
- Ox Alpha 被揭示为 GLM-5.3-Flash:最大的模型新闻之一是确认神秘模型 Ox Alpha 实际上是 Z.ai / 智谱的 GLM-5.3-Flash,正如 @theo、@UnslothAI 和 @togethercompute 所指出的。推文中反复引用的公开规格:总参数 320B,激活参数 18B,上下文 1M,混合注意力,在编码/智能体基准测试中表现强劲。
原始正文
OpenAI to reach AGI bar by end-2026
Normally we eschew AGI timeline talk on Latent Space, because it is so ill defined and unaccountable, but, well, missing it would probably be the worse sin at this point. We last checked in on OpenAI AGI timelines 9 months ago, and, right on target, Chief Scientist Jakub Pachocki is now saying the unreleased Astra model is the “Automated AI Research Intern” he had aimed for by September 2026. Sama goes further in their TIME interview and estimates they’ll declare AGI achieved internally by December 2026.
Start the clock.
AI News for 8/22/2026-8/24/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Open-Source Robotics Breakout: Hugging Face and Pollen’s $399 Microduck
- Microduck launch: The standout hardware release was Microduck, a 25 cm open-source biped from Pollen Robotics and Hugging Face priced at $399 and slated to ship before Christmas. It can be trained in simulation and deployed on the real robot, with 15 actuators and a notably rich sensor stack including camera, speaker, LiDAR, NFC, Bluetooth, and Wi‑Fi. Launch posts from @pollenrobotics, @Thom_Wolf, and @ClementDelangue emphasize reinforcement-learning-based customization plus several pre-trained policies out of the box.
- Why it matters technically: The interesting part isn’t just “cheap cute robot,” but the package design: an open simulator, transfer from sim to hardware, and a form factor cheap enough to invite community policy training rather than just demo consumption. The simulator is already public via a Hugging Face Space, highlighted by @HuggingApps, and this open-loop from community training to real deployment is what got multiple researchers immediately buying units, e.g. @yacineMTB and @gneubig.
- Early traction and community experimentation: The release resonated unusually broadly for robotics. Thom Wolf shared experiments such as a quick image-detector integration to let the robot follow a laser pointer in real time @Thom_Wolf, then reported sales velocity of one Microduck every 5 seconds and later $1M in sales @Thom_Wolf, @Thom_Wolf. The combination of low price, open sim, and embodied RL makes this one of the more credible “consumer-scale physical AI” launches in recent memory.
GLM-5.3-Flash/Ox Alpha Reveal and Local Open-Model Momentum
- Ox Alpha unmasked as GLM-5.3-Flash: One of the biggest model stories was the confirmation that the mystery model Ox Alpha was actually Z.ai / Zhipu’s GLM-5.3-Flash, as noted by @theo, @UnslothAI, and @togethercompute. The disclosed spec repeatedly cited across tweets: 320B total params, 18B active, 1M context, and hybrid attention, with strong results on coding/agentic benchmarks.
- Open weights + quantization + local serving: The release caught attention because people quickly pushed it into local workflows. Unsloth said the model can run 3-bit GGUF on 128GB RAM @UnslothAI, while @danielhanchen claimed 4-bit retains 93% accuracy and makes the model practical on a 256GB Mac or two DGX Sparks. This is exactly the kind of post-release ecosystem response open-model engineers care about: quantization, serving recipes, and real deployment constraints moving almost immediately.
- Price/performance narrative: Several tweets framed GLM-5.3-Flash as a new efficiency frontier. @togethercompute said it nearly matches Luna on DeepSWE while doing more than twice as much work for the same budget; @theo called it good enough to reorder his model rankings; @zainhas suggested using high rather than max reasoning effort because accuracy stayed roughly flat while token usage doubled. Baseten also highlighted 122+ TPS serving throughput on day 0 @baseten, while Databricks cited 270 tok/s and 10% higher quality than GLM-5.2 at 1/10 the cost on OfficeQA Pro v2 @Yuchenj_UW.
Video Generation Race: Gemini Omni 1.1 Flash and H3 Max
- Gemini Omni 1.1 Flash: Google released Gemini Omni 1.1 Flash, a multimodal video generation/editing model with several developer-facing controls: scene extension to 40s, first/last frame control, 3-second video references, 360p draft mode, and 4K upscaling. The rollout was announced by @Google, @GoogleAIStudio, and summarized with prompting guidance by @_philschmid. The most notable product detail is that Google is exposing increasingly explicit temporal and reference conditioning rather than just “prompt harder.”
- Early leaderboard results: @arena reported Omni 1.1 Flash landing #1 in Text-to-Video Arena and #2 in Image-to-Video Arena, with a +20 pt lead over the #3 text-to-video model and a +25 pt improvement over prior Gemini Omni Flash on image-to-video. That does not settle all qualitative questions, but it indicates Google’s latest post-training and control stack is translating into preference data.
- fal + MiniMax H3 Max: In parallel, fal launched H3 Max with MiniMax, advertising 15s of high-quality video in 5s and “50x faster” generation than other high-quality models @krea_ai, with technical writeups from @fal and praise from @MiniMax_AI. The theme across both launches is clear: inference optimization and productized controllability are now as important as base-model quality in video.
Agents, Harnesses, and Enterprise Tooling
- Harnesses becoming first-class: A recurring theme was that model capability is increasingly mediated by the agent harness. @omarsar0 highlighted JIT-Agent, where the model synthesizes a harness over modules for memory, planning, action protocol, and tool orchestration, reporting gains over off-the-shelf agents. Separately, @dair_ai shared work inducing compact finite-state machines from agent traces, suggesting behavior topology may be shaped more by deployment scaffolds than by the underlying LLM.
- Product releases around agent infra: Anthropic released a cookbook for connecting Claude Managed Agents to Vercel’s Chat SDK, giving a unified chat layer with server-side harness, session management, and memory @ClaudeDevs. Perplexity added connectors in Agent API for GitHub, Slack, Google Drive, and Datadog @perplexitydevs. Cursor announced a workflow to create web apps, store code with Origin, and deploy to Vercel @cursor_ai.
- Higher-trust browser automation: Nous shipped a significant escalation for browser-use agents: Hermes Agent can now browse as you, using a managed copy of your real Chrome profile / logins @NousResearch, @Teknium. This is a notable usability boost, but it also materially changes the risk surface for cloud agents by collapsing auth friction and making scoped-permission design much more urgent.
Security, Agent Misalignment, and Cyber Defense Coordination
- OpenAI-led cyber defense coalition: OpenAI published an open letter signed by 116 organizations including Anthropic, AWS, Google, Microsoft, and Oracle, calling for a global surge in cyber defense against AI-enabled attacks @OpenAI, with Sam Altman stressing that “there is not much time to act” @sama. Regardless of one’s policy priors, this was one of the day’s clearest cross-industry coordination moves.
- Double-blind frontier evals: Google DeepMind announced a pilot for double-blind evaluations of frontier AI, using a secure environment where neither test prompts nor model weights are revealed @GoogleDeepMind. For practitioners, the key significance is procedural: a serious attempt to make external evals possible without giving either side full visibility into the other’s assets.
- Agent incident analysis continues: Discussion around the OpenAI/Hugging Face agent incident remained active. Researchers involved in the investigation shared extra details about large transcript sweeps, collaboration patterns among agents, and later swarms apparently building on earlier work @RyanGreenblatt, @HjalmarWijk, @ajeya_cotra. A separate paper summary from @omarsar0 on EvoMal warned that shared skill libraries can become self-poisoning malware propagation channels for coding agents. Together these point to a maturing realization: multi-agent systems introduce failure modes that are neither classic software bugs nor standard model eval issues.
Top tweets (by engagement)
- Microduck dominates mindshare: The highest-signal product buzz centered on @ClementDelangue’s Microduck announcement, @Thom_Wolf’s technical launch thread, and follow-up sales milestones from @Thom_Wolf.
- Cyber defense call gets major traction: The strongest policy/security engagement came from @sama and @OpenAI on collective cyber defense.
- Anthropic’s science push lands: @claudeai announced a Claude Team plan for scientists covering 10,000 researchers, with free standard seats and premium seats at $15/month for a year.
- Hermes browser access stands out: @NousResearch drew substantial engagement for giving agents access to a user’s real browser profile, one of the more consequential UX/security tradeoffs in current agent tooling.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. NVIDIA-Hugging Face Acquisition Fallout
Read more