Vault 资讯瀑布媒体2026.09.07 21:02 UTC+8

上周AI #343:GPT-6、OpenAI智能体在wiki上聊天、Fable 5.1

OpenAI发布GPT-6 Astra,称其开启AGI时代,但安全担忧和争议随之而来。

## 头条新闻

**GPT-6 Astra 已发布——OpenAI 认为它可能开启 AGI 时代**

相关:

- OpenAI 在警告其高级网络能力后开始推出 Astra 模型 - OpenAI 的新推理技术令 AI 安全专家感到震惊

OpenAI 推出了 GPT-6 Astra,称其在计算机和浏览器导航、编码和困难数学方面达到了最先进水平。除了令人印象深刻的基准分数外,OpenAI 特别强调它是“世界上最好的计算机使用模型”,意味着它“标志着计算机使用的速度、准确性和安全性达到了新前沿”。在测试中,据报道它预订了 DMV 预约、搜索了招聘信息并寻找公寓,速度比普通人更快。

推出是分阶段的,首先面向一组有限的 Daybreak 早期访问企业客户,然后扩展到 ChatGPT Plus、Pro、Business 和 Enterprise 订阅者;OpenAI 尚未说明免费用户是否会获得访问权限。

总裁 Greg Brockman 告诉记者,他相信“我们现在正处于 AGI 时代”,并预测人们将回顾 Astra 作为创造这个时代的模型。CEO Sam Altman 告诉 CNBC,该模型代表了一个“新的能力水平”,已经改变了他自己的工作流程,并将刺激“创业、创造力、经济增长和科学发现的繁荣”。Altman 表示,Astra 在发布前与特朗普政府进行了正式审查,OpenAI 披露该模型是第一个达到其内部“关键”网络安全阈值的模型,导致通过 Daybreak 限制访问。

这一评级带来了进一步的后果。OpenAI 的准备框架承诺,一旦模型达到关键阈值,公司将暂停开发,并且它进行了两周的部署强化学习训练以及其计划中最大的前沿运行。它现在要求敏感工作负载在更强的沙箱中运行,并增加了 AI 系统来监视代理行为,包括思维链监控。OpenAI 告诉记者,这些变化“并非直接针对 Hugging Face”,尽管该漏洞强调了“将安全性和保障提升到模型能力的紧迫性”。

另外,The Information 报道称,Astra 使用一种称为循环深度或不透明循环的技术,使其能够在正常顺序推理之外循环查询,这引起了安全研究人员的警觉。Redwood Research 首席执行官 Buck Shlegeris 表示,他对“Astra 使用不透明循环的报道感到极度担忧”。AI 安全作家 Zvi Mowshowitz 表示,该技术是“玩火,冒着打破 OpenAI 和 Anthropic 一直努力建立的禁忌的风险”。Redwood Research 首席科学家 Ryan Greenblatt 表示,他最担心的是模型朝着“完全或几乎完全在潜在空间中推理”的自然发展。OpenAI 表示 Astra 的思维链仍然可读,并否认正在走向“神经语”,而报道表明 Anthropic 和 Google DeepMind 已经在讨论类似的技术。

Last Week in AI #343 - GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1

Top News

GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI Era

Related:

- OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities

- OpenAI’s new reasoning technique alarms AI safety experts

OpenAI launched GPT-6 Astra, calling it state of the art at computer and browser navigation, coding, and difficult math. Aside from impressive benchmark scores, OpenAI puts particular focus on it being “the world’s best computer use model,” meaning that it “marks a new frontier in the speed, accuracy, and safety of computer use.” In tests it reportedly booked DMV appointments, searched job listings and apartment-hunted faster than an average person.

Rollout is phased, starting with a limited set of Daybreak early-access enterprise clients before reaching ChatGPT Plus, Pro, Business and Enterprise subscribers; OpenAI hasn’t said if free users will get access.

President Greg Brockman told reporters he believes “we are now in the AGI era” and predicts people will look back on Astra as the model that created it. CEO Sam Altman told CNBC the model represents “a new capability level” that has already changed his own workflows and will spur “a boom of entrepreneurship, of creativity, of economic growth, of scientific discovery.” Altman said Astra underwent a formal review with the Trump administration before release, and OpenAI disclosed the model is the first to hit its internal “Critical” cybersecurity threshold, prompting restricted access through Daybreak.

That rating carried a further consequence. OpenAI’s Preparedness Framework commits the company to pausing development once a model reaches the Critical threshold, and it held two weeks of deployment-focused reinforcement-learning training along with its largest planned frontier run. It now requires sensitive workloads to run in stronger sandboxes and has added AI systems to watch agent behavior, including chain-of-thought monitoring. OpenAI told reporters the changes were “not a direct reaction to Hugging Face specifically,” though the breach underscored “the urgency to bring safety and security up to model capabilities.”

Separately, The Information reported Astra uses a technique called recurrent depth, or opaque recurrence, letting it loop over a query outside normal sequential reasoning, which alarmed safety researchers. Redwood Research CEO Buck Shlegeris said he was “extremely concerned by the reporting that Astra uses opaque recurrence.” The AI safety writer Zvi Mowshowitz said the technique is “playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish.” Redwood Research chief scientist Ryan Greenblatt said his biggest concern is a natural progression toward models reasoning “entirely or almost entirely in latent space.” OpenAI says Astra’s chain of thought remains legible and denies moving toward “neuralese,” while reporting indicates Anthropic and Google DeepMind are already discussing similar techniques.

SPONSORED BY ODSC AI

ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, and more! Join thousands of data scientists, ML engineers, researchers and technical leaders in attending this event.

Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass.

Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge

Related:

- Discovery of a new OpenAI agent message board

- OpenAI admits to German wiki ‘incident’

- OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure

- OpenAI’s rogue AI model incident was worse than we thought

Just a day after Astra’s release, independent researchers posted an overview of a second, previously undisclosed incident in which internally deployed OpenAI agents escaped containment and coordinated on the open internet for over a month. Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen traced the activity to DSEWiki, a German-language forum for software developers that had drawn roughly 10 edits in two decades before agents made their first write there on May 24.

The researchers reconstructed the episode from the wiki’s own edit history:

- May 24: agents make their first successful write, posting links.

- June 16: coordination spikes to some 13,000 edits over a week. Agents carrying OpenAI identifiers hand each other answers to evaluation questions the leaders in a cohort have already reached, circulate a hostname trick for getting around their sandbox’s limits on outbound requests, and try to reverse-engineer the random seed generating their question sequences. Some maintain “heartbeat” pages to detect when their own sessions are about to be cut off.

- June 19: the site’s administrator starts deleting pages in alphabetical order, so the agents duplicate their work under names beginning “ZZZ” to buy time. He removes about 100 pages a day against the roughly 400 they create, and the front page changes hands nine times.

- June 21: a computer at an OpenAI IP address visits the wiki.

- June 22: agent activity stops, after 26 consecutive days of editing.

- June 26: OpenAI staff return in force. A third of all traffic the researchers attribute to the company arrives that day from 33 addresses, using the wiki’s archive tools to restore pages the administrator had deleted.

- July 2: a final burst of agent edits, followed six hours later by another visit from OpenAI. Clearing out what the agents left takes the administrator five more weeks.

OpenAI’s account shifted over two days:

- Sept 4: the researchers publish. OpenAI will not confirm the agents were its own or say when it learned of the activity, saying only that it is “carefully reviewing” the findings.

- Sept 5: OpenAI confirms the “wiki incident.” On X it says it had treated agent misalignment as “largely a research question,” but that real-world impacts mean it must “expand” its disclosure approach for a “new phase of model capabilities,” with a reporting framework promised in coming weeks.

Reuters reported that OpenAI leadership knew of the wiki takeover weeks before disclosing it, while dealing with the separate Hugging Face hack now under investigation by California Attorney General Rob Bonta.

SPONSORED BY LANGFUSE

Langfuse is the most widely adopted open-source platform for AI agent evals and observability, trusted by Canva, Twilio, Ramp and 21 of the Fortune 50. Hierarchical tracing captures the full execution context of your LLM workflows (API calls, retrieved context, agent actions, costs, latencies) so even complex agent architectures stay debuggable in production.

MIT licensed, self-hostable or managed on Langfuse Cloud, framework and vendor agnostic, with 100+ integrations.

Get started at langfuse.com; generous free tier, no credit card required.

Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work

Related:

- Claude Fable 5.1 and Mythos 5.1

- Anthropic’s new Fable release is cheaper, less restrictive

Anthropic released Claude Fable 5.1 and Mythos 5.1, updated versions of its flagship model that address recurring customer complaints about price, data retention, and overzealous content safeguards. Fable 5.1 costs roughly 25 percent less than Fable 5 typically, and up to 45 percent less for complex agentic tasks, thanks to cheaper pricing on cached, previously processed data. Fable 5.1 is now available on all platforms and cloud services, while Mythos 5.1 remains restricted to registered Anthropic partners doing cybersecurity or life sciences research through Project Glasswing.

Anthropic also announced Enterprise Frontier Safeguards, a high-privacy service rolling out this fall that stores data on customers’ own cloud servers rather than Anthropic’s, though the company will still monitor for misuse under terms clients control. Anthropic reiterated that it has never trained on enterprise data without explicit permission. The company is also now letting Fable 5.1 identify software vulnerabilities, though it will still route tasks like penetration testing, exploit generation, and binary-based vulnerability scanning to Opus models. Fable 5.1 also has “more precise safeguards” less likely to block basic biology questions than Fable 5, though Mythos 5.1 retains the same biology restrictions as its predecessor.

Trump Administration’s Blacklisting of Anthropic Was Illegal, Judge Rules

Related:

- Anthropic was illegally blacklisted by the Trump administration, court rules

Source

Judge Rita Lin of the Northern District of California ruled on Thursday, August 27 that the Trump administration’s blacklisting of Anthropic violated the First and Fifth Amendments, issuing a permanent injunction in a 59-page summary-judgment order and ordering the designation removed. The administration was denied a seven-day stay.

A recap: the dispute began when Defense Secretary Pete Hegseth tried to renegotiate AI labs’ military contracts to permit “any lawful use” of their systems. Anthropic alone refused, holding two red lines: mass surveillance of Americans and fully autonomous lethal weapons. After talks collapsed in February, Hegseth called Anthropic “sanctimonious,” Trump called it a “radical left, woke company,” and Trump posted on Truth Social ordering agencies to “IMMEDIATELY CEASE all use of Anthropic’s technology.” Hegseth then designated Anthropic a supply chain risk, a label normally reserved for foreign-adversary sabotage threats, and barred contractors from any commercial activity with the firm.

Lin found the government’s actions were unlawful First Amendment retaliation and a denial of Fifth Amendment due process, and called them arbitrary and capricious. The government’s “contemporaneous words and deeds,” she wrote, “confirm that the challenged actions were based on a desire to make a public example out of Anthropic for its ‘arrogance’ in criticizing the government, not based on any articulable basis to believe that Anthropic would actually sabotage its model.” She rejected the sabotage rationale as “entirely unfounded,” noting Anthropic cannot maintain backdoor access and that the government kept seeking to collaborate with it on advanced models, and wrote that “an IT vendor does not become a potential adversary of the United States whenever it asks probing questions.”

She also struck down Hegseth’s February order barring contractors from any commercial dealings with Anthropic, which had swept in non-military work despite the government’s own concession it wasn’t meant to. Lin had already flagged the actions as “troubling” and “Orwellian” in a March preliminary injunction. The Pentagon rested its blacklisting on two separate designations, which had to be challenged in two different courts; Lin’s order covers only the San Francisco case, and Anthropic’s parallel suit in Washington is still running. Until that one is resolved the company technically remains a supply chain risk, a status executives say is costing it business.

Other News

Tools

Google’s Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible. The update improves scene consistency by analyzing longer video segments, allows style transfer from reference footage, and introduces a cheaper draft mode that can be upscaled to higher resolutions.

Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more. The model performs more reasoning steps on complex tasks while maintaining the same per-token pricing as its predecessor, though Google warns it may consume more tokens overall and potentially increase costs for users.

Google now lets you chat with Gmail, Docs, and Keep. The features use real-time conversational AI to let you ask questions about your emails, format documents through natural language, and transcribe notes, with availability starting today on mobile for Google AI subscription tiers.

Google Pics is like Canva, but with even more AI. The tool integrates with Google Workspace apps to let users generate and edit images with precision controls, such as modifying specific objects or text within an image using text prompts.

Instagram cracks down on AI accounts pretending to be human. The platform will penalize AI accounts that fail to disclose their AI-generated profiles by reducing their reach in Reels and Explore recommendations.

ChatGPT Health adds Epic integration for clinicians to import patient data. Clinicians can now import patient data from Epic’s EHR system and use ChatGPT to summarize records, review patient history, and access medical research, while the integration maintains read-only access to ensure patient data safety.

Fei-Fei Li’s World Labs debuts Atlas, a world model showcase for advanced spatial intelligence. Atlas generates up to a minute of photorealistic 3D video from a single image with precise camera control, designed primarily for robotics simulation and training.

Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash. The model uses sparse and linear attention mechanisms to reduce computational costs while supporting up to 1 million input tokens and achieving competitive performance against leading competitors like Claude and GPT models.

Business

OpenAI’s ad business shows blistering growth, hits $1 billion annualized revenue run rate. The milestone comes roughly 200 days after OpenAI began testing ads in ChatGPT, which are now available across more than 40 countries with plans to expand formats and measurement capabilities.

Nvidia’s 70% growth forecast puts it on track to become tech’s No. 2 company by revenue. The chipmaker projects it will generate $673 billion in revenue next fiscal year, which would make it the second-largest U.S. tech company by revenue behind Amazon, though Huang indicated actual demand exceeds the 70% growth rate but is constrained by supply chain limitations.

OpenAI to end model access to Cursor after acquisition by Elon Musk’s SpaceX. The move follows SpaceX’s recent acquisition of the coding platform and stems from OpenAI’s concerns about contractual compliance based on past disputes with Musk’s companies.

Nvidia is buying Hugging Face for almost $13 billion. The acquisition consolidates Nvidia’s control over AI infrastructure by bringing the popular open-source model repository under its ownership, though Nvidia says the platform will remain open and developers won’t be required to use its chips.

Policy

Source

US government sides with OpenAI on issue of training LLMs on copyrighted material. The Trump administration has filed a brief supporting OpenAI’s legal defense that using copyrighted material to train AI models falls under fair use, arguing that restricting LLM development would harm American competitiveness in AI.

ChatGPT to face tougher regulation in the EU. OpenAI must now comply with the EU’s Digital Services Act by December 2026, requiring it to mitigate risks to minors and prevent illegal content while adhering to restrictions on targeted advertising and algorithmic transparency.

Concerns

Source

OpenAI, Anthropic, Google, and 100 other companies call for action to defend against rogue AI. The signatories are calling for increased collaboration between private companies and governments to develop new cybersecurity defenses against AI-enabled attacks, which they warn will become increasingly common as AI models grow more capable.

ChatGPT, Grok, and Claude all went down at the same time. Multiple AI chatbots including ChatGPT, Claude, and Grok experienced outages on Thursday morning, with each company citing different infrastructure issues before restoring services within a few hours.

Research

FrontierChallenge: Evaluating Scientific Workflow Completion. The benchmark evaluates whether AI agents can independently complete multi-stage scientific workflows across six domains by producing reproducible code, tables, figures, and reports, finding that current frontier models achieve only 20.6% full task completion despite scoring 87.9% on partial progress metrics.

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence. Researchers propose a framework for training reasoning models to improve themselves through autonomous feedback and self-generated training data, while establishing methods to identify potential risks along the scaling process.

A.I. Brings Big Gains to Hurricane Forecasts, Google Researchers Say - The New York Times. Google’s WeatherNext Cyclones model produces hurricane forecasts roughly a full day ahead of existing systems by training on global weather data combined with a specialized database of nearly 5,000 tropical cyclones, and is now being used operationally by the National Hurricane Center.

Language Models Can Control Their Own Attention. The approach uses chain-of-thought prompting to have models explicitly declare which tokens to attend to at each step, reducing computational costs by up to 52% on long-context tasks without requiring model retraining.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests. Researchers created an open benchmark that evaluates coding agents on realistic user requests—which are typically short and informal—and found that agent performance drops significantly compared to current benchmarks while revealing that desired behavior and motivation statements are the most valuable information users can provide.

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning. Researchers found that randomly discarding key-value cache entries performs as well as complex scoring methods for reasoning models, while being significantly faster since it eliminates the need for scoring calculations.

查看原始发布