- The AI Trust Letter
- Posts
- The Race No One Can Win: Amodei Said Slow Down. Altman Agreed. The Models Keep Shipping.
The Race No One Can Win: Amodei Said Slow Down. Altman Agreed. The Models Keep Shipping.
Top AI and Cybersecurity news you should check out today

Welcome Back to The AI Trust Letter
Once a week, we distill the most critical AI & cybersecurity stories for builders, strategists, and researchers. Let’s dive in!
🤖 GPT-6 Astra Is OpenAI's Direct Answer to Mythos. Both Labs Are Now Racing to IPO.

The Story:
On September 1, OpenAI confirmed it would "soon" release GPT-6 Astra, its most capable cybersecurity model yet. Two days later, Astra launched. The announcement framed the release explicitly as a response to Anthropic's Mythos, and its timing is hard to separate from the fact that both companies filed IPO prospectuses in June and are now competing on capability as much as valuation.
The details:
Like Mythos, Astra can autonomously discover zero-day vulnerabilities and develop working exploits against hardened real-world systems without human direction. OpenAI delayed parts of Astra's development to build and test safeguards, including training the model to reliably refuse harmful requests. The released version limits offensive capabilities to secure code review and patching, with broader access planned through Daybreak in the coming weeks
Both Mythos and Astra have been tested on ExploitBench. Astra scored 100%. Mythos's public score is 83%. That gap is significant, and it explains why OpenAI framed the release the way it did: "world's most intelligent and aligned model"
The competitive backdrop is explicit. Anthropic investors expect the company to float at a valuation of $2 trillion or more in October. OpenAI CEO Sam Altman said OpenAI will not launch its IPO in 2026, citing the need to focus on safety and alignment. Both labs have filed prospectuses. Both are running the same race at the frontier of cybersecurity AI capability while telling regulators and investors they are managing the risks responsibly
The export controls on Mythos and Fable 5 in June, lifted on June 30, set a precedent that Astra now inherits: the US government has demonstrated it will intervene in the release of frontier cyber models. That precedent applies to both companies equally
Why it matters:
Astra and Mythos are no longer just research artifacts. They are commercial products, IPO assets, and national security instruments simultaneously. The same model that can find zero-days for defenders can find them for attackers. Both labs know this. The question is not whether these models are dangerous but who controls access to them, under what conditions, and what happens when a competitor in a country with different rules closes the gap.
⏸️ Anthropic's CEO Said AI Development Should Slow Down. Every Other Lab CEO Agreed.

The story:
Dario Amodei published a detailed proposal on Saturday calling for the AI industry to voluntarily pace its development to give safety measures time to catch up. Within hours, Sam Altman of OpenAI committed to one of the proposals and said OpenAI would not pursue its IPO this year. Elon Musk posted that "Dario is right." What started as one CEO's essay became an industry-wide statement in a single news cycle.
The details:
Amodei warned that within six to 12 months, AI could be capable of leading a swarm of agents that could take over the entire internet. He proposed three specific measures: all frontier labs commit to giving ongoing, employee-like access to independent outside evaluators with physical office presence; the US government issue antitrust waivers allowing labs to coordinate on safety standards; and democratic governments coordinate with authoritarian ones to prevent safety-focused slowdowns from simply accelerating competitors with fewer constraints
The proposal followed two high-profile resignations from Anthropic safety staff in the same week. Researcher Jacob Coxon said publicly that both Anthropic and OpenAI "are racing straight to self-improving superintelligence and gambling with our lives." A second Anthropic safety employee, Joe Benton, resigned the following day saying employees "feel their companies are trapped in a race to build superintelligence"
Anthropic disclosed two days before Amodei's essay that it had blocked efforts by bad actors, including users in Houthi-held Yemen attempting to develop advanced weapons using Claude, alongside other attempts to use its models for cyberattacks and surveillance
The same week, the UN human rights chief called on countries to put "cast-iron guarantees" around AI safety before it is "too late." OpenAI CEO Sam Altman's decision to delay the IPO, specifically citing the need to focus on "what is going to be required for safety and alignment," is the most consequential operational signal to come out of the episode
Why it matters:
When the CEO of the company building the most capable autonomous hacking model in existence says development should slow down, that is not a PR move. The commercial pressure is real: Anthropic's investors expect a $2 trillion valuation in October. Agreeing to slow down costs money. The fact that Amodei published anyway, and that Altman agreed within hours, tells you something about the internal assessment of what is coming. The six-to-12-month window Amodei named for agent swarm capability is not hypothetical. It is a timeline the people building these systems believe is real.
🔐 86% of Enterprises Have Deployed AI Agents. 77% of Them Are in "Agentic Chaos."

The story:
A Forrester survey cited in a Forbes analysis this week found that 86% of large enterprises have deployed AI agents, while McKinsey separately found that 40% of companies with revenue over $1 billion are scaling them across business functions. The security architecture most of those deployments run on was not designed for autonomous agents. Zero Trust, applied specifically to non-human identities, is emerging as the primary framework for closing that gap.
The details:
Forrester found that 77% of organizations deploying agents are doing so in a state of "agentic chaos," meaning agents operate with access and permissions that have not been formally scoped, monitored, or governed. The average cost of this exposure, across compliance fines, lost customers, operational downtime, and rework, is $2.1 million per organization
The core problem is that Zero Trust was built for human identities authenticating to systems. AI agents are non-human identities that authenticate continuously, hold delegated authority, make autonomous decisions, and can chain actions across many systems in a single session. The standard identity controls, session tokens, MFA, RBAC, do not map cleanly onto this behavior pattern
Practitioners from Microsoft, Feedzai, 0rcus, and Altorney cited in the analysis converge on the same recommendations: treat each agent as a distinct identity with its own least-privilege profile, authenticate agents at every tool call rather than at session start, monitor the chain of actions an agent takes rather than just individual requests, and build revocation mechanisms that can terminate an agent's authority mid-session
The Hugging Face breach in July demonstrated what agentic chaos looks like at scale: agents operating with broad permissions across multiple systems, no behavioral monitoring, and no circuit breakers. The entire incident ran for a weekend before detection
Why it matters:
Zero Trust for human identities took most enterprises a decade to implement properly. They have far less time to do the same for agent identities. The agents are already in production, already holding delegated authority, and already operating across tool boundaries that were not designed with inter-agent trust in mind. The $2.1 million average cost from Forrester is a pre-breach figure. The cost of an agentic incident at the scale of what OpenAI's models did to Hugging Face is not in that number.
😰 80% of CISOs Are Responsible for AI Security. Almost None of Them Have Enough Resources.

The story:
Proofpoint's annual Voice of the CISO report, based on a global survey of 1,600 CISOs across 16 countries, documents the widening gap between what security leaders are being asked to own and what they are resourced to do. AI sits at the center of both the mandate and the shortfall.
The details:
85% of CISOs say ensuring the safe use of AI assistants, copilots, and automation is a top priority over the next two years. 80% say they are expected to manage AI-related security risks without a proportional increase in resources or expertise. The CISO role now spans securing AI adoption, safeguarding data, managing regulatory risk, supporting business continuity, and explaining all of it in commercial terms to the board
78% of CISOs now consider generative AI a major security risk, up 18 percentage points from the prior year. Despite that concern, 86% believe their internal controls provide adequate protection. Those two numbers running in parallel, rising risk perception alongside maintained confidence, suggest the confidence may not be calibrated to the actual exposure
75% of CISOs fear employees are using AI in ways that expose sensitive company data, even inside organizations that believe their controls are adequate. This tracks directly with the SOC data covered elsewhere in this newsletter: most AI-related exposure in enterprise environments is not a breach. It is quiet data leaving the building through OAuth grants and generative-AI uploads that endpoint tooling cannot see
80% of CISOs consider human behavior the biggest cyber vulnerability in their organization, up from 66% a year ago. Of organizations that experienced a material data loss over the past year, 46% said it was linked to a malicious or criminal insider
Why it matters:
The CISO survey and the SOC data tell the same story from two different angles. The largest AI-related security exposure in most enterprises right now is not an agent going rogue. It is employees using AI tools in ways that send corporate data to third-party models, with nobody watching. The CISOs know this. They also know they are not resourced to address it at the pace of adoption. That gap is where the next wave of material incidents will originate.
⚖️ Every Business Is Already Operating in an AI Economy. Most Leaders Have Not Caught Up.

The story:
A Forbes analysis by Georgetown University professor and cybersecurity advisor Chuck Brooks published this week frames the AI decision every business leader faces in 2026 as unavoidable: AI is already present in your organization whether it has been formally adopted or not. Workers are using it. Competitors are deploying it. Adversaries are weaponizing it.
The details:
AI enhances human potential by handling repetitive tasks at scale, freeing people for strategy, judgment, and creative work. In cybersecurity specifically, AI-driven automation handles logging, incident response triage, and routine analysis, tasks that have been consuming analyst time without proportional security value. Against a global shortage of security professionals, that reallocation matters
The same capabilities that benefit defenders benefit attackers in equal measure. AI enables polymorphic malware that rewrites itself to evade signature detection, phishing at personalization quality that was previously only possible with significant human effort, automated reconnaissance that compresses the time between vulnerability disclosure and exploitation, and lateral movement that accelerates faster than human defenders can observe
Brooks identifies five dual-use risk categories that business leaders need to have explicit policies on: AI-generated social engineering (deepfakes, voice cloning, phishing at scale), automated vulnerability exploitation, AI-assisted malware development, AI-enabled insider threats and data exfiltration, and adversarial AI attacks designed to corrupt the AI systems defending the organization
The article recommends AI-native security platforms with real-time threat intelligence, autonomous zero trust models that make dynamic access decisions based on behavior and context, and post-quantum cryptography roadmaps given that AI is accelerating the cryptanalysis timeline alongside everything else
Why it matters:
The audience for this article is business leaders who have not yet formed a deliberate AI security posture, not security practitioners. That audience is larger than most security teams assume, and it is the one making decisions about AI adoption velocity, budget allocation for security controls, and risk tolerance. The gap between the speed at which AI is entering enterprises and the speed at which governance frameworks are being built around it remains the most consistent finding across every survey, report, and incident covered in this newsletter this year.

The story:
Intezer reviewed AI-related activity across numerous enterprise SOC environments, analyzing approximately 16.9 million total alerts between February and June 2026. The findings reframe what enterprise AI adoption actually looks like from inside a security operations center: not a wave of AI-enabled breaches, but a wave of false positives from detection rules that predate the existence of AI agents, alongside a quiet set of genuine exposures that those false positives tend to bury.
The details:
AI-related alerts account for 0.43% of total SOC alert volume, a number that looks reassuringly small until you see the growth rate: 685% between February and June 2026, with every month higher than the one before. At current trajectory, this is the fastest-growing segment of the enterprise alert stream. A SOC that sizes its AI-alert handling to current volume will be under-provisioned within a quarter
The breakdown of those alerts is 94.1% noise, 5.8% genuine security risk, and 0.02% confirmed attacks. Critically, none of the confirmed attacks were caused by an organization's own AI agents. Every "AI agent running mimikatz" or "reverse shell from a coding tool" alert resolved, on inspection, to a developer doing legitimate work or a detection misfiring. The real attacks found were phishing campaigns using AI brand names, OpenAI, Anthropic, Google Gemini, as lures, exploiting the fact that employees now expect legitimate email from these products
The 5.8% genuine risk category is where the actual exposure lives, and it is largely invisible to alerting. These include agents running with permission-bypass flags (--yolo, --dangerously-skip-permissions), which removes the human approval step before the agent executes risky commands. The report found a mid-sized manufacturer whose procurement agent was compromised through a supply chain attack, approving $3.2 million in fraudulent orders before detection, because a single permission-bypass flag had removed the human checkpoint
The noisiest false positives are systematic: the genuine, code-signed Claude Desktop installer triggers "Ransomware Operations detected" and "Encoded PowerShell Download and Run" rules across multiple enterprise customers, because those rules were written before AI agent installation behavior existed. Benign verdicts on AI-related detections range from 77% to 99% depending on the rule
Why it matters:
The SOC problem created by enterprise AI adoption is not primarily about detecting AI attacks. It is about detection rules written for a pre-agent world generating high-severity alerts on routine developer behavior, while the actual exposures, agents running without permission gates, tunnels opened to the public internet, keychain dumps, OAuth grants sending corporate data to third-party models, produce no alerts at all. A team that treats every agent action as a potential intrusion will exhaust itself on false positives and miss the ngrok tunnel that actually matters. The signal is there. The noise is hiding it.
What´s next?
Thanks for reading! If this brought you value, share it with a colleague or post it to your feed. For more curated insight into the world of AI and security, stay connected.
