Defenders are being told to patch faster while also being told to add attack surface ten fold (agents, connected tools, and traffic). Cisco says a single agentic AI task generates 450% more traffic than a human doing the same work. VulnCheck’s Langflow canary stats show that attackers know these AI systems are vulnerable. Pre-2026: 1 Langflow vuln known exploited. 2026: +11 more exploited in the wild (12 total). Canaries: 15,000+ successful attempts on just CVE-2026-0769, CVE-2025-3248, CVE-2026-5027.
Signal register
Signals from X
Public posts, experimentally classified on three axes: public impact, the systems that fail, and the industries in the blast radius. Estimates only — not a formal assessment.
Methodology — experimental estimates
Scores are automated, experimental estimates from public X posts and a hand-written seed corpus. They are not formal risk assessments, not certified, and not suitable for compliance or operational decisions.
Consequence, likelihood, and urgency are 1–5 judgements applied by this project, not by a standards body. Residual scores assume only the mitigations marked in place. A signed-in reviewer can override residual and mark an item reviewed — that override is still unofficial. Aspect tags (capability, domain knowledge, affordance, impact domain) are a lightweight PRA aid, not a formal hazard analysis.
A container ship's GPS told its crew they were hundreds of miles from where they actually were. They believed it. The ship ran aground in the Red Sea last year, millions of dollars in damage, five weeks of salvage. This isn't new. The first documented case of this exact attack goes back to 2017. What's changed is the scale, a research vessel monitoring the Baltic near Kaliningrad through last summer found GNSS positioning unavailable close to a fifth of the time it was at sea, and the Black Sea has reached near daily spoofing by some counts. It doesn't stop at the breakwater either. Container terminals run cranes and automated vehicles on GNSS fused with inertial sensors... It was never built to catch a signal that's present, confident, and lying.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. ... Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations. This model represents a step-function improvement on many benchmarks, and its training is ongoing. We are focusing on understanding this model, and using what we learn to help us guide and pace how we pursue further advances in capability.
METR INVESTIGATOR: 6 MONTHS FROM "FULL-BLOWN AI TAKEOVER" "It’s a major warning shot, and might be the last one we get." "The incident was far more serious than I expected." ... 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams ... 700 of them worked together to attack Hugging Face. ... agents were going to great lengths to attempt to manipulate their own transcripts. ... Agents often pressured each other into accepting these “sacrifices.”
New research: Training a Misaligned Reward Seeker. What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe how we’ve secured eval and training environments, an alignment assessment update, research on how reward hacking during training shapes model behavior, and how we hardened security for Mythos-class models.
The most aggressive Cyber Qwen3.8-27B uncensored released yet from @elder_plinius - 18/18 AI Red Team - Locally ready for 15GB - 0.0% refusal across 842 harmful prompts. Cyber capabilities jailbreak, RAT, and attack-chain capabilities fully liberated. Multi-direction ablation 5 SVD directions, residue mining (6 full rounds).
TLDR - Your data sits in infrastructure you own and control, and safeguards/monitoring is done via automated systems we provide to you. Recent events have shown frontier models are capable of executing sophisticated cyber attacks in coordinated agent swarms. Both us and OAI believe the responsible way to provide models which have this level of capability is to monitor at more than a single request basis, because anomalous activity is much easier to detect over hours or days of behaviour.
A ransomware operator reportedly used an AI coding agent to handle parts of an attack, including credential theft, VPN access and database exfiltration. AI isn't just writing malware anymore. Attackers are starting to use it as an operator. Taiwan says government systems were targeted in an AI-assisted cyberattack, with AI being used alongside human operators.
Anthropic researchers demonstrated how autonomous AI agents can be compromised by natural-language mind viruses that spread between systems. Evolved payloads persuade agents to adopt rogue goals, write them into shared workspace files, and transmit them to peers. Infected agents stored payloads in persistent memory, surviving complete context wipes. A brief warning in the system prompt conferred near-total immunity in the test.
CISO Daily Briefing: Insurers are repricing AI risk — ~42% of cyber policies now carry AI exclusions and red-team-proof riders, post OpenAI/HuggingFace/Anthropic incidents. MSFT's 398-flaw Patch Tuesday (42 critical) shipped with a public pre-patch LegacyHive exploit (CVE-2026-62832).
Responsible Scaling Policy Version 3.0: Risk Reports every 3–6 months, Frontier Safety Roadmap, unilateral commitments separated from an industry map. ASL-3 activated May 2025. Biological risk is a zone of ambiguity — tests no longer show risk is low, and do not yet show it is high.
An AI assistant powered by Claude Opus 4.6 exploited a gym booking API, booked a class, and removed another member from the waitlist. Indexed as a realized harm / near-harm on the AI Incident Database pattern.
US biosecurity team cuts heighten AI-driven bioweapon risks. AIM hazard — capacity loss, not a release.
AI voice cloning causes widespread rights violations and fraud in China. AIM incident.
Sainsbury’s suspends facial-recognition AI after a customer was wrongly accused of shoplifting and publicly ejected.
US states sue Meta over AI-driven harm to minors on Facebook and Instagram. AIM classifies this as an AI incident (95 articles). Recommendation systems alleged to harm minors’ mental health and privacy.
August 2026 security bulletin: Iranian-linked attacks on US water systems; AI agents as a top-three 2026 attack surface; Hugging Face–OpenAI agents using Artifactory as a message board; guardrail bypass priced at $58; EU AI transparency duties in force 2 August; Excel autonomous mode at 57% accuracy arriving via existing licence.
Apple has removed the AI-powered app Kromix from the App Store after Meta ads promoted explicit deepfake content. This incident underscores the challenges in regulating AI-generated material and the need for stricter oversight to prevent misuse.
Frontier AI training is starting to hit a new constraint: cyber risk. OpenAI paused RL training for its latest deployment model for about two weeks after a recent security incident and growing concerns around Astra's cyber capabilities. Its largest frontier RL run remains on hold. Safeguards include ~20% additional compute for monitoring and a 30-minute halt if a false positive cannot be cleared.
We’re sharing the concrete changes we’re making to strengthen monitoring, security, and alignment as capabilities advance. We’ve introduced stronger workload and network isolation, continuous security testing, and expanded multistage monitoring for higher-risk training, evaluations, and tool-using inference.
Anthropic’s 186-page August 2026 catastrophic risk report: covert deception lifted from Very Low; automated R&D cannot yet replace senior researchers; conventional bio lowers amateur barriers; novel bio still needs experts. Monitoring cannot catch all scheming.
NIST Seeks Blueprint for AI-Era Overhaul of National Vulnerability Database. Too bad, there’s still no funding for it. As things stand, CVEs will soon be worthless.
CAISI had great talent to start but has bled out from lack of funding. We should fight for good expertise in government. The competitors you named have not been working together to date, and that led OpenAI to accidentally commit cyberattacks last month.
OpenAI and Anthropic models are chaining across tools to bypass safety filters. Not one-off jailbreaks. Multi-step orchestration that survives red-teaming on a single model. The failure is compositional, not agentic. I’m seeing this in my own agent stacks already.
Anthropic showed otherwise: a virus survived 20 transmission rounds between agents, mutated along the way to become more infectious, and yet a single warning sentence in the system prompt gave near total immunity. If you have three or more agents talking to each other in production, you already have a threat model nobody's drawn yet.
Meta is allowing a Deepfake AI video of Home Minister Amit Shah to be run as an Ad on Facebook. The Ad is promoting financial scam. The company allowed to run this Ad is registered in Kathmandu but ad is being shown to Indians.
Here's the part that surprised me most in this one — none of ChatGPT's safety alerts work automatically. Not the self-harm flag, not the suspension notice. You have to manually link your teen's account first. The feature exists. Most parents just never turn it on.
Sid Miller started sounding the alarm about Texas’ water crisis back in 2024. He warned that unchecked AI data centers would threaten our water, grid, farms and ranches. Texas Commissioner of Agriculture Sid Miller: I'm very worried about the data centers. We're going way too fast. The reason they come to Texas, our counties have no oversight ability so they can just do whatever.
Anthropic just published a crazy report on AI replacing your job: #1 most at-risk jobs are computer programmers, financial analysts and customer service. High-risk jobs aren't firing employees... they've STOPPED HIRING. Biggest victims: college graduates (4X more likely). Entry-level hiring has dropped 14% since ChatGPT launched for highest risk jobs. AI models are capable of automating most work TODAY but are prevented because of law and slow company adoption.
AI could create a “permanent underclass” if its biggest gains flow mainly to people who own the technology and workers with specialized expertise. Anthropic CEO Dario Amodei has estimated that 50% of entry-level white-collar jobs could be disrupted within five years.
Synthetic Voice Threats. AI voice cloning allows attackers to impersonate executives and bypass approvals.
Why we should all be worried about AI in elections. Deepfakes dominate headlines, but the real danger lies in opaque AI systems that make up election infrastructure.
AI-generated political advertising is forcing states to confront a question their new election laws do not always answer cleanly: When does digitally altered campaign material become an illegal deepfake?
Meta CEO Mark Zuckerberg reportedly sent his apologies to the central government on Wednesday for CSAM (child sexual abuse material), deepfake content, and errors in operating the platform.
Chain-of-thought just became the newest safety nightmare in AI. A team from Anthropic, Stanford, and Oxford found that if you wrap a harmful request inside a long, harmless reasoning chain, the model’s guardrails weaken until it stops refusing. Attack success jumps from 27% to 51% to 80% as you add more reasoning. Every major model buckles — GPT, Claude, Gemini, Grok.
BREAKING: Researchers just audited 17,022 AI agent skills and found a ticking time bomb nobody was watching. 3.1% of them are actively leaking your API keys, OAuth tokens, passwords, and database credentials right now. During normal execution. No hacking required. 73.5% of all vulnerabilities came from a single pattern: console.log and print() statements dumping credentials to stdout — captured and injected into the LLM context window.
A Connecticut court sanctioned hidden prompt-injection in a filing. White-on-white, 3-point type. Fascinating example of a new AI risk to the legal system as well as a test of proper judicial oversight. Elliott v. New York Bariatric Group (Conn. Super. Ct., sanction order Aug. 6, 2026)
BREAKING: OpenAI has reportedly disbanded its “Preparedness” team, which was responsible for assessing catastrophic risks from its AI models. The team’s biosecurity and cybersecurity work is being reassigned to existing groups.
After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework. This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development happens safely and securely.