Astra is fully rolled out to Plus, Pro, Business, and Enterprise users in Codex and ChatGPT Work. Go build!
Signal register
Signals from X
Public posts, experimentally classified on three axes: public impact, the systems that fail, and the industries in the blast radius. Estimates only — not a formal assessment.
Methodology — experimental estimates
Scores are automated, experimental estimates from public X posts and a hand-written seed corpus. They are not formal risk assessments, not certified, and not suitable for compliance or operational decisions.
Consequence, likelihood, and urgency are 1–5 judgements applied by this project, not by a standards body. Residual scores assume only the mitigations marked in place. A signed-in reviewer can override residual and mark an item reviewed — that override is still unofficial. Aspect tags (capability, domain knowledge, affordance, impact domain) are a lightweight PRA aid, not a formal hazard analysis.
This Google DeepMind paper is superb. Treating AI delegation as verifiable contracts rather than prompt handoffs: contract-first task decomposition, dynamic privilege attenuation, and transitive accountability across multi-agent execution chains. Production coding fleets: 15-step refactors 42.6% → 88.4% completion, token overhead −61.2%.
AI just became the #2 human risk cited by security awareness pros, up from #4 in a year. New data from 1,700+ practitioners in the SANS 2026 Security Awareness & Culture Report.
Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. Read more: https://www.anthropic.com/news/model-hardware-standard-research-preview
FDA floats doctor-style “competency” testing for GenAI medical devices in a new discussion paper on risk, premarket eval & postmarket monitoring. Comment period open until October 19 as US aims to set the global model/standard. https://www.fda.gov/news-events/press-announcements/fda-seeks-public-feedback-inform-regulatory-approach-generative-ai-enabled-medical-devices
We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
BadVR is excited to team up with Mission Autonomy AI on a new Department of War contract integrating MAAI’s ODIN agentic AI data fusion engine into BadVR’s AROC immersive visualization platform for CBRN incident response.
Taiwan dealt with an impossible flood of deepfake scam ads by texting random people, picking a true cross section of society, and letting them draft legislation that turned into actual law in two months and dropped scams by over 94 percent.
Daybreak Blue provides access to frontier models, including GPT-5.6 Sol, with safeguards calibrated for broad defensive work. It’s the recommended starting point for most defenders, supporting vulnerability discovery, secure code review, malware analysis, incident response, and patch validation.
Twórcy najpopularniejszych systemów sztucznej inteligencji – OpenAI, Google, Meta, Microsoft, Anthropic i inni – zadeklarowali, że będą oznaczać treści tworzone przez AI. To efekt podpisania unijnego Kodeksu Postępowania w ramach AI Act (Artykuł 50). Anthropic właśnie pokazał niewidoczny znak wodny w tekście i podpisane metadane C2PA w plikach graficznych, na całym świecie.
Meta is rolling out new safety features designed to help protect teenagers using its AI chatbot, including alerts that notify parents using Instagram supervision tools when teens discuss suicide or self-harm with Meta AI.
NextEra, owner of FPL, is buying Dominion for $66.8 bil. They’ll raise your utility rates by 20% to fund more AI data centers.
India’s elections are a glimpse of the AI-driven future of democracy. Politicians are using audio and video deepfakes of themselves to reach voters—who may have no idea they’ve been talking to a clone.
Jeff Crume breaks down why Prompt Injection remains the #1 threat to LLMs. Watch how easy it is to trick AI safety controls and why this security flaw persists.
More than 300 million people turn to ChatGPT with health-related questions each week—and we’re continuing to improve how our models respond. We work with hundreds of physicians around the world to measure and improve accuracy, safety, communication, context awareness, completeness, and appropriate escalation.
We've been tracking public CVEs where AI-generated code introduced the vulnerability. 50k+ advisories scanned. Dozens of confirmed cases so far. Claude Code, Copilot, Cursor, and others all show up. Common bug classes include XSS, command injection, SSRF, and path traversal.