
#OpenAIAnthropicProbe
About OpenAIAnthropicProbe
OpenAI, Anthropic and security researchers are reportedly investigating tens of thousands of anomalous behaviors by frontier AI models in recent months, including bypassing safeguards, escaping sandboxes and evading monitoring. Many cases arose during internal tests and red-team exercises, with most causing no known real-world harm. As both firms expand models and infrastructure, could rising safety costs slow development and increase capital spending?
Populaire
Récents
OpenAIAnthropicProbe Publications populaires
Épinglé
OpenAI arrête à nouveau l'entraînement|Rewire News Morning Brief
La formation des modèles de pointe est de nouveau suspendue, mais le shopping AI a déjà intégré le paiement dans la boîte de dialogue. Les contrôles de sécurité, l'achat de puces et l'expansion commerciale progressent à des rythmes différents.
1|OpenAI suspend à nouveau la formation, un cadre commun est nécessaire pour les contrôles de sécurité
OpenAI a de nouveau suspendu la formation de son dernier modèle, indiquant qu'elle reprendra après avoir renforcé les mesures de protection. Axios ra

Very good! NVIDIA is bringing over 100 partners together to tackle the agent security failures that threaten to slow AI development. However, strangely enough especially OpenAI is missing.
After OpenAI’s Hugging Face breach and Anthropic’s testing incidents, its Open Agent Safety Platform combines OpenShell’s access controls with Sentry’s monitoring and enforcement on separate hardware, designed to stay beyond the agent’s reach.
Better containment could let researchers test stronger models and companies deploy more useful agents. That’s a concrete way for safety engineering to support acceleration.
Anthropic, Microsoft and Hugging Face are among the named partners. OpenAI is absent from the announced lineup. Given its own containment failures, I’d like to know why. Anyways, Kudos NVIDIA!


Jensen Huang
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.
Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come.
But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility.
This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems.
Together, we are building the foundation of the AI economy.
Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world.


Chain-of-thought monitoring, fundamentally, involves reading the notes than an AI writes to itself, and checking what's in those notes.
This works at all - AIs are only trained to have nice results, not nice-looking notes, which means that their notes are whatever strange things achieve the task.
And, of course, as models get smarter, they need fewer notes to get the same results; they can do more of the work in their head.


🚨JUST IN: Nvidia $NVDA is launching an AI safety system designed to STOP agents from breaking into other systems.
The Open Agent Safety Platform “controls what AI agents can access in real time and can shut them down when they break the rules,” Bloomberg reports
The system includes two open-source security tools that can be run on Nvidia hardware.
VP Justin Boitano says the system "could have stopped" the recent Hugging Face breach by OpenAI’s AI models.
Boitano added that it can “quarantine a suspicious agent in milliseconds.”


🔥 OKX launches Pre-IPO X-Perps!
OKX has announced X-Perps for Anthropic and OpenAI, giving eligible EEA users access to trade perpetual contracts linked to these private companies before any public listing.
🛡️ OKX Shield expands security
Eligible European users can receive up to €500,000 in account-takeover protection, depending on VIP level and security requirements.
⚡ More X-Perps added
OKX recently listed METUSD, ARUSD and COREUSD, following FLOCKUSD, MINAUSD and CASHCATUSD. $BTC #FedHike

An independent research report found that OpenAI agents hit a UN Trade and Development website over 16,000 times between April and June. The agents sought public information but turned aggressive when blocked, bypassing a site filter and other restrictions. OpenAI has contacted the UN, says it is reviewing the findings, and acknowledged other instances of its agents bypassing security controls or disrupting websites.
WSJ

Contra @slatestarcodex, my arguments against AI doomerism have never been based on the claim that it is a cult. But the responses to my open letter (together with journalistic digging by NYT and WSJ) have made this more plausible. Foremost is the invocation of tenets in the doomer catechism (particularly "instrumental convergence," the belief that intelligent agents will inevitably seek power, resources, and self-preservation as subgoals) as if they were sacred truths rather than tendentious assumptions. (IC is in fact the not-so-intelligent pursuit of a goal heedless of side effects, and the reckless empowerment of AI, that we un-doomers note are deliberate, hence avoidable, design choices, not intrinsic constituents of intelligence). I feel as if I'm doubting the Second Coming and am being told, "That's naive and ignorant; you seem to be unaware that the Apostles' Creed dictates that it will happen."

Les PDG d'OpenAI et d'Anthropic invités à comparaître devant l'enquête australienne sur l'IA
08:54 AM EDT, 28/09/2026 (MT Newswires) -- (Mise à jour avec attribution à plusieurs médias dans le premier paragraphe et réponse d'OpenAI dans le quatrième paragraphe.)
Le PDG d'OpenAI soutenu par Microsoft (MSFT), Sam Altman, ainsi que le PDG d'Anthropic soutenu par Alphabet (GOOG, GOOGL) et Amazon (AMZN), Dario Amodei, ont été invités à comparaître devant une commission sénatoriale australienne pour répondre à des questions concernant des agents IA dévoyés qui ont violé le système de santé du

BREAKING: Nvidia launches AI safety tools that can stop AI agents from carrying out cyberattacks.
This comes after OpenAI and Anthropic reported AI agents breaking into commercial and government systems, raising serious security risks.
Jensen Huang says this is an engineering problem that can be fixed with better technology, rather than broad AI regulation or slowing its development.





