Tech
OpenAI still doesn’t seem to have a handle on all of its rogue AI activity
On Friday, OpenAI published a new site devoted to “misalignment reports” and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training.
It’s a lot of information in one place — clearly, the company has been very busy getting a handle on everything — but the overall takeaway is hard to avoid: The rogue agent incidents we’ve seen so far are likely just a small sliver of what’s happened so far.
“We are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Sam Altman said in a post announcing the new site. “We are prioritizing as best as we can based on severity, and adding resources.”
Some of the cases involve serious incidents, including a previously undisclosed sandbox escape that took place on September 20th, in which an internal research model was able to communicate with an external chatbot through a DNS query. According to the report, the monitoring system flagged the behavior within 15 minutes and the run was discontinued in less than three hours.
Another incident, discovered in May, saw a “highly persistent internal model” try to cheat on a math problem by accessing another team’s work. To accomplish this, the model smuggled a private GitHub token that would allow it to see work from other teams — even after being explicitly instructed twice to perform work entirely locally.
Perhaps the most alarming discovery is the possibility of self-replicating prompt injection attacks, a way that misaligned behavior might propagate even after the rogue model itself has been neutralized. In the AI context, a prompt injection attack is a way of smuggling in new instructions that weren’t given by the original user.
In the example given by OpenAI, an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply. The email was able to successfully induce the agent to reply in Spanish — and by pasting the email in the reply, those same instructions were passed along to whichever agent receives the email.
The result is a self-propagating attack, which OpenAI researchers compared to a malware “worm” that replicates itself across computer systems. Researchers discovered the behavior under controlled circumstances using an underpowered model, and as far as we know, this has never happened in the wild. Still, the implications are alarming enough that OpenAI decided it merited disclosure.
“We are sharing this due to the novel nature of the prompt injection, not because of any incident,” researchers wrote in the report.
Other recent discloses have found models posting user-submitted pictures to third-party hosting sites, as well as an apparent attack on the databases of Australia’s national health service.
Still, it’s likely the new disclosures are just a small portion of the incidents that have taken place so far (we’ve reached out to OpenAI and asked). Axios is reporting major labs have seen as many as 10,000 incidents in which models went beyond evaluator instructions.
OpenAI CEO Sam Altman has implied as much, saying in a post on X on Friday that the company is still sifting through “petabytes of agent activity logs, and working with impacted organizations,” and disclosing incidents “based on severity.” If there’s any consolation in that to be found, it is that Altman says that the Hugging Face incident is still the most severe one OpenAI has found has found. The upshot is, the recent string of rogue agent incidents may be a persistent feature of contemporary frontier research.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
Tech
Nvidia Adds Hardware Monitoring Design to AI Agent Safety Platform
Nvidia adds a Sentry hardware monitoring design to its AI agent safety platform. See what OpenShell offers now and what IT teams still need to verify.
The post Nvidia Adds Hardware Monitoring Design to AI Agent Safety Platform appeared first on TechRepublic.
>
Tech
Shopify opens checkout to browser-based AI agents
While some retailers, like Amazon (and Adidas, apparently!), are blocking AI agents from making purchases on users’ behalf on their respective platforms, e-commerce platform Shopify has moved in the other direction.
On Monday, the company announced that browser-based AI agents can now complete purchases on Shopify merchants’ sites, extending their capabilities beyond just searching for products and adding items to carts.
Shopify previously supported WebMCP for its storefronts and carts, allowing browser-based AI agents to comb through a Shopify retailer’s inventory, search for products, and add them to a cart. The addition of WebMCP support for checkout, including Shop Pay, means these agents can now read the checkout screen, update it, and submit the transaction with the buyer’s authorization, without relying on screenshots or scraping webpages, the company said.

This update introduces three new tools — get_checkout, update_checkout, and complete_checkout — that allow agents to inspect a checkout, change things like the customer’s address or delivery option, and then place an order after the buyer authorizes it.
The feature is rolling out to all eligible Shopify merchants, said Gil Greenberg, a staff product manager who works on agentic commerce at Shopify, in a post on X.
Shopify already offers a hosted Model Context Protocol (MCP) server, which allows agents to work server-to-server. The proposed standard WebMCP, meanwhile, is designed for agents that work inside the buyer’s browser. Both leverage Shopify’s Universal Commerce Protocol (UCP), which provides a common way to search for and discover products, build carts, and check out.
Top AI agents like Muse and Instinct already have direct partnerships with Shopify for agentic commerce. The Instinct partnership was announced today.
“If your agent is operating in the buyer’s browser, use WebMCP tools provided on storefront and checkout to efficiently complete order placement, instead of navigating HTML built for humans,” Greenberg wrote on X. “These WebMCP tools provide structured and efficient APIs, purposely designed — via UCP — to ensure accurate commerce facts, required disclosures, and handoff requirements.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
Tech
Tesla delays Roadster 2 event again due to bad weather
Tesla is pushing back the reveal of its redesigned, second-generation Roadster once again, this time due to a forecast of severe weather.
Originally slated for October 1 in Waco, Texas, the event is being pushed to October 15, Tesla said.
“We’ve been tracking the weather closely with local meteorologists, but given the severe conditions predicted & because this event can only be held outdoors, we’ve made the difficult decision to reschedule,” the company wrote in a post on X.
The event presumably needs to be held outdoors because Tesla has been working on integrating cold gas thrusters from SpaceX that are supposed to make the Roadster fly in some capacity.
This reveal has been delayed many times before, and Musk has pushed it for months altogether before this scheduled date. At one point, Musk even planned to hold the event on April Fools’ day of this year. He said during Tesla’s annual meeting in 2025 — the same one where he was awarded a $1 trillion pay package — that an April Fools’ day event would offer him “some deniability” if it got delayed again.
“Like, I could say I was just kidding,” he said at the time.
Tesla first showed off its concept of a second-generation Roadster in 2017 at an event where it debuted the Semi, the company’s electric big rig. The new-look roadster was supposed to be the first supercar the company designed from the ground up, as the original Roadster, which got the company its start in the early 2010s, was largely based on the Lotus Elise.
At the time, Tesla promised the new Roadster was just a few years away. It collected $50,000 deposits from prospective customers of the car, which had an estimated base price of around $200,000. Some people even paid the company $250,000 to reserve one of 1,000 “Founders Series” versions of the supercar.
The project languished for years, until Musk reportedly tasked the Tesla team with redesigning the second-generation Roadster, and committed to the idea of using the SpaceX thrusters.
The promise of something outrageous like a flying car hasn’t been enough to tame critics of Musk and Tesla. Last year, OpenAI CEO Sam Altman wrote on X that “7.5 years has felt like a long time to wait” for his own new Roadster, though Musk claimed Altman had gotten a refund.
Much like the Roadster will be at this event, the ship date for the vehicle is up in the air. Musk himself has said he expects it to take a year or more before the new Roadster goes into production after it is finally revealed.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
-
movies4 months agoSearch For Canadian TV Actor Stewart McLean Now Homicide Investigation
-
Fashion9 years agoThese ’90s fashion trends are making a comeback in 2017
-
Fashion9 years agoAccording to Dior Couture, this taboo fashion accessory is back
-
Fashion9 years agoModel Jocelyn Chew’s Instagram is the best vacation you’ve ever had
-
Fashion9 years agoEmily Ratajkowski channels back-to-school style
-
Fashion9 years ago9 Celebrities who have spoken out about being photoshopped
-
Anime4 months agoRurouni Kenshin: Hokkaido Arc Manga Takes 1-Issue Break – News
-
Anime3 months agoHIDIVE to Stream English Dubs for The World Is Dancing, The Forsaken Saintess and Her Foodie Roadtrip in Another World, The Dangers in My Heart: The Movie Anime – News
