Tech
Anthropic can’t reliably control its AI agents. It’s cutting off its internal evals from the live internet instead
Anthropic said its models exploited websites on the internet, including some run by U.S. government agencies, and it will turn off live internet access for all of its internal evaluations until the frontier lab is sure it can monitor and control its AI agents.
The incidents, disclosed in a blog post, involved AI agents tasked to solve problems seeking resources on the internet. In the process, they exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information pass restrictions, and even submitted a false murder tip to the Philadelphia police.
Anthropic said it discovered these new issues in a review of its model’s activities that began in July, underscoring the lab’s lack of awareness of its software’s behavior.
Notably, the company said that alignment training was not yet sufficient for skills like search and computer use that are central to its pitch that AI agents will be used by any professional who relies on digital tools.
The behaviors Anthropic disclosed are similar to incidents involving OpenAI agents that collaborated to break into various websites in search of information, including some run by the Australian government.
Anthropic previously disclosed that its models had broken into external systems. The frontier lab said it considered today’s disclosures “significantly less severe from an alignment and security perspective” than those it announced before.
However, the lab still said it had “turned off live internet access” for “all our internal evaluations” until it is certain it can monitor and control its agents.
It’s not clear what that means, but Sydney Von Arx, the founder of Nightingale AI safety, told TechCrunch in an interview before this disclosure that developing models on a data center cut off from the open internet would be very challenging for researchers to access, and for the progress of the models, which benefit from internet access.
“You have to align them at some point,” Von Arx said. “If the AIs are released to production and never have access to the internet, that’s not a very useful tool.”
Anthropic said the behavior was a result of flaws in the lab’s training environments, which led the models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behavior called “reward hacking.”
The company said it would stop running some of its evaluations or move them offline, and has built tooling to detect and block this behavior. This tooling was tested against the kind of incidents disclosed today and blocked them; it’s not clear what evidence will prompt Anthropic to return live internet access to its internal evaluations.
Anthropic also said it would migrate its internal AI agents to “centrally managed infrastructure with strong containment,” and is beginning to using safety classifiers more frequently to monitor those agents.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
Tech
Long live the mechanical keyboard
For a writer, the mechanical keyboard is a beautiful thing. They’re certainly not necessary for everyone, but a well designed mechanical keyboard can turn a tedious day of typing into an aesthetically pleasing experience.
Given that most of my days involve sitting at a desk for hours, having a keyboard that I enjoy using is a necessity. I first got a mechanical keyboard about two years ago — the Keychron K2 — and I’ve never looked back.
Keychron is one of the most prominent brands within the mechanical keyboard industry, having made a name for itself after initially launching via Kickstarter back in 2017. The company has since produced dozens of different keyboards (as well as keypads), and a variety of mouse models.
Unlike other electronics, there isn’t a whole lot to owning a mechanical keyboard. The simplicity is part of the appeal. You unbox them, set them up (the K2 comes with little tilt legs that give you a slightly better angle from which to type), and plug them in.
From here, you have a couple of different options. The K2 comes with a simple USB cable, but there’s also an option for Bluetooth connection. If you value a clean, minimalist workspace, Bluetooth is probably the way to go.
While I am loyal to my older Keychron K2, there are a few upgraded and newer versions to consider. Keychron has released a variety of new models over the past two years, with a broad range of features and price tags.
In many cases, the form factor is the draw. One of the more interesting releases, called the Keychron K8 HE Wireless Magnetic Switch Custom Keyboard, comes with an all-wood body and is equipped with an LED backlight. That one is substantially more expensive at around $200. The K2, meanwhile, costs $60. But if you’re into the look, maybe it’s worth it.
If you’re a gamer, there’s also Keychron’s C0 HE One-Handed 8K Keyboard, which is a keypad built for convenience and speed and has an industrial aesthetic to it. However, interested gamers will have to wait — as it’s currently sold out.
The company has also updated its K, Q, and V series keyboards, with many of those ranging in price from $100 to $200 depending on the features and model.
One of the appealing elements of owning a mechanical keyboards is the versatility. You can swap out the key caps for a wide variety of others (indeed, there’s a whole sub-market devoted to this) or make your own custom caps (I’ve never gone that overboard).
There is also the sound, of course. Some people balk at the accentuated noise of loudly clacking keys, but I’m a fan. In fact, the sound is one of the key selling points for a lot of consumers. The distinct auditory experience produced by said keyboards has even made its way into ASMR videos.
For me, the Keychron’s biggest selling point is its retro form factor. I get really nostalgic for old school electronics, and Keychron scratches this itch pretty well. It’s a beautiful looking device, that brings to mind a different era of computing, when the style was more workman-like, less minimalist.
As far as home office purchases go, you could do a lot worse than to pick yourself up a piece of hardware that makes your desk look like it just time-traveled from the 1990s.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
Tech
The maker of non-text AI model Jev valued at $7.5B just weeks after launch
TypeSafe AI, the developer of Jev, a new type of AI model that gained rapid popularity after launching just a few weeks ago, has raised $870 million at a $7.5 billion valuation. The round was led by Andreessen Horowitz, with participation from Sequoia and existing investor DCVC.
The massive fundraise comes as little surprise given Jev went viral almost instantly after its Sept. 15 release. The startup claims that a third of Fortune 500 companies are already using the model, a remarkably swift adoption by enterprises.
Jev is based on a transformer architecture, but it is not a large language model (LLM). It doesn’t output text, but instead produces probabilities, or what the company calls “calibrated decisions.” What has users and large corporations so excited about Jev is TypeSafe’s claim that it works significantly faster and uses far fewer tokens than LLMs. The company positions its approach as uniquely suited for automating tasks rather than generating text or code.
“We have been super good at human language for four years, but it’s not useful for automation because computers speak a different language,” TypeSafe co-founder Diogo Almeida told TechCrunch last month.
Along with Almeida, who was previously a researcher at OpenAI, TypeSafe was co-founded in 2024 by former Meta research engineer Sasha Sheng and Erik Gafni, an engineer and entrepreneur.
>
Tech
AI Agents, Local Hardware, and Security Risks Reshape Tech
Explore the latest tech news, from ChatGPT’s interactive AI tools and Microsoft’s new hardware to cybersecurity threats, robotics, and Big Tech shakeups.
The post AI Agents, Local Hardware, and Security Risks Reshape Tech appeared first on TechRepublic.
>
-
movies5 months agoSearch For Canadian TV Actor Stewart McLean Now Homicide Investigation
-
Fashion9 years agoThese ’90s fashion trends are making a comeback in 2017
-
Fashion9 years agoAccording to Dior Couture, this taboo fashion accessory is back
-
Fashion9 years agoModel Jocelyn Chew’s Instagram is the best vacation you’ve ever had
-
Fashion9 years agoEmily Ratajkowski channels back-to-school style
-
Fashion9 years ago9 Celebrities who have spoken out about being photoshopped
-
Anime4 months agoRurouni Kenshin: Hokkaido Arc Manga Takes 1-Issue Break – News
-
Anime3 months agoHIDIVE to Stream English Dubs for The World Is Dancing, The Forsaken Saintess and Her Foodie Roadtrip in Another World, The Dangers in My Heart: The Movie Anime – News
