Connect with us

Tech

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Published

on

In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world. 

Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research unprecedented access to the company’s systems. CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.

Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out — and ideally backed by legislation — if they’re to know whether they will function as truly independent watchdogs or vendors operating on the AI companies’ terms. 

That deeper access is becoming more important as models get better at recognizing when they’re being evaluated, raising the risk that they’ll behave well during testing while concealing problematic behavior. Researchers say clues to that behavior can be missed when testing the finished model, but uncovered by investigating how it behaved throughout training. 

“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”

Historically, AI companies brought in outside reviewers to test finished models shortly before their release. Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training. Adam Gleave, CEO of Far.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed. 

Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Neither company has shared which evaluators they’ll work with, when they will be embedded, how many they’ll bring on, exactly what systems and information they will be able to access or what can be disclosed to the public, despite repeated questions from TechCrunch.

Looking under the hood like this matters because models that perform well on safety tests aren’t necessarily safe if they’ve learned specifically how to pass that test. Steidley pointed to an example of a “shutdown resistance benchmark” that measures if the AI will resist being shut down in certain circumstances.

“It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.

Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what happened internally. 

Amodei did outline a fairly comprehensive proposal that might give evaluators the kind of access they think is necessary, including the right to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.”

But evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Previous efforts at independent evaluations suggest that that surrender will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly.

Gleave said Far.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm’s independence. By default, he said evaluators are treated like ordinary contractors: bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published.

The time limit

There’s also the question of whether reviewers will get enough time and access to do the work they’re being asked to do. When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations. 

A similar issue occurred during the pre-release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet. According to Apollo Research’s contribution to the model card, the firm was given only three days to test Astra, which made it difficult to draw firm conclusions.

“Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the firm wrote in its evaluation. 

That track record leaves evaluators with a basic question: Why should this time be different? 

“It’s certainly possible that Dario and Sam just had a change of heart, and they’re going to be very open about this,” Gleave said. “But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.” 

Several researchers who spoke to TechCrunch called for a transparent framework that they all agree to publicly. Part of the framework, says John Steidley, head of strategy at Palisades Research, should involve standards for what kinds of auditors companies can rely on, lest they try to sidestep the issue by shopping for evaluators that either aren’t qualified or aren’t interested in assessing the most concerning risk. 

Henry Papadatos, executive director of Safer AI, says the problem, even with a public framework, is that voluntary measures are always dependent on a company’s goodwill. 

“Ideally, we would have good regulation mandating this…because then companies cannot change their mind tomorrow if they have a big PR crisis,” Papadatos told TechCrunch, noting that it’s also a good means of pushing all companies to adhere to the rules, not only the most willing. 

Not everyone has signed on. So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks. 

Some laws are already forming around the idea of third party evaluators. California’s SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized “independent verification organizations” with expertise assessing AI risks. 

In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts. 

For now the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they will submit to. Papadatos said voluntary self-regulation is better than nothing, but ultimately, companies can’t demand the freedom to control their own safety rules while also asking the public to trust that they’re following them. 

“You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules,” Papadatos said.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Tech

Musk’s long-time backer is giving SpaceX stock to its investors

Published

on

VC firm Valor Equity Partners, founded by Antonio Gracias, a long-time Elon Musk backer and current SpaceX board member, has opted to simply give away a chunk of its SpaceX stock to its limited partner investors, according to an SEC filing spotted by Bloomberg.

Valor made a killing on SpaceX after investing in it over decades, with entities controlled by Gracias owning more than 500 million shares at the time of the IPO. This was second only to Musk (who owned over 6 billion shares at the time). But instead of cashing out and issuing returns to LPs, Valor handed over 8.5% of its holdings to them, worth about $8.5 billion, Bloomberg estimates. Valor will still own more than 460 million shares after the giveaway, the SEC disclosure form says.

Transferring ownership of the shares could give those Valor limited partners a tax advantage. More importantly, it avoids dumping a giant tranche of shares into the open market. Such a dump could cause a glut of available shares and a corresponding dip in price. SpaceX is already down about 10% since its blockbuster IPO day.

>

Continue Reading

Tech

Al Gore has a surprisingly calm take on the AI data center backlash

Published

on

Al Gore has spent more than two decades as one of the most recognizable voices in the climate movement, so when he talks about AI, it’s worth trying to understand what he considers truly scary compared with what’s disturbing people so much that they are protesting in the streets.

In an interview with TechCrunch this week alongside Lila Preston, head of growth equity at their investment firm Generation Investment Management, Gore said emissions from AI data centers aren’t what should keep people up at night, despite the mounting opposition over their environmental impacts. What should concern people more, in his view, are the warnings coming from inside the AI industry itself, which he takes them more seriously than many other investors.

It’s not a new position for Gore, even if it’s something we haven’t discussed with him in our annual conversations, dating back several years. This past May, for example, he described AI data centers to another outlet as a “cause for deep concern, but not panic,” and in our conversation this week, his central point about emissions was about scale.

“If you look at the emissions of all of the AI data centers put together, it’s only a fraction of the emissions from uncovered landfills in the world,” he said, adding that a comparable build-out is tied to something that gets far less attention: air conditioning.

It’s a valid observation (and one the AI labs should be using in their discussions with cities and states if they aren’t already). The International Energy Agency estimates that global air conditioning already consumes more electricity annually than the entire European Union, and with ownership still under 15% across the hottest, fastest-growing parts of the world, that demand is expected to triple by 2050, putting far more pressure on the grid than AI data centers over the same period.

Gore isn’t dismissive of concerns about how these data centers are being powered, by the way.

“It is a matter of deep concern that some of the hyperscalers are jumping into new methane turbines,” he said.

But what worries him about methane turbines isn’t the AI workload driving demand so much as the prospect that building new gas plants to meet that demand locks in decades more of fossil fuel generation.

“I much prefer those that are supplying their energy needs with renewables and batteries,” he continued, saying he expects more companies to move inexorably in that direction simply because renewables are increasingly the cheapest option available.

In fact, Gore thinks the public backlash at planning-commission meetings has less to do with carbon emissions than with anxiety over automation upending society in the not-too-distant future.

“I think that part of the reason for the growing bipartisan opposition to data centers in so many parts of the U.S. is being driven by the underlying concerns about job losses and some of the other threats that experts inside OpenAI and inside Anthropic have been trying to alert the public to,” he said.

Naturally, while we had him on the phone, we pressed Gore on the wave of warnings from AI industry leaders that has dominated headlines over the last week. Though President Trump and Nvidia CEO Jensen Huang shared a laugh at their expense on Monday during a conference in Los Angeles (Trump suggested that Anthropic CEO Dario Amodei’s public call to slow the pace of capability gains, quickly seconded by Sam Altman and Elon Musk, was a “hoax”), Gore said he takes the warnings at face value.

“I think we went through a phase where some cynically suspected that the apocalyptic warnings were some kind of bizarre marketing strategy. If that was ever the case, I don’t think it is now,” he said. “I do think that Dario Amodei is very sincere, and I think that Sam Altman and Elon Musk were sincere in seconding [Amodei’s] warning of a few days ago.”

Gore being Gore (he often reaches for literature, scripture, and pop philosophy to make a point), he invoked Maya Angelou’s well-known line: “When someone tells you who they are, believe them the first time.”

Part of what’s changed his calculus, he said, is recent, concrete behavior from AI systems themselves — pointing to “the recent example of models escaping confinement, collaborating secretly, covering their tracks, engaging in deceptive behavior,” and noting that Anthropic has said it stopped instances of Claude being used to help develop biological weapons in different countries.

That said, Gore doesn’t treat AI risk and AI’s climate potential as being in opposition. He cited a recent study from economist Nicholas Stern of the London School of Economics projecting that AI applications aimed at efficiency and waste elimination could drive global emissions down by 6% to 9% per year starting next decade. Unsurprisingly, he suggested that the best thing for humanity would be if the U.S. and China could cooperate specifically on “solving of the climate crisis together and establishing appropriate regulations for the frontier models in AI,” even as he said he doubts the current U.S. administration will be the one to make that happen.

If Gore’s focus was risk, Preston’s was where capital is already moving in response. She pointed to several fronts where the AI build-out itself is generating investable opportunities, not just consuming energy.

Among these, the firm is focused on opportunities to decouple compute from energy intensity at every layer of the so-called stack, from the power supply itself down to green cement and green steel used in construction, storage optimization software, and database design. The firm is also focused on grid resilience and flexibility, as the mix of power sources grows more complex. Here, Preston noted, Generation has invested in Volue, a European company helping utilities pull more renewables into the grid, and Gridware, which places sensors on utility poles to monitor and help protect the grid from wildfire risk.

Gore tied the AI moment to a broader argument made in Generation’s 10th annual sustainability trends report: this year marks the second time in four years the world has been “brutally reminded” that fossil fuels are a volatile source of energy, first with Russia’s invasion of Ukraine and now with the disruption to the Strait of Hormuz.

But unlike past years, when Gore’s exasperation with the pace of the clean energy transition was obvious, he doesn’t see these events as a setback but instead as reasons for countries to embrace it at long last.

Indeed, Gore noted that globally, clean energy investment now runs at roughly twice the level of fossil fuel investment. Of all new electricity generation capacity added worldwide last year, 86% came from renewables, and in the U.S. specifically, that figure was even higher, at 91%, “in spite of Donald Trump’s best efforts to slow it down,” he said.

China, in particular, received high marks from Gore, who, among other things, applauded Beijing for wanting to be measured on emissions reductions rather than improvements in carbon intensity.

Asked what’s changed more than he expected a decade into publishing this report, Gore immediately pointed to solar power.

“For me personally, the scope and scale of the solar revolution is the breakout star of the sustainability transition,” he said. After observing that it’s at, or near, the cost floor for new electricity generation right now — roughly tied with wind and meaningfully cheaper than gas, coal, or nuclear — he closed with a line from a friend in Tennessee that had stuck with him.

“If God had intended us to have a limitless supply of cheap, clean energy, He — or she — would have put a fusion reactor in the sky.” Gore said with a chuckle, “The joke makes itself.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Tech

US automakers could soon be forced to include AM radio for free

Published

on

It turns out that the one thing that can apparently bridge Washington’s partisan divide is AM radio. U.S. automakers may be forced to build AM radio into new vehicles within the next year.

The House of Representatives, in a rare instance of bipartisan support, overwhelmingly approved legislation on Wednesday that would require new cars, trucks, and SUVs to be built with AM radio at no additional charge to consumers. The AM Radio for Every Vehicle Act, which would direct the National Highway Traffic Safety Administration to require automakers to include AM radio in new passenger vehicles as standard equipment, will now head to the Senate.

Until recently, proponents of AM radio have waged an unsuccessful lobbying effort for a mandate, arguing that it offers critical information during emergencies because its signal can travel long distances. But a groundswell of support from Democrats and Republicans, the broadcast industry, emergency responder organizations and even former administrators of the Federal Emergency Management Agency is pushing this legislation closer to becoming law.

And there is reason to believe the Senate will follow the House’s lead on the bill, which was co-sponsored by Sens. Ed Markey (D-Mass.) and Ted Cruz (R-Texas). Earlier this year, the senators secured 60 co-sponsors for their version of the bill, a critical threshold to overcoming a filibuster.

Cruz, who chairs the Senate Commerce, Science, and Transportation Committee, and Markey hailed the decision in a joint statement Wednesday.

“This vote sends a clear message to car manufacturers that AM Radio is a lifeline that must be protected in new vehicles. From emergency response to sports, entertainment, and news, AM radio is an essential communication tool for tens of millions of Americans. It is now time for the Senate to pass the AM Radio for Every Vehicle Act and for this legislation to become law so AM radio remains a trusted and essential resource for commuters and communities across the country.”

A growing number of automakers have moved away from the century-old broadcast band in favor of software-defined vehicles that offer streaming services like TuneIn and SiriusXM. EV makers have led the charge because the electromagnetic interference generated by electric motors can affect the frequencies used by AM radio. Automakers have argued that this interference degrades the AM signal and audio quality.

A growing number of automakers no longer have traditional AM receivers, including BMW, Rivian, Tesla, and Volvo. Ford had also removed AM radio from some vehicles but has since reversed course. Tesla couldn’t be reached for comment; Rivian declined to comment for this story.

Tesla doesn’t have AM receivers in the base versions of its Model 3 and Model Y vehicles. Rivian’s new R2 SUV doesn’t have an FM or AM receiver, however, the vehicle model offers AM/FM radio through the digital radio platform iHeartRadio for free. The Rivian R1S and R1T models allow customers to access FM content digitally or via built‐in receiver.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Trending

Copyright © 2017 Zox News Theme. Theme by MVP Themes, powered by WordPress.