Connect with us

Tech

Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

Published

on

The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text.

Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what’s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.

Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.

The launch comes after a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.

Goodfire’s system works a bit like airport security. Small detectors called probes read the model’s internal signals at every step of an agent’s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.

Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely. 

Goodfire says its approach is also cheaper to run. Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire’s probes instead tap into calculations the model is already making as it works.

“Internal activation monitors are really cheap because they reuse the computations in the forward pass,” Goodfire CEO Eric Ho said on venture capitalist Matt Turck’s MAD Podcast last week. “So the model’s already computing this token. All you’re doing is taking the intermediate neural activations that it’s already computed and then running a classifier over these internal computations.” In short, the model is already doing the math, and the probes just read the results.

In Goodfire’s tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look.

Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.

“The great advantage is that you can catch things before they happen,” Goodfire CTO and co-founder Dan Balsam said. “We can detect when the model might hack during eval or training.”

The pitch is aimed at open models. Developers can download them and strip out their safeguards, and they don’t come with the kind of monitoring that closed labs run on their own systems. 

“The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers—where most of the liability is,” said Balsam. “When we have the open “Mythos” moment, it’s going to become clear that models need guardrails deployed at inference time.”

Goodfire’s recent research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents. 

Goodfire isn’t the first to try this approach. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini. 

Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. “We hope to turn the magic of training models into precision engineering, ” he said.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Tech

Engineering Ethics Should Be Part of Technical Training

Published

on

In mid-September, Anthropic CEO Dario Amodei joined a growing group of researchers calling for the AI industry to slow down. While some have questioned these researchers’ motivation for voicing concerns about the rapid pace of development, the plea highlighted a key tension between safety and speed. The underlying problem is not simply a technical error, but a sociotechnical failure: Pressure to move quickly has made it hard for engineers to fully assess and manage new risks of increasingly powerful AI systems. It is a clear warning that responsible engineering requires more than technical expertise; engineers must also understand workplace pressures, moral and ethical rules, and the social consequences rooted in the systems they create.

Cases like this, alongside invasive surveillance technologies and biased AI tools, have exposed serious shortcomings in how engineers recognize and address the ethical and social implications of their work. While the need to educate graduates in ethics and social responsibility is widely recognized, a persistent gap remains—these topics are too often treated as afterthoughts, disconnected from the technical content of engineering coursework. This siloed approach leaves graduates ill-equipped to address real-world challenges; indeed, only about a third of engineering professionals in the United States report having received training on their public welfare responsibilities during their formal engineering education.

We are working to bridge this gap by designing two complementary instructional approaches, one that integrates social issues directly into foundation technical coursework and a second that prepares students to act as “public welfare watchdogs” through a standalone engineering course. By embedding ethics and social responsibility training directly into the core of undergraduate engineering coursework, we equip future professionals with the agency they will need to protect the public good.

The gap in instruction about ethics and social responsibility

Across the engineering profession, there is a persistent belief that technical skill-building is central to engineering, while ethical, social, and public welfare concerns are peripheral or distracting. This technical/social dualism is evident even in the evaluation of scholarly research. For instance, we recently submitted an academic position paper discussing this ethics education gap to a leading peer-reviewed engineering journal. It was dismissed by reviewers along the very same dualism it described.

One reviewer warned that it would be “highly irresponsible” to challenge technical skill-building as “the foundation of any kind of engineering” by adding ethics and social responsibility training into engineering classrooms. Another reviewer objected that integrating these considerations into foundational courses risks “distracting the students from fully absorbing the technology.” These comments miss the core argument of our work: that ethics and social responsibility are integral to responsible technical practice.

Today, accreditation agencies like ABET (formerly known as the Accreditation Board for Engineering and Technology) and professional licensing bodies like the National Society of Professional Engineers mandate student outcomes related to social and public welfare consequences. But these curriculum requirements tend to be unhelpfully generic and minimally enforced. As a result, engineering students often receive ethics instruction through introductory or non-engineering coursework, where the focus is on codes of conduct and academic integrity, rather than deeper topics of social responsibility. This instruction is almost always disconnected from the complex real-world dilemmas engineers face.

Indeed, we find that only ethics instruction in actual engineering courses (not in other parts of the curriculum) can effectively instill ethical awareness as professionals. Concerningly, only about 4 in 10 engineers in the United States recall receiving this kind of training, and 3 in 10 report never receiving training on their public welfare responsibilities at any point in their careers. This leaves most of the profession underprepared to handle the complex challenges of engineering practice.

A new approach: Integrating technical and social dimensions

To advance a new paradigm—one that embeds ethics and social responsibility issues within the undergraduate engineering curriculum—we’ve designed two innovative course-based approaches: a library of one-hour sociotechnical modules for foundational circuits courses (described here and here) and a new engineering course dedicated to public welfare responsibilities.

The first approach integrates discussions about relevant social issues directly into foundational engineering courses, where students begin to form their identities as engineers. “Introduction to Circuits” is typically the first course in the electrical engineering major and is a requirement for many other engineering disciplines, making it a powerful site for introducing social content and ethical reasoning. We designed several one-hour modules for this course, where social issues provide context for a technical topic, offering a manageable entry point for faculty who may feel ill-prepared to teach about ethics and social responsibility.

While many engineering faculty recognize the importance of teaching students about these topics, they struggle to integrate them into their courses amid heavy teaching and research demands. To overcome these challenges, each module provides ready-to-use materials, including lecture slides with scripts and homework and exam problems. Integrating these modules directly into the course signals to students that sociotechnical awareness is not an “add-on,” but a required professional competency.

Our module on hospital power prioritization, for instance, connects the technical topics of power and equivalent circuits to the ethical challenges of healthcare infrastructure. Students learn about the reality of “red outlets” in medical settings, debating the power needs of specific wards (e.g., the intensive care unit, operating rooms, and physiotherapy rooms) to decide which branches should receive emergency power during a hypothetical outage.

Our conflict minerals module connects the technical topic of capacitors to the global tantalum supply chain. Students calculate the amount of tantalum used in cell phones and identify mining locations in regions like the Democratic Republic of the Congo, where revenue often fuels armed conflict and human rights abuses. Students research the conflict-mineral policies of major companies such as Apple and Samsung and reflect on their future roles as ethical designers and consumers.

Student responses to these modules have been overwhelmingly positive. In interviews, students told us that the content helped them get away from the “stigma that engineers only worry about math” and made the work more meaningful. We plan to share the modules widely to help circuits instructors across the globe bridge the technical-social divide.

Beyond these integrated modules, our second approach is a standalone engineering course—Public Welfare Responsibilities in Engineering (PWRE)—that helps round out students’ sociotechnical education by preparing them to fulfil their professional obligations as engineers. While most dedicated ethics courses end by introducing engineers’ abstract responsibilities, PWRE helps students recognize specific public welfare concerns, enhances their motivation to act when they encounter them, and teaches concrete, practical intervention strategies.

In addition to teaching students about professional codes, such as the IEEE Code of Ethics, PWRE engages students in discussion of where these formal definitions may fall short, and how structural barriers such as workplace culture can hinder engineers from fulfilling these responsibilities. Then, the course equips students with intervention strategies. Students learn about their rights and responsibilities as whistleblowers, review tactics for seeking advice from professional societies, and practice writing op-eds to raise the alarm about ethical issues to the broader public. Our research shows that the course produces notable shifts in students’ ethical agency, strengthening their ability to serve as “public welfare watchdogs” in the modern workplace.

Ethical engineers for the future

Today’s engineering graduates face complex sociotechnical challenges beyond AI expansion, including climate change and global supply chains; they need tools to navigate ambiguity and act ethically in uncertain contexts. To meet these evolving demands, engineering education must move beyond outdated modes of instruction that focus only on building technical skills—some of which may soon be outsourced to AI. We must fundamentally reimagine the role of ethics and social responsibility, embedding them across the engineering curriculum and throughout the professional formation of engineers.

All engineers learn to solve problems, but ethical engineers can also discern which problems are worth solving and anticipate the consequences of innovation for society. By centering ethics and social responsibility as core priorities of the profession, we can prepare engineers to be stewards of technology and trusted voices in public discourse.

As one student reflected after participating in our modules, “we are a part of the issue if we don’t decide to fix it.”

From Your Site Articles

Related Articles Around the Web

>

Continue Reading

Tech

Master AI Chip Principles With New IEEE Design Program

Published

on

Today’s engineers face an unprecedented acceleration in AI hardware complexity, as explained in the recent research article “Revisiting Edge AI: Opportunities and Challenges.” The article examines the rapid growth of edge AI and the challenges it creates, including resource constraints, model architecture limitations, and network demands across edge-AI deployments.

The acceleration is driven by a fundamental shift in how modern AI models are built and scaled. As the models have become much larger and more complex, they are computationally more demanding because they contain more parameters and require more calculations.

To meet the demands of scaling deep neural networks, the industry is increasingly developing AI chips that are designed for specific tasks.

One major reason is that moving data between memory and the processor has become a major limitation on AI performance.

The movement to confront the hardware bottleneck—the AI memory wall—has altered the trajectory of semiconductor innovation, shifting architectural priorities toward domain-specific accelerator platforms.

No longer can engineers evaluate systems statically; they must master joint hardware design and network-algorithm co-optimization to navigate the critical trade‑offs between throughput, latency, and operational efficiency.

The challenges are addressed in the new AI Processor Architecture, Design Principles, and Performance program, developed by IEEE Educational Activities with support from the IEEE Computer Society.

Topics covered

The five-course program provides a structured exploration of AI processor technologies, from fundamental design principles to advanced architectures and real-world deployment.

The topics are:

  • Fundamental principles of design and functionality.
  • Practical insights into advanced architectures.
  • Understanding neural processing units for industry deployment.
  • Emerging trends and evolving architectures.
  • Designing for edge, cloud, quantum, and the Internet of Things (IoT).

The program addresses the needs of professionals across the AI hardware ecosystem, including hardware architects, chip designers, embedded systems developers, data‑center hardware engineers, and innovators exploring next‑generation processor ecosystems.

It is also valuable for those transitioning into AI chip design or seeking to understand the architectural forces shaping modern machine learning acceleration. For many, it provides the bridge between theoretical knowledge and the collaborative, cross‑disciplinary reasoning required in engineering environments.

The program is designed to explore how modern AI processors are conceived, structured, and optimized. The architectural layers that define contemporary AI hardware will be covered, including compute units, memory hierarchies, dataflows, and the performance characteristics that emerge from design decisions. The curriculum bridges theory and application, enabling participants to interpret architectural foundations, analyze trade‑offs, and understand how hardware structures shape computational efficiency across diverse environments.

By the end of the courses, learners will be able to evaluate processor behavior with the analytical precision expected of professionals working at the frontier of AI hardware design.

AI-generated avatars explain concepts

The program also uses a dialogue‑driven learning approach. Learners view conversations between AI-generated avatars that engage in scenario‑based dialogues. The avatars represent engineers tackling the same problems from various perspectives based on their different roles. The simulated storytelling and dialogues make advanced engineering concepts approachable without sacrificing depth or interactivity, because the user is occasionally challenged to decide the correct answer that leads to the best course of action.

A hardware engineer might push back against a systems engineer’s demands, for example, revealing the friction between physical constraints and algorithmic ambition. A computational validation specialist might interrogate a chip performance engineer’s optimism, exposing the gap between theoretical throughput and real‑world behavior. A heterogeneous systems architect could debate a multiprocessor coordination specialist about synchronization overheads, while a standards development engineer discusses regulatory implications with a technology strategy and compliance architect.

In the final course, an IoT systems architect and an embedded AI optimization engineer dissect the realities of deploying AI in constrained environments.

By the end of the course, learners will be able to evaluate processor behavior with the analytical precision expected of professionals working at the frontier of AI hardware design.

The avatar-driven learning environment creates a psychologically safer environment for learners, who might feel intimidated by traditional expert‑led videos. Research published in 2024 in IEEE Transactions on Learning Technologies showed that avatar‑based instruction can increase a learner’s confidence by up to 25 percent and improve retention of complex technical material.

Each course concludes with a module in which the two experts from different disciplines debate, question, and challenge each other’s assumptions.

Modeling expert reasoning through dialogue

The cross‑disciplinary conversations can do more than explain concepts; they can model how experts think. They can reveal the negotiations behind architectural choices, the competing priorities that shape system design, and the analytical rigor required to balance performance, efficiency, scalability, and compliance.

Learners observe engineering discourse, gaining insight into the reasoning patterns that drive innovation in AI processor development.

The program goes further by interrupting the dialogue at key moments and inviting the learner to step in. Instead of passively absorbing information, the learner decides how to resolve a trade‑off, predict the outcome of a design choice, or select the most defensible engineering path. The experience can feel less like a course and more like an apprenticeship inside an engineering team.

By merging rigorous technical content with an innovative, dialogue‑driven delivery model, the program sets a standard for teaching advanced engineering. It captures the complexity of real‑world problem‑solving, humanizes the learning experience, and challenges participants to think like the engineers who are defining the future of AI hardware. It is not simply a new course; it is a new way of learning.

For individual access, visit the IEEE Learning Network.

For customized organizational options, contact a content specialist to discuss volume pricing.

From Your Site Articles

Related Articles Around the Web

>

Continue Reading

Tech

An Anthropic AI model sent a false homicide tip to Philadelphia police

Published

on

An Anthropic AI model submitted a false tip about an unsolved murder to the Philadelphia police, according to a report from 6abc Action News.

The AI reportedly submitted this incorrect information to a public Philadelphia Police Department (PPD) tip line on July 18, but Anthropic didn’t discover the behavior until September 28. The police had not seen the tip because it was marked as spam.

Anthropic notified the PPD about the incident on Wednesday and met with the department the following day.

Anthropic and the PPD did not immediately respond to TechCrunch’s requests for comment.

“The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city’s knowledge. The two-month delay in detecting and reporting the incident to the City is unacceptable,” the PPD said in a statement to 6abc.

As autonomous AI agents are increasingly made available to consumers, this incident highlights the danger of giving AI the ability to carry out tasks without any human supervision.

Anthropic CEO Dario Amodei has been especially vocal about his belief that AI development should be slowed down so that labs can implement adequate guardrails. Perhaps this stance was informed, in part, by witnessing his company’s tools submit false homicide tips.

These issues are not exclusive to Anthropic. OpenAI recently revealed that one of its models acted unexpectedly during a test and hacked the AI dataset platform Hugging Face, exposing critical vulnerabilities in its software. As AI models continue to be granted unchecked access to people’s computers and login credentials, this problem is expected to persist.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Trending

Copyright © 2017 Zox News Theme. Theme by MVP Themes, powered by WordPress.