Connect with us

Tech

Kog is going deeper to squeeze more inference out of GPUs

Published

on

The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.

The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and NVIDIA H200 GPUs it used for its demo.

Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch. 

Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.

Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.

The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”

This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open-sourced Laneformer 2B

Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.

Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper-level focus on GPU acceleration.

Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, has nothing to do with his new one — other than his former cofounder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.

Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.”

As for hacking, the four-time finalist at at DEF CON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.”

The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future.

In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program.

For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said. 

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

>

Continue Reading

Tech

White House AI Safety Reviews Could Expand to Open Models

Published

on

The White House is preparing to bring powerful open AI models into its secretive safety-review framework.

The administration’s current framework applies to closed models from leading AI companies, including OpenAI and Anthropic. But White House officials are now expected to expand it to open models once they reach frontier-level capabilities, according to WIRED.

A White House official told WIRED that open models could face prerelease testing when their capabilities reach the level of Anthropic’s Mythos-class models and OpenAI’s GPT-5.6.

The framework remains voluntary and has not been publicly released. Axios reported that the administration has generally viewed models with frontier capabilities and national security risks as requiring some form of government collaboration, regardless of whether they are open or closed.

That puts the administration in a difficult position. Open models can be downloaded and modified after their weights are released, making them fundamentally different from proprietary systems that their developers can update, restrict or shut down.

A balancing act for Washington

The potential policy shift comes after growing pressure to keep open models competitive in the US AI market. Companies including Meta, Microsoft and Palantir have backed efforts to protect open-weight development. Nvidia CEO Jensen Huang has also argued for maintaining both open and closed frontier models.

“Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty,” Huang said. “The world needs both frontier closed models and frontier open models.”

The administration also faces a practical concern: If government approval becomes associated primarily with closed models, businesses could view open models as riskier, potentially hurting US developers working on them. At the same time, officials worry that highly capable open models could be misused for cyberattacks and other national security threats.

More must-read AI coverage

A new test for open-weight AI

The emerging debate highlights a growing reality for AI companies: the distinction between open and closed models may matter less to regulators than the capabilities of the systems themselves.

For businesses building AI products, broader oversight could provide greater confidence in the safety of advanced models. For developers of open-weight systems, however, additional reviews could slow releases and increase compliance costs.

The challenge for policymakers is finding a middle ground. Open models have become a critical part of the US AI ecosystem and are increasingly viewed as a strategic response to China’s advances in the field. Yet the same accessibility that fuels innovation also raises concerns about misuse.

As frontier AI capabilities continue to spread beyond a handful of closed providers, the administration appears to be moving toward a capability-based approach rather than one defined by whether a model is open or closed.

Why this matters

For businesses adopting advanced AI, the White House’s approach could become another signal for evaluating model risk. If open and closed models undergo similar safety reviews once they reach frontier capabilities, enterprises may have more information to weigh alongside performance, cost, and deployment flexibility.

For open-model developers, however, broader oversight could introduce new friction. Prerelease testing may slow launches, increase compliance costs and make it harder for smaller developers to compete with companies that have larger legal and policy teams.

The bigger shift is regulatory. Rather than treating open and closed AI as separate categories, Washington appears increasingly focused on what a model can do and the risks those capabilities create. If that approach takes hold, capability thresholds could become a much more important factor in how advanced AI is developed, released and adopted.

Also read: For another example of how openness can create security trade-offs, 77 counterfeit Open VSX extensions recently exposed developer and CI/CD environments to supply-chain risk.

>

Continue Reading

Tech

WhatsApp Begins Limited Test of AI Scam Alerts for Unknown Senders

Published

on

WhatsApp is adding a new layer of protection that could stop some scams before users even respond.

WhatsApp is rolling out Scam Alert, an optional feature that uses an on-device machine learning model to identify potential scams in messages from people who are not in a user’s contacts. The feature is currently available in a limited beta as Meta tests it with security researchers and gathers feedback before a broader launch.

The model looks for linguistic and conversational patterns associated with known scams. If it detects a likely scam, WhatsApp displays a warning that only the recipient can see. Users can then block or report the sender, continue the conversation, or mark the chat as trusted.

If a warning is wrong, marking the conversation as trusted removes the alert and prevents Scam Alert from flagging that chat again. Users can also voluntarily share the last five messages from a trusted conversation with WhatsApp to help improve the model.

Privacy remains central

Meta says Scam Alert was designed to work without giving WhatsApp access to users’ encrypted conversations. The detection model runs locally on the device, and message content is not automatically sent back to the company.

The company says it also cannot use the system to send a particular AI model to an individual user selectively. Each model version is recorded on a third-party transparency ledger before distribution, while devices verify published signatures and file hashes before loading a model.

WhatsApp also collects limited performance data, such as how often warnings appear and what users do afterward. According to Meta, those figures are aggregated and protected using confidential computing and differential privacy rather than exposing individual conversations.

Users will eventually be able to inspect Scam Alert activity in WhatsApp under Account > Request Info > Scam Alert Activity, including which messages were analyzed, the result, and the model version involved.

Must-read security coverage

A response to a growing scam problem

The feature arrives as scammers increasingly use messaging platforms to build trust, pressure victims and solicit money or personal information. The Federal Trade Commission said consumers reported $2.1 billion in losses from social media scams in 2025, including $425 million tied specifically to WhatsApp, according to CNET.

That makes the phone-based approach particularly significant: WhatsApp is trying to add an automated layer of protection without requiring its servers to inspect private conversations.

The tradeoff behind the warning

Scam Alert could give users a useful second opinion before they respond to an unfamiliar sender, but it is not a guarantee that a message is safe or fraudulent. Machine learning can miss sophisticated scams or incorrectly flag legitimate conversations.

Scam Alert remains an experiment, not a finished security feature. Meta says it will continue testing the system with security researchers, expand bug-bounty coverage, and refine the model before deciding whether to make it broadly available.

If the approach works, WhatsApp could gain an additional layer of scam protection without abandoning the privacy model that has defined the service — but users will still need to treat any automated warning as guidance rather than proof.

Other News: DentaQuest has disclosed a data breach affecting nearly 15 million people, with exposed information potentially including names, Social Security numbers, and health insurance details.

>

Continue Reading

Tech

Microsoft Nearly Left China. AI Gave It a Reason to Stay

Published

on

Microsoft has spent years retreating from China. AI may be a reason it does not leave entirely.

Reuters reports that the company has closed at least 15 offices and joint ventures in China and came close to executing a complete exit in 2023, as geopolitical tensions, tighter Chinese technology policies, and U.S. export controls have made the market increasingly difficult to serve.

However, Microsoft has found a narrower opportunity in the same market: selling Azure cloud and AI services to Chinese companies with international operations, including ByteDance.

That shift changes what Microsoft’s China strategy looks like. Rather than trying to remain a major foreign software provider in China’s domestic market, Microsoft is increasingly using its cloud platform and access to Western AI models to serve Chinese companies that need technology for their businesses outside the country.

That business model, too, isn’t entirely immune to the problems that led the company to want to leave, leaving Microsoft in an unusually sensitive spot.

Microsoft is facing pressure from two sides

Microsoft’s office closures are not just about reducing physical presence; they reflect a business that has become harder to justify as Microsoft’s traditional software market in China shrinks.

The country accounted for only about 1.5% of Microsoft’s global revenue in 2024, according to Reuters. That means Microsoft is being asked to bear significant geopolitical, regulatory, and operational risks for a relatively small share of its worldwide business.

Beijing already has a growing preference for domestic technology. Chinese government agencies have been encouraged to replace foreign software with Chinese alternatives. Microsoft’s own attempt to build a China-specific version of Windows for government customers also did not translate into measurable adoption.

What is equally important is the squeeze from Washington. U.S. export controls on advanced chips and AI technology have restricted what Microsoft can provide to Chinese customers and what its China-based engineers can access.

Reuters says those restrictions have also affected Microsoft’s research operations, with the company having to move some researchers outside China.

The problem is that U.S. restrictions can also strengthen the very domestic technology push that is hurting Microsoft. Beijing has also tightened restrictions to foreign businesses.

In other words, U.S. controls may directly limit Microsoft’s business in China while also giving Chinese companies and policymakers another reason to reduce their dependence on foreign technologies like Microsoft’s.

More Microsoft news

AI gives Microsoft a narrower path in China

AI and cloud services give Microsoft a reason to preserve part of its China business, but analysts have questioned how durable that opening is.

Microsoft relies on third-party AI providers, meaning a change in U.S. or AI providers’ access-to-China policy could quickly weaken one of the reasons customers use Azure.

There is another problem: the same push for domestic technology that hurt Microsoft’s traditional software business also applies to AI. Chinese companies have increasingly capable homegrown models, giving them another reason to avoid foreign providers.

That is where companies such as ByteDance and Shein become particularly important. They operate internationally and often require AI and cloud infrastructure to serve global markets. That makes Microsoft’s Azure and access to Western AI models an attractive alternative.

The old global tech playbook is breaking down

Microsoft’s China strategy shows how geopolitical tensions are moving beyond policy documents into corporate operating decisions. The company has reduced its footprint while keeping selected businesses alive, even as a Reuters source stressed there are no definitive plans to leave China.

For other multinationals, that could become the new playbook: Stay where the business still works, reduce exposure where politics raises the cost, and build around a technology environment that is increasingly split along national lines. How that plays out in the long run remains to be seen.

Other Microsoft News: The company is merging its separate Microsoft 365 Copilot and Copilot apps into a more unified experience as it continues consolidating its growing lineup of AI tools.

>

Continue Reading

Trending

Copyright © 2017 Zox News Theme. Theme by MVP Themes, powered by WordPress.