Tech
Kog is going deeper to squeeze more inference out of GPUs
The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.
The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and NVIDIA H200 GPUs it used for its demo.
Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch.
Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.
Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.
The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”
This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open-sourced Laneformer 2B.
Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.
Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper-level focus on GPU acceleration.
Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, has nothing to do with his new one — other than his former cofounder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.
Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.”
As for hacking, the four-time finalist at at DEF CON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.”
The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future.
In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program.
For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
Tech
Gemini Becomes Google’s 14th Product to Reach 1B Monthly Users
Gemini has joined the billion-user club, giving Google its strongest sign yet that its AI assistant has moved into the mainstream.
Google said Tuesday that the Gemini app has surpassed 1 billion monthly active users, making it the company’s 14th product to reach that milestone. The figure puts Gemini among the largest consumer AI services in the world.
ChatGPT has also crossed the billion-user mark, although OpenAI reports weekly rather than monthly users, making the figures unsuitable for a direct comparison. Together, the milestones show how quickly generative AI assistants are moving from novelty to mass-market infrastructure.
Gemini took acceleration seriously
Gemini’s 1 billion monthly users are impressive, but the speed at which Google reached that number is arguably more significant.
Google’s CEO, Sundar Pichai, announced the development on his X page.
According to The Verge, the platform went from 750 million MAUs in February to 950 million last month, then added another 50 million, bringing the total to 1 billion.
That growth is also showing up in how people use the service. In its July earnings call, Google said Gemini’s daily active users have tripled over the past year, while newer features such as Daily Brief and Gemini Spark are pushing the product beyond simple question-and-answer interactions and toward more persistent assistance.
The result is a product that is not only reaching more people but becoming a more regular part of how some of those users interact with AI. That distinction matters because a large monthly user base is less meaningful if people rarely return or use the product only once.
A billion users, but with a revenue question
A billion monthly users gives Google enormous reach, but reach is only part of the business equation. Google has not disclosed how much revenue Gemini itself generates, making it difficult to tell how effectively that audience is being converted into paid subscriptions or other revenue.
OpenAI offers a useful comparison because it has already demonstrated that a large AI user base can translate into substantial revenue. OpenAI’s annualized revenue reached $25 billion by the end of February, according to Reuters. However, that figure reflects OpenAI’s broader AI business, not ChatGPT alone.
That leaves Google with a different question after reaching 1 billion users: how much is each of those users worth? The answer will matter more over time than the headline user count itself, particularly as Google continues investing heavily in the infrastructure required to run Gemini.
The underlying moat behind Gemini’s 1 billion MAUs
The biggest advantage behind Gemini’s growth may not be Gemini itself, but where Google can put it.
The company controls Android, Chrome, Search, YouTube, Workspace, and several other products, giving Gemini access to products and devices that already have enormous audiences rather than requiring every potential user to seek out a separate AI service.
That changes the mechanics of user acquisition. Someone does not necessarily have to decide to download Gemini or visit its website, as Google can introduce the assistant inside products they already use.
ChatGPT, on the other hand, has built an enormous audience without owning an operating system or search engine. That means while OpenAI is building integrations and partnerships, Google is making Gemini part of the software people already use every day.
For enterprises looking to expand, Google’s advantage shows why owning the channels where customers already spend their time can matter as much as the product itself.
Related News: Google is pushing Gemini deeper into the Pixel 11, adding new AI features designed to make the smartphone a more proactive and personalized assistant.
>
Tech
Apple in Talks to Pay Publishers for Siri AI Content
Apple is discussing multiyear deals with publishers that would let Siri AI use their content for current news and information, according to The Wall Street Journal. Apple has proposed variable compensation, with publishers paid when their content is used, and has discussed a possible nine-figure budget.
The publishers involved and exact payment terms remain undisclosed.
The talks arrive as Apple prepares a broader Siri AI rollout later this year. Apple has confirmed that the assistant can retrieve up-to-date information from the web, but it has not publicly explained how publisher sources would be cited or linked in those answers. Apple declined to comment on the reported negotiations.
Siri AI is being built around current web information
Apple says Siri AI can combine web retrieval with onscreen awareness, personal context, and broad world knowledge to answer questions on virtually any topic. The assistant is available for developer testing across iOS 27, iPadOS 27, macOS 27, and visionOS 27, with a user beta planned for later this year.
Siri AI will work across iPhone, iPad, Mac, Apple Watch, and Apple Vision Pro, although supported devices and regional eligibility vary.
The Wall Street Journal reported that Apple’s proposal differs from common AI licensing agreements, which offer publishers guaranteed fees for broad access to content. Apple’s model would instead compensate participating publishers when their content is used. The company previously paid publishers in agreements that included AI training rights and has commercial relationships through Apple News+.
The Journal also noted a source-quality concern from Apple’s past. In late 2024, Apple tested AI-generated news summaries that produced erroneous headlines, prompting the company to disable the feature.
Publishers are negotiating as AI changes referral traffic
Apple’s discussions enter a market where publishers are reconsidering how AI companies use their work.
Reddit has reportedly reviewed parts of its Google AI partnership as AI-generated search answers change the economics of web traffic, while Google AI Overviews have faced legal and regulatory scrutiny over publisher content, attribution, and traffic.
Brookings research published in June describes an AI licensing market split between direct deals with major publishers and an intermediary layer offering content marketplaces, bot controls, and pay-per-use models. It found that publishers with direct licensing agreements initially received a click-through advantage from AI interfaces, but that premium had largely disappeared by the fourth quarter of 2025 amid a sixfold decline in AI click-through rates.
For publishers, unresolved terms include attribution, linking, usage reporting, and how variable payments would be calculated. IT and AI-governance teams have a separate concern: whether policies for AI-assisted research account for assistants that retrieve and synthesize live web content across managed devices.
Apple’s planned user beta may provide more visibility into Siri AI’s citation and source-link behavior. It is less likely to reveal the private commercial terms behind any publisher agreements.
Also read: Apple’s iOS 27 beta 2 adds Write with Siri, RCS upgrades, and other changes ahead of the full software rollout.
>
Tech
Apple proposes to take a 15% cut of purchases made outside the App Store
After trying and failing to delay the matter, Apple on Thursday submitted its proposal for the commissions it wants to charge on purchases made using external links inside apps on its iOS devices.
In a new filing in the U.S. District Court of Northern California, Apple proposed new commissions of 15% for standard apps, with further discounts for developers who are enrolled in special Apple programs.
Under the proposed structure, small business developers would pay a 5% commission on payments, while those in the Video Partner Program, News Partner Program, and Mini Apps Partner Program would pay 10%. Subscription renewals would also be reduced to 10%, the filing states.
The iPhone maker has been engaged in a years-long legal battle with Epic Games over its alleged anti-competitive policies regarding App Store commissions, and had been attempting to stall its answer to the court’s request for this part of its commission structure.
Apple tried to argue that these lower court proceedings should wait until the Supreme Court ruled on another matter related to the case: Whether or not Apple was in contempt of a court order when it imposed a new 27% commission on purchases made through external links, and imposed rules that restricted how developers could present those links to customers.
The Supreme Court on Thursday rejected Apple’s bid to pause further action in the lower court’s case, forcing the company to reveal its planned commission structure.
Apple’s position is that it should be permitted to charge fees on in-app purchases made by users of its devices as a means to recoup its investments in the tools, technology, and services that allow it to maintain its App Store and software.
The company also compared its link-out fees to those on Google Play, which charges 20% link-out rates for standard apps, 15% for apps in special programs, and 10% for subscription renewals, noting that Epic Games had agreed to these rates.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
>
-
movies3 months agoSearch For Canadian TV Actor Stewart McLean Now Homicide Investigation
-
Fashion9 years agoThese ’90s fashion trends are making a comeback in 2017
-
Fashion9 years agoAccording to Dior Couture, this taboo fashion accessory is back
-
Fashion9 years agoModel Jocelyn Chew’s Instagram is the best vacation you’ve ever had
-
Fashion9 years agoYour comprehensive guide to this fall’s biggest trends
-
Fashion9 years ago9 Celebrities who have spoken out about being photoshopped
-
Fashion9 years agoEmily Ratajkowski channels back-to-school style
-
Fashion9 years agoA photo diary of the nightlife scene from LA To Ibiza
