Últimas Noticias de IA

Kog is going deeper to squeeze more inference out of GPUs
The race for fasterAI inferenceis on, and markets gave Cerebras and its purpose-built chipsa warm welcomein its IPO debut in May. But French startupKogis betting that there’s a lot more power to be squeezed out of conventional GPUs. The startuphit the front page of Hacker Newsin May with atech previewaimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and Nvidia H200 GPUs it used for its demo. Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attractedmore than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch. Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and chargesa price multiplefor Claude’s Fast Mode. Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said. The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.” This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, thenow open sourcedLaneformer 2B. Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked. Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box.ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University labHazy Research, with an even deeper-level focus on GPU acceleration. Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alumStribe, has nothing to do with his new one — other than his former co-founder turned VC Kamel Zeroual, whose firmVarsity VCco-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background. Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.” As for hacking, the four-time finalist atDEFCON’s CTF tournamentsaid it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.” The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future. In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is alreadysupported by Scalewayand backed by France’sBpifranceandFrench Tech 2030’s program. For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said.
View

Does Mark Zuckerberg really believe AI is ‘for everyone’?
Loading the player… Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongsidea letter from Mark Zuckerbergarguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s hosts point out, the vision comes withsome asterisks. On this episode of TechCrunch’sEquitypodcast, Kirsten Korosec, Anthony Ha, and Rebecca Bellan take a look at Glimmer, Zuckerberg’s 6,500-word manifesto, and more of the week’s headlines, from the true cost of the AI industry’s energy needs to a $250M acquisition gone very wrong. Subscribe to Equity onYouTube,Apple Podcasts,Overcast,Spotifyand all the casts. You also can follow Equity onXandThreads, at @EquityPod.
View

Google Gemini 3.7 Flash Launched Globally With Enhanced Capabilitiesfor For Web Development and Software Engineering
Gemini 3.7 Flash, the latest AI model from the Mountain View-based tech giant, was launched globally on Thursday, the company announced. The new AI model succeeds the Gemini 3.6 Flash, which was released earlier this year, in July, along with the Gemini 3.5-Lite model. The company claims that the new Gemini 3.7 model delivers enhanced capabilities in terms of software engineering, knowledge work, and web development workflows. Currently being rolled out for developers, enterprises, and individual users, the tech giant is offering Gemini 3.7 Flash at an introductory price. The company is also upgrading Gemini Spark with the new Gemini 3 series model.
View

DeepSeek Releases V4 Pro and Open-Source Harness
The China-based AI lab has also increased its prices for the V4 family of models.
View

Samsung Strengthens India HVAC Manufacturing with New FläktGroup Facility in Pune
The new facility will manufacture cooling equipment for AI data centres, with FläktGroup planning annual production capacity of up to 6,500 HVAC units.
View

Z.ai Launches GLM-5.3 With Gains in Coding and Cybersecurity
Z.ai said the model weights will be publicly released two weeks after launch, following safety evaluation and hardening.
View

Was Oracle Right in Banning AI Contributions for OpenJDK?
An interim policy states that contributions to OpenJDK “must not include content generated, in part or in full,” using AI.
View

Karnataka Launches Tathyakosh to Make 1.5 Million Public Datasets Easier to Find
The platform developed by the Karnataka government’s Centre of Excellence in AI brings data from 458 sources.
View

Indian Banks are Rushing to Appoint Chief AI Officers
Where artificial intelligence initiatives once sat quietly within IT departments or analytics teams, lenders are now consolidating that responsibility into a single, senior role reporting directly to the top of the organisation.
View

Blacksmith Raises $45 Mn in Series B Led by Peak XV at $550 Mn Valuation
The startup said it will use the capital to expand into a broader suite of coding tools and help developers write, validate and merge software faster.
View

TCS, Vodafone Business Team Up to Drive AI-Led Digital Transformation for UK Enterprises
The partnership builds on VodafoneThree's £11 billion investment to create the UK's best network for business.
View

HCLTech, NetApp Expand Partnership to Offer Hybrid Cloud Storage-as-a-Service
The companies have expanded their partnership to offer hybrid cloud storage-as-a-service, combining consumption-based infrastructure with pay-as-you-go storage.
View
