Kog Speeds AI Inference by Reworking GPU Software
Paris-based Kog, founded in 2023 by École Polytechnique graduate Gaël Delalleau, is betting that software rather than silicon is preventing conventional data-center GPUs from reaching their inference potential. The challenge is becoming more important as coding agents and other generative-AI systems execute long, sequential loops in which every planning, testing and revision step depends on the last. Faster single-request decoding could shorten response times and reduce the infrastructure cost of such professional workflows.
Kog released a technology preview of the Kog Inference Engine, or KIE, on May 28, 2026. The company said a 2B model produced 3,000 output tokens a second on eight AMD MI300X GPUs and 2,100 on eight NVIDIA H200 chips, using FP16 without speculative decoding. Kog has raised $5 million from Varsity VC and Bpifrance’s Deep Tech Program. Its pitch has drawn growing enterprise interest as it works to extend the system to larger third-party mixture-of-experts models.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →