• Kog says software can unlock 30x faster LLM inference on existing GPUs
  • Morgan Stanley Boosts Bitcoin and Ethereum Fund Holdings, Trims Coinbase Stake in Q2
  • Canton Strategic Holdings, Inc. Releases Second Quarter 2026 Financial and Operational Results
  • Colombia Retail Sales Jump 14.7% in June, Beating Market Expectations
  • Deeper Insights, Better Analytics: Bybit Options Close the Data Gap Between Retail and Institutional Traders
2026-08-14
Coins by Cryptorank
Bitcoinworld Bitcoinworld
Bitcoinworld Bitcoinworld
  • Crypto News
  • AI News
  • Forex News
  • Sponsored
  • Press Release
  • Media Kit
  • Advertisement
  • More
    • About Us
    • Learn
    • Exclusive Article
    • Reviews
    • Events
    • Contact Us
    • Privacy Policy
Bitcoinworld
  • Crypto News
  • AI News
  • Forex News
  • Sponsored
  • Press Release
  • Media Kit
  • Advertisement
  • More
    • About Us
    • Learn
    • Exclusive Article
    • Reviews
    • Events
    • Contact Us
    • Privacy Policy
Skip to content
Home AI News Kog says software can unlock 30x faster LLM inference on existing GPUs
AI News

Kog says software can unlock 30x faster LLM inference on existing GPUs

  • by Keshav Aggarwal
  • 2026-08-14
  • 0 Comments
  • 3 minutes read
  • 0 Views
  • 17 seconds ago
Facebook Twitter Pinterest Whatsapp
Close-up of a GPU server in a data center with glowing lights, representing high-performance computing.

French startup Kog is betting that software optimization can dramatically accelerate large language model inference on standard data center GPUs, claiming up to 30x speed gains without requiring new hardware. The company, which emerged from stealth in May 2025, says it has already attracted over 200 business leads, signaling strong enterprise demand for faster AI response times.

Why inference speed matters

As AI models become integral to professional workflows, the time it takes for a model to generate a response has become a critical bottleneck. Slow inference not only hampers productivity but also increases operational costs, especially for applications like AI-assisted coding, where users may wait minutes or even hours for results. Anthropic, for example, charges a premium for its ‘Fast Mode’ on Claude, underscoring the value of speed.

Kog’s approach targets this pain point directly. The company’s technology, dubbed the Kog Inference Engine (KIE), is designed to optimize the decoding process on existing GPUs from NVIDIA and AMD, such as the H200 and MI300X. In a technical preview that reached the front page of Hacker News, Kog demonstrated 3,000 tokens per second on a single request using a small 2-billion-parameter model, which they open-sourced as Laneformer 2B.

From small models to large language models

While the initial demo used a small model, Kog’s CEO Gaël Delalleau says the company is now focused on scaling its techniques to large language models, a significant technical leap. “We’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen,” Delalleau told Bitcoin World. The company has learned that most potential customers are not interested in fine-tuning small models for their specific tasks, making support for mainstream LLMs essential.

Delalleau is confident that the same optimization principles will work with larger models, despite skepticism from some quarters. He argues that modern GPUs have abundant memory bandwidth that is underutilized by current inference software. “GPUs have a bright future,” he said, rejecting the notion that they are ill-suited for decoding.

How Kog’s approach differs

Kog is not alone in exploring software-based inference acceleration. ZML, another French startup, has developed hardware-agnostic software that bypasses NVIDIA’s CUDA to support fast inference across competing chips. However, Delalleau draws a distinction, comparing Kog’s work to Stanford’s Hazy Research lab, with a deeper focus on GPU-level engineering.

This deep-level focus stems from Delalleau’s background in solid-state physics and offensive cybersecurity. He explains that this combination of understanding the laws of hardware and reverse-engineering at the assembly level allows his team to push GPUs beyond their intended performance envelopes. But this approach is time-consuming; for each new GPU, the team dedicates weeks or months to detailed engineering research.

Market readiness and future plans

Kog’s initial market focus is on software engineering, where slow inference is a known pain point. The company also has design partners in the app and game generation space, where faster output translates directly to revenue. However, the market for ultra-fast inference is still maturing, and Kog is adapting its roadmap accordingly.

The company, which has a team of 11, is backed by French institutions including Bpifrance and the French Tech 2030 program, and is supported by cloud provider Scaleway. Delalleau anticipates that demonstrating a major model running at 10x speed, possibly by September 2025, will be key to securing a Series A round. “Once we’ve implemented our first major model at 10x speed, we’ll be able to start demonstrating customer traction,” he said.

Conclusion

Kog’s promise to unlock more inference performance from existing GPUs addresses a pressing need in the AI industry. If successful, it could offer a cost-effective alternative to specialized hardware, giving enterprises a way to accelerate AI workloads without significant capital expenditure. The coming months will be critical for the startup to prove its technology at scale.

FAQs

Q1: What is Kog’s core technology?
Kog’s Kog Inference Engine (KIE) is a software layer that optimizes the decoding process of large language models on standard data center GPUs, aiming to deliver up to 30x faster inference speeds compared to conventional methods.

Q2: How does Kog achieve faster inference?
Kog uses deep-level GPU engineering techniques, including low-level reverse engineering and hardware-aware optimizations, to better utilize the memory bandwidth and compute capabilities of existing GPUs from NVIDIA and AMD.

Q3: When will Kog’s technology be available for large models?
Kog is currently focused on scaling its approach to large language models. The company expects to demonstrate a major model running at 10x speed by September 2025, with broader availability likely following a successful Series A round.

Disclaimer: The information provided is not trading advice, Bitcoinworld.co.in holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Related Reading

  • France Inflation Matches Forecasts: July CPI (EU Norm) Rises 0.6% Month-on-Month
  • France Inflation Ex-Tobacco Rises to 0.6% in July, Reversing June Dip
  • Databricks wanted to raise $1B, investors offered $15B — it settled on $5B at a $190B valuation
  • OpenAI unveils Ultrafast mode for GPT-5.6 Sol, claiming 14x speed boost
  • Cognition reportedly in talks to raise at $40B valuation, three months after $1B round

Tags:

AI inferenceFranceGPU optimizationKogStartups

Share This Post:

Facebook Twitter Pinterest Whatsapp
Avatar photo

Keshav Aggarwal

Co- Founder
Keshav Aggarwal is the Co-Founder & CEO of BitcoinWorld, a Google News - indexed publication covering crypto, AI, and forex markets since 2020. A blockchain investor and trader with over six years in the digital-asset space, he built one of India's most active crypto investor communities and has guided thousands of retail participants through their first investments in the asset class. At BitcoinWorld, he sets editorial direction across the newsroom and reports on the business of crypto, AI, and Web3 - tracking the funding rounds, product launches, and regulatory shifts shaping the future of finance and frontier technology.
Next Post

Morgan Stanley Boosts Bitcoin and Ethereum Fund Holdings, Trims Coinbase Stake in Q2

Categories

92

AI News

Crypto News

Bitcoin Treasury Ambition: The Blockchain Group Seeks Staggering €10 Billion

Events

97

Forex News

33

Learn

Press Release

Reviews

Google NewsGoogle News TwitterTwitter LinkedinLinkedin coinmarketcapcoinmarketcap BinanceBinance YouTubeYouTubes

Copyright © 2026 BitcoinWorld | Powered by BitcoinWorld