Kog is going deeper to squeeze more inference out of GPUs
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
Researched and edited by Kiran Ch and the WhatIsFuture editorial team. Reviewed for factual accuracy before publication.
Kog is Going Deeper to Squeeze More Inference Out of GPUs
The prevailing industry consensus suggests that modern tensor processing architectures are fundamentally mismatched with autonomous agentic workloads. While GPUs excel at dense, highly predictable matrix multiplications across batch-dense training runs, agent loops introduce dynamic branching, erratic context windows, multi-turn tool calling, and high-frequency state lookups. This architectural friction has led many platform teams to prematurely blame hardware constraints for the latency spikes and exorbitant compute bills plaguing their production deployments.
French infrastructure startup Kog is taking a contrarian stance by proving that GPUs are not inherently inefficient for agentic workloads—the abstraction layers built on top of them are. By bypassing high-level inference runtimes and rewriting scheduling, memory paging, and kernel dispatch logic from bare metal upward, systems engineers can extract unprecedented throughput from existing accelerator clusters without waiting for specialized ASICs.
Join Our Tech Community
Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just signal.
The Architecture Fallacy: Why Agent Loops Choke Modern Silicon
To understand why autonomous agent loops degrade performance on modern silicon, we must look at the mathematical nature of transformer execution. LLM inference operates in two distinct phases: the prefill phase (which is compute-bound, processing prompt tokens in parallel using high-performance General Matrix Multiply or GEMM operations) and the decode phase (which is memory-bound, generating one token at a time using Vector Matrix Multiply or GEMV operations). In classic chatbot applications, a single prefill is followed by a predictable stream of decode steps.
When multi-agent architectures stall, however, the failure mode is almost always memory-bound rather than compute-bound. Traditional inference serving frameworks were optimized for standard user-facing request-response cycles with linear token generation. In contrast, compound AI systems generate complex directed acyclic graphs (DAGs) of execution: parallel sub-agent queries, tool payload evaluations, self-correction loops, and recursive retrieval passes. As observed when Anthropic set AI agents loose on the same task, token consumption explodes while GPU compute utilization frequently dips below 25% due to scheduling gridlock.
This gridlock occurs because agent loops are fundamentally asynchronous and non-linear. An agent does not simply generate text; it pauses to call an external database, waits for a Python sandbox to execute a script, inspects the error logs, adjusts its internal prompt, and continues generating. During these idle tool-use pauses, the context must either be held in precious GPU memory (starving other active processes) or swapped out to host memory and recompiled later. This constant, unpredictable cycle of context thrashing turns the GPU into an incredibly expensive waiting room.
The root cause lies in Key-Value (KV) cache fragmentation and DRAM bandwidth saturation. In naive agent harnesses, every tool-use step or context append triggers full context recompilation or inefficient cache eviction. High-end accelerators like the NVIDIA H100 SXM5 boast 3.35 TB/s of memory bandwidth, but when small, fragmented batch sizes force continuous context swapping between high-bandwidth memory (HBM3) and host RAM, memory interconnects choke. Modern GPU compute engines sit idle waiting for memory transfers, masking fundamental software architecture deficiencies as a hardware ceiling.
Kernel-Level Surgery: Squeezing Low-Latency Compute from Bare Metal
Kog's engineering intervention focuses on the execution runtime rather than model parameter counts. Instead of accepting generic CUDA graph captures that break whenever context length dynamically shifts, specialized inference engines are implementing granular token-level scheduling and dynamic prefix caching tailored specifically for agent branching. CUDA graphs are traditionally used to bypass the CPU overhead of launching individual kernels by capturing a static sequence of operations. However, because agent workloads change their execution path and context size dynamically with every external tool call, static CUDA graphs break, forcing slow recompilations or fallback to CPU-controlled kernel launches.
To resolve this, Kog bypasses standard runtime layers like Triton or PyTorch-level abstractions altogether. They build custom, low-level scheduling loops directly inside bare-metal CUDA C++ and Triton codebases. When multiple agents share a common system prompt, base tool definitions, or retrieved codebase index, the underlying kernel must treat that memory region as immutable and persistent across concurrent worker threads. This means the engine does not duplicate the shared prompt in HBM for every running instance; instead, it uses pointer-level sharing, routing multiple parallel execution streams through the exact same physical memory layout.
Similar to enterprise efforts where organizations deploy an upgraded harness to contain token costs and manage execution overhead, optimizing the lower-level hardware abstraction yields dramatic ROI. By deploying custom Triton and CUDA kernels designed for non-linear graph execution, systems can perform in-place cache mutation without triggering memory reallocations. This eliminates pipeline bubbles—moments when the tensor cores stop computing while waiting for the control CPU to schedule the next batch of instructions—and keeps tensor cores saturated even during asynchronous tool invocation pauses.
Furthermore, this kernel-level optimization targets warp-level parallelism (where groups of 32 parallel threads, or "warps," execute instructions simultaneously on a Streaming Multiprocessor). In typical transformers, threads within a warp operate in tight synchronization. However, when branches occur within agent logic—such as when one token triggers a stop sequence for tool usage while another continues generating—divergent execution paths occur. Kog's custom memory paging allocates memory blocks explicitly mapped to the warp architecture, ensuring that divergent branching does not lead to serialized instruction execution, preserving the massive parallel computing advantages of the underlying silicon.
"The narrative that general-purpose GPUs are inadequate for autonomous agents stems from treating the GPU as a black-box token generator rather than a programmable parallel runtime. If you control memory layout and instruction scheduling down to the warp level, you unlock massive headroom without swapping silicon."
Latency Asymmetry and the Agentic Feedback Loop
In single-turn conversational AI, user experience is governed primarily by inter-token latency (ITL)—the time elapsed between subsequent tokens appearing on a screen. Human readers require a steady, moderate flow of text (around 20 to 50 tokens per second) to feel that a system is responsive. In agentic loops, however, human presence is removed from the immediate loop. Consequently, Inter-Token Latency is irrelevant during internal reasoning steps. Instead, Time-To-First-Token (TTFT) and overall end-to-end task completion time dominate the architecture.
When an autonomous workflow requires 30 sequential reasoning steps and 15 programmatic function calls, a 300ms latency floor per call compounds into multi-minute execution times. For enterprise use cases like automated customer service, autonomous software engineering agents, or real-time trading copilots, multi-minute execution times are unacceptable. Proprietary cloud providers have pushed speed aggressively—such as when OpenAI introduces Ultrafast mode to accelerate token generation—but private infrastructure engineers operating in secure virtual private clouds (VPCs) must solve this within their own bare-metal constraints.
Bridging this gap requires rethinking continuous batching and speculative decoding for agent trees. When an agent generates a deterministic payload like a JSON schema or SQL query, speculative drafting algorithms can predict tool invocation syntax with near-zero compute cost. Instead of running a massive 70-billion parameter model to generate standard syntax like {"action": "query", "parameters": {}}, a lightweight draft model (often under 1 billion parameters) predicts these highly structured tokens. The primary GPU model is then used only to verify these draft tokens in a single, parallelized forward pass, rather than generating them step-by-step.
This decoupling of compute-heavy reasoning from deterministic structural overhead reduces execution latency across multi-step execution graphs. By aligning the speculative model specifically to the target tool's JSON schema, the system can bypass traditional generation entirely for structural tokens, accelerating the path to tool execution and reducing overall loop latency by up to 5x.
Deep Dive: Defragmenting the KV Cache in Non-Linear Trees
To grasp the engineering depth of Kog's approach, one must look closely at how the Key-Value (KV) cache behaves in non-linear agent tasks. During transformer-based generation, key and value states are computed for each token and stored in memory to prevent the model from recalculating them at every new step. In a standard linear chat session, this cache grows sequentially, like a single thread of string.
In an agent loop, however, the path of reasoning branches like a tree. The agent might try Method A, find that it fails, backtrack to the original prompt, and try Method B. In a conventional inference server, backtracking requires either throwing away the cache of Method A and recompiling the original context (wasting compute), or keeping both Method A and Method B caches in separate, unlinked memory segments (wasting HBM).
Kog solves this through a physical-to-virtual memory mapping layer that mimics the virtual memory systems of classic operating systems. Instead of contiguous physical memory allocations, the KV cache is broken up into tiny, fixed-size physical blocks (e.g., block sizes representing 16 tokens). A virtual lookup table translates the model’s logical view of context into these physical pages.
When the agent branches from Method A to Method B, the system does not copy any memory. Instead, it marks the virtual page table of Method B to point to the exact same physical blocks as Method A up to the point of divergence, creating a new virtual branch that shares the underlying parent physical blocks. When Method B writes its first unique token, a "copy-on-write" mechanism allocates a new physical block specifically for that branch. This prevents memory bloat, ensures near-zero overhead for context backtracking, and maintains low latency during complex tree-search algorithms like Monte Carlo Tree Search (MCTS).
Key Architectural Heuristics for Agent Inference
Optimizing hardware performance for autonomous loops requires shifting from general-purpose LLM hosting patterns to agent-first infrastructure designs. Teams looking to run low-latency agent workloads inside their own enterprise clouds should prioritize the following four heuristics:
- Persistent Global Prefix Caching: Retain base agent system prompts, schemas, and static context directly in L2 cache or physical HBM. This eliminates prefill computation overhead by up to 70% during continuous multi-turn loops, ensuring that the model never recomputes unchanged system instructions.
- Asynchronous Dynamic Batching: Decouple tool execution wait states from active token generation pipelines. When an agent pauses to execute an API call or SQL query, its memory slot must be suspended immediately, freeing active compute warp resources for other running tasks without forcing a full host-to-device memory eviction.
- Kernel-Level Speculative Verification: Utilize highly specialized, open-weight speculative draft models to predict highly structured syntactic tool calls. Validate these drafts against the main model in single-pass forward sweeps, bypassing the need for autoregressive decoding of predictable formatting characters.
- Warp-Level Context Swapping: Minimize host-to-device memory roundtrips through pinned unified memory strategies. By bypassing standard OS virtual memory allocations and utilizing direct memory access (DMA) paths over PCIe Gen5, physical context pages can transition between host RAM and GPU HBM3 at line speed.
The Bottom Line
The race to deploy resilient, fast, and production-grade autonomous agents will not be won by simply throwing more H100s or B200s at unoptimized inference pipelines, nor by prematurely abandoning GPUs for unproven, specialized custom hardware. The actual bottleneck is not silicon throughput, but rather the inefficient software wrappers that hide hardware control from application developers.
Software teams that take direct control of low-level memory allocators, continuous scheduling runtimes, and dynamic prefix caching will slash their serving latency and operating expenses by an order of magnitude. For enterprise systems architects, the mandate is clear: stop treating inference runtimes as commoditized wrappers, and start engineering the hardware-software interface as a core competitive edge.
Supercharge Your Workflow with Claude AI
The AI assistant used by professionals worldwide. Write, code, analyse — all in one place.