KV Eviction

Long-context AI, without the GPU tax.

We build infrastructure that shrinks the memory LLMs need at inference — so startups can ship longer-context products on smaller, cheaper GPUs without giving up quality.

The market problem

Context is the new bottleneck.

Every serious AI product wants longer context — docs, chats, agents, reasoning. But memory grows with every token, and GPU spend scales with it. That puts long-context features out of reach for most teams.

The winners won’t just have bigger models. They’ll run more context per dollar.

What we’re building

A layer that makes long context affordable.

KV Eviction compresses an LLM’s working memory during generation so products can keep longer prompts practical on real cloud hardware — same answers, far less memory.

  1. 01

    Cut infrastructure cost

    Run longer prompts on the GPUs you already have — or the same workload on smaller, cheaper instances.

  2. 02

    Ship without quality tradeoffs

    Cost savings only matter if users still get correct answers. Quality under tight memory budgets is the product bar.

  3. 03

    Built for production workloads

    Aimed at documents, multi-step agents, and long generations — the use cases that actually burn GPU budget.

Why AWS

We need GPU scale to ship this.

Validating memory savings across real models and workloads takes serious GPU capacity. AWS Activate credits fund the compute to benchmark, harden, and move from research prototype to production-ready infrastructure.

Category

AI infrastructure · inference efficiency

Outcome

Longer context · lower GPU cost

Stage

Early startup · active R&D

Team

Building the memory layer for LLMs.

KV Eviction is founded out of Universidad de San Andrés to make long-context transformers cheaper to serve. Technical details stay private until we’re ready to publish and ship.

Founder

Alex Bodner

Computer Science · Universidad de San Andrés

me@alexbodner.com