KV Eviction
Long-context AI, without the GPU tax.
We build infrastructure that shrinks the memory LLMs need at inference — so startups can ship longer-context products on smaller, cheaper GPUs without giving up quality.
The market problem
Context is the new bottleneck.
Every serious AI product wants longer context — docs, chats, agents, reasoning. But memory grows with every token, and GPU spend scales with it. That puts long-context features out of reach for most teams.
The winners won’t just have bigger models. They’ll run more context per dollar.
What we’re building
A layer that makes long context affordable.
KV Eviction compresses an LLM’s working memory during generation so products can keep longer prompts practical on real cloud hardware — same answers, far less memory.
-
01
Cut infrastructure cost
Run longer prompts on the GPUs you already have — or the same workload on smaller, cheaper instances.
-
02
Ship without quality tradeoffs
Cost savings only matter if users still get correct answers. Quality under tight memory budgets is the product bar.
-
03
Built for production workloads
Aimed at documents, multi-step agents, and long generations — the use cases that actually burn GPU budget.
Why AWS
We need GPU scale to ship this.
Validating memory savings across real models and workloads takes serious GPU capacity. AWS Activate credits fund the compute to benchmark, harden, and move from research prototype to production-ready infrastructure.
Category
AI infrastructure · inference efficiency
Outcome
Longer context · lower GPU cost
Stage
Early startup · active R&D
Team
Building the memory layer for LLMs.
KV Eviction is founded out of Universidad de San Andrés to make long-context transformers cheaper to serve. Technical details stay private until we’re ready to publish and ship.