# How We Scaled Zep 30x in 2 Weeks (and Made It Faster)

By Daniel Chalef · Nov 06, 2025 · https://www.getzep.com/blog/scaling-agent-memory-zep-30x/ · Tags: Engineering, Product updates, Retrieval

Our infrastructure broke under 30x growth. Six weeks later, we made Zep faster than before: 10x better latency, 92% faster processing.

This summer, our largest enterprise customers landed all at once. Service usage increased 30x in two weeks. We went from handling thousands of hourly requests to millions overnight.

Our infrastructure didn't just bend under the load. It broke. Hard. 😬

Context retrieval latency spiked from a P95 of 200ms to over 2 seconds. Episode processing crawled to 60 seconds. LLM costs exploded to 3-5x our provisioned capacity, with rate limit errors cascading to customer-facing failures. Throwing more hardware at our graph database couldn't stop the system from melting down.

![Timeline of P95 latency: 200ms at baseline, over 2000ms during the summer crisis (10x degradation), and back to 150-200ms six weeks later.](https://www.getzep.com/blog/scaling-agent-memory-zep-30x/summer-of-fun-1.svg)

Context retrieval sits on the critical path for every agent interaction. When it takes 2 seconds instead of 200ms, end users wait. When it fails, agents hallucinate. We were breaking production for customers with millions of users.

Six weeks later, we not only stabilized Zep but made it significantly faster than before the crisis. Graph search now returns results in 150ms (P95), down from 600ms. Context retrieval, which is a retrieval pipeline with many search components, dropped to 200ms (P95). And episode processing latency improved 92%, from around to 4 seconds. And we cut LLM token usage in half while maintaining accuracy above 80% on the LongMemEval benchmark, up 10% since our [early 2025 paper](https://www.getzep.com/blog/state-of-the-art-agent-memory/).

Here's how we did it.

[Zep Realtime Demo](https://www.youtube.com/embed/zOhzk0Gdyv0?feature=oembed)

Our virtual Staff Engineer. A realtime application built with Zep and [Anam AI](https://anam.ai/). 💥

## The Challenge: When Growth Exposes Bottlenecks

Over eight weeks this summer, we onboarded our biggest enterprise customers. These weren't gradual rollouts. Customers exceeded their own growth forecasts. Service volume increased 30x in just two weeks.

The rapid scaling revealed bottlenecks we didn't know existed:

**Our graph database was doing too much.** We'd been using our graph database to handle graph operations, vector search, and BM25 full-text search. We were adding new vectors and texts too fast, all day long. Under 30x load, search latencies spiked dramatically. The system was trying to be a Swiss Army knife when we needed specialized tools.

![Line chart of LLM token throughput spiking to millions per second during the load surge, then settling near provisioned capacity.](https://www.getzep.com/blog/scaling-agent-memory-zep-30x/image-2.png)

Millions of Tokens Per Second 🤯

**LLM costs spiraled out of control.** We use LLMs for graph construction, entity extraction, relationship inference, and fact deduplication. Burst usage hit 3-5x our provisioned throughput, triggering rate limit errors that cascaded to customer applications. We were dramatically underwater on margins.

![Kubernetes memory chart with sawtooth lines climbing toward the 2 GB limit and resetting, showing LLM gateway pods repeatedly leaking.](https://www.getzep.com/blog/scaling-agent-memory-zep-30x/image-3.png)

Too Many Tokens, Too Many Memory Leaks

**Our Python-based LLM gateway choked.** We used an LLM Gateway to route calls, program fallbacks, and for granular cost tracking. The proxy ran on 8-12 pods with high memory footprints and at times significant CPU load. The final straw were memory leaks resulting in service outages every few hours. The operational overhead was crushing us.

These weren't isolated issues. Everything broke simultaneously, and our customers' production systems were on the line. We had non-negotiable SLA targets. Missing these meant breaking production for customers with millions of users.

## The Solution: A Multi-Pronged Technical Overhaul

We took a systematic approach to rearchitecting under fire. The strategy came down to three principles: separate concerns early, question every LLM call, and build for burst traffic instead of average load.

**Specialized search infrastructure.** Our graph database was handling graph operations, vector search, and BM25 full-text search all at once. We needed to separate these concerns and use purpose-built tools for each job. We considered OpenSearch but estimated it would cost over $50K per month in infrastructure. Instead, we moved to a high-performance, very high-scale retrieval infrastructure optimized for our use case. We offloaded all vector and BM25 search to this dedicated service, stripped content storage from the graph database, and let our graph database focus solely on what it does best: graph operations. The result: 65ms P75 search latency at massive scale, a fraction of the cost, and far lower operational overhead.

![Query latency chart for a hot namespace over several hours: median near 12ms, with p95 and max spiking but mostly staying below 100ms.](https://www.getzep.com/blog/scaling-agent-memory-zep-30x/image.png)

This looks much better! 💜

**Classical NLP and Information Retrieval techniques over LLM calls.** We realized we'd been treating LLMs as the solution to every problem, even when simpler approaches would work better. We replaced expensive LLM calls with Shannon Entropy for information density scoring, TF-IDF for content deduplication, and LSH (Locality-Sensitive Hashing) for fast similarity matching. We compressed context before LLM calls by stripping low-signal content and optimized prompts to reduce token overhead. This cut LLM token usage by 50% while keeping costs within our prepurchased inference capacity.

![Python code from Graphiti defining a name entropy function using Shannon entropy over characters to defer low-entropy names to the LLM.](https://www.getzep.com/blog/scaling-agent-memory-zep-30x/image-1.png)

Scoring Shannon Entropy. Excerpt from [Graphiti](https://github.com/getzep/graphiti), Zep's open source graph framework.
