Benchmarks and evaluation
Measuring accuracy, latency and token cost.

Evaluating Nemotron 3 Embed for agent memory
Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

Sycophancy is a design choice
Writer's "Recalling Too Well" paper says memory systems amplify sycophancy. Its own data traces the amplification to two design decisions — one in Writer's experiment itself, one in a competitor's memory product.

Evaluation and Control: Evaluation Framework, zepctl CLI, and Dashboard Overhaul
An evaluation framework for testing Zep against your data, a redesigned dashboard with analytics, and zepctl — a CLI for administering Zep projects.

The Retrieval Tradeoff: What 50 Experiments Taught Us About Context Engineering
Zep builds temporal knowledge graphs from conversations and business data, then automatically retrieves relevant context when your agent needs it.

Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA in Agent Memory?
Mem0 claims State-of-the-Art in Agent Memory, but Zep outperforms it by 24%. We unpack why.

GPT-4.1 and o4-mini: Is OpenAI Overselling Long-Context?
We put OpenAI’s latest models through the LongMemEval benchmark—here’s why raw context size alone isn't enough.

Zep Q1 Product Round-up
Graph Explorer, an upgraded Playground, and a practical Cookbook—giving developers more control, better graph quality, and improved usability.

Zep Is The New State of the Art In Agent Memory
Setting a new standard for agent memory with up to 100% accuracy gains and 90% lower latency.

Zep: A Temporal Knowledge Graph Architecture for Agent Memory
We introduce Zep, a novel memory layer service for AI agents that outperforms the current state-of-the-art systems.

A Survey of Embedding Models (and why you should look beyond OpenAI)
User experience can be severely impacted when AI chat apps and agents are slow to respond. Embedding model performance can contribute to this problem. Is OpenAI's embedding API always the best fit?
Sign up for Zep’s Newsletter
Writings on Zep, LLMs, and AI ecosystem tools.