We're hiring! Come build with us →
Zep

Benchmarks and evaluation

Measuring accuracy, latency and token cost.

10 posts

Agent memory

Evaluating Nemotron 3 Embed for agent memory

Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

Benchmarks and evaluation

Sycophancy is a design choice

Writer's "Recalling Too Well" paper says memory systems amplify sycophancy. Its own data traces the amplification to two design decisions — one in Writer's experiment itself, one in a competitor's memory product.

Benchmarks and evaluation

Evaluation and Control: Evaluation Framework, zepctl CLI, and Dashboard Overhaul

An evaluation framework for testing Zep against your data, a redesigned dashboard with analytics, and zepctl — a CLI for administering Zep projects.

Benchmarks and evaluation

The Retrieval Tradeoff: What 50 Experiments Taught Us About Context Engineering

Zep builds temporal knowledge graphs from conversations and business data, then automatically retrieves relevant context when your agent needs it.

Benchmarks and evaluation

Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA in Agent Memory?

Mem0 claims State-of-the-Art in Agent Memory, but Zep outperforms it by 24%. We unpack why.

Benchmarks and evaluation

GPT-4.1 and o4-mini: Is OpenAI Overselling Long-Context?

We put OpenAI’s latest models through the LongMemEval benchmark—here’s why raw context size alone isn't enough.

Product updates

Zep Q1 Product Round-up

Graph Explorer, an upgraded Playground, and a practical Cookbook—giving developers more control, better graph quality, and improved usability.

Benchmarks and evaluation

Zep Is The New State of the Art In Agent Memory

Setting a new standard for agent memory with up to 100% accuracy gains and 90% lower latency.

Agent memory

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

We introduce Zep, a novel memory layer service for AI agents that outperforms the current state-of-the-art systems.

Benchmarks and evaluation

A Survey of Embedding Models (and why you should look beyond OpenAI)

User experience can be severely impacted when AI chat apps and agents are slow to respond. Embedding model performance can contribute to this problem. Is OpenAI's embedding API always the best fit?

Sign up for Zep’s Newsletter

Writings on Zep, LLMs, and AI ecosystem tools.

No spam. Unsubscribe anytime.