We're hiring! Come build with us →
Zep

Retrieval

Search, reranking, context assembly and latency.

13 posts

Agent memory

Agent memory placement can cut your token bill up to 2x

Agent memory has to refresh every turn, but putting it in the system prompt breaks prompt caching and re-bills the whole conversation each turn. Here is the one-message fix, and the token savings it led to in an experiment.

Retrieval

Smart Context Assembly: Fewer Tokens, Better Quality

Today we're announcing Smart Context Assembly, an upgrade to how Zep's default Context Block is built: higher accuracy from fewer tokens, with no code changes.

Context engineering

Zep's 5 Context Types: How to Use and Combine Each One

Zep produces five distinct types of context from a user's graph. Each captures something different. Here's when to reach for each, and how to combine them in one prompt.

Governance and security

Context You Can Trace, Filter, and Trust

Every fact in your agent's context graph came from somewhere. Zep's provenance architecture traces facts back to their source data — and lets you filter retrieval by origin.

Context engineering

3 Decisions That Shape Every Agent's Context Architecture

Every agent context architecture comes down to three decisions: scope, data sources, and retrieval strategy. A framework for reasoning about persistent context for AI agents.

Benchmarks and evaluation

The Retrieval Tradeoff: What 50 Experiments Taught Us About Context Engineering

Zep builds temporal knowledge graphs from conversations and business data, then automatically retrieves relevant context when your agent needs it.

Engineering

How We Scaled Zep 30x in 2 Weeks (and Made It Faster)

Our infrastructure broke under 30x growth. Six weeks later, we made Zep faster than before: 10x better latency, 92% faster processing.

Context engineering

Zep v3: Context Engineering Takes Center Stage

Context engineering > prompt engineering: Zep v3 assembles memory & business data for agents that work.

Retrieval

The One-Token Trick

How single-token LLM requests can improve RAG search at minimal cost and latency.

Retrieval

How do you search a Knowledge Graph?

Can you build graph search that's both elegant & powerful? Here's how we did it in Graphiti, Zep's open source temporal Knowledge Graph library.

Context graphs

Building A Russian Election Interference Knowledge Graph

How we built a Knowledge Graph-based app for exploring Russian interference in the run-up to the 2024 US elections.

Agent memory

Beyond Chat Memory: Making AI Interactions More Personal

Zep now connects user conversations and business data to help AI agents understand and serve users better.

Retrieval

Introducing Zep Hybrid Search and Custom Metadata

Zep now supports both vector search over message text and filtering on message metadata, including system metadata such as Named Entities and creation dates.

Sign up for Zep’s Newsletter

Writings on Zep, LLMs, and AI ecosystem tools.

No spam. Unsubscribe anytime.