Engineering
Architecture, infrastructure and performance internals.

Why we built a graph database service for agent memory
Konig, Zep's graph database service, is the data plane beneath our agent memory platform. This post covers why we built it, how it works, and what we learned along the way.

Evaluating Nemotron 3 Embed for agent memory
Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

Agent memory placement can cut your token bill up to 2x
Agent memory has to refresh every turn, but putting it in the system prompt breaks prompt caching and re-bills the whole conversation each turn. Here is the one-message fix, and the token savings it led to in an experiment.

Building Agents in Go Without a Framework
A production agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human. That shape fits Go's runtime. This post explains why, surveys the Go framework options, and shows how to build an agent without one.

The Batch API: Load Large Datasets into Agent Memory
Zep's Batch API loads large datasets into agent memory faster, in batches up to 50,000 items, with a progress dashboard and no impact on real-time ingestion.

How We Scaled Zep 30x in 2 Weeks (and Made It Faster)
Our infrastructure broke under 30x growth. Six weeks later, we made Zep faster than before: 10x better latency, 92% faster processing.

Scaling LLM Data Extraction: Challenges, Design decisions, and Solutions
How we made building Knowledge Graphs faster and more dynamic

HNSW Indexes and Custom Prompts!
Zep v0.13.0, released today, now includes support for HNSW indexes and the ability to set custom prompts for summarization tasks.

Session Metadata 👾, Custom OpenAI Endpoints 🔧, & Kubernetes Deployment 🔥
Zep now includes support for adding arbitrary metadata to Sessions and more deployment options

Introducing Open Source Embeddings! 🔥
Zep now has support for open source embedding models. Search your LLM App's chat history faster and more cheaply.

A Survey of Embedding Models (and why you should look beyond OpenAI)
User experience can be severely impacted when AI chat apps and agents are slow to respond. Embedding model performance can contribute to this problem. Is OpenAI's embedding API always the best fit?
Sign up for Zep’s Newsletter
Writings on Zep, LLMs, and AI ecosystem tools.