We're hiring! Come build with us →
Zep

Engineering

Architecture, infrastructure and performance internals.

11 posts

Engineering

Why we built a graph database service for agent memory

Konig, Zep's graph database service, is the data plane beneath our agent memory platform. This post covers why we built it, how it works, and what we learned along the way.

Agent memory

Evaluating Nemotron 3 Embed for agent memory

Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

Agent memory

Agent memory placement can cut your token bill up to 2x

Agent memory has to refresh every turn, but putting it in the system prompt breaks prompt caching and re-bills the whole conversation each turn. Here is the one-message fix, and the token savings it led to in an experiment.

Engineering

Building Agents in Go Without a Framework

A production agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human. That shape fits Go's runtime. This post explains why, surveys the Go framework options, and shows how to build an agent without one.

Product updates

The Batch API: Load Large Datasets into Agent Memory

Zep's Batch API loads large datasets into agent memory faster, in batches up to 50,000 items, with a progress dashboard and no impact on real-time ingestion.

Engineering

How We Scaled Zep 30x in 2 Weeks (and Made It Faster)

Our infrastructure broke under 30x growth. Six weeks later, we made Zep faster than before: 10x better latency, 92% faster processing.

Graphiti

Scaling LLM Data Extraction: Challenges, Design decisions, and Solutions

How we made building Knowledge Graphs faster and more dynamic

Engineering

HNSW Indexes and Custom Prompts!

Zep v0.13.0, released today, now includes support for HNSW indexes and the ability to set custom prompts for summarization tasks.

Engineering

Session Metadata 👾, Custom OpenAI Endpoints 🔧, & Kubernetes Deployment 🔥

Zep now includes support for adding arbitrary metadata to Sessions and more deployment options

Engineering

Introducing Open Source Embeddings! 🔥

Zep now has support for open source embedding models. Search your LLM App's chat history faster and more cheaply.

Benchmarks and evaluation

A Survey of Embedding Models (and why you should look beyond OpenAI)

User experience can be severely impacted when AI chat apps and agents are slow to respond. Embedding model performance can contribute to this problem. Is OpenAI's embedding API always the best fit?

Sign up for Zep’s Newsletter

Writings on Zep, LLMs, and AI ecosystem tools.

No spam. Unsubscribe anytime.