Agent memory
Memory as a type of context: what an agent retains across sessions, built on Zep's context layer.

Defending Agent Memory Against Poisoning
One poisoned message, web page, or document can shape every later session that reads the same memory. Persistent memory makes prompt injection durable. We explain how memory poisoning works and which controls contain it, in your application and in Zep.

Why we built a graph database service for agent memory
Konig, Zep's graph database service, is the data plane beneath our agent memory platform. This post covers why we built it, how it works, and what we learned along the way.

Memory for every agent your team uses
A knowledge worker's agents don't share what they know: Claude and ChatGPT often can't reach company context, and the agents you build never see what Claude and ChatGPT learn. The Memory MCP Server gives each user one agent memory across all of them, governed by policy.

Coding agents can design your Zep implementation
Zep now ships one plugin for Claude Code, Codex, and Cursor. It gives your coding agent the Zep documentation MCP server and the new building-with-zep skill, which encodes how to design and evaluate an agent memory implementation around the use case you need to deliver.

Evaluating Nemotron 3 Embed for agent memory
Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

How Zep tracks provenance in agent memory
Agent memory is synthesized: an LLM derives facts from chat histories, documents, and business data. A derived fact matches no source word-for-word, so nothing ties it back to where it came from. Zep records that lineage; debugging, source-scoped retrieval, and compliance all depend on it.

Securing Agent Memory with ABAC
Attach policies to any Zep API key to control the exact endpoints it can call and the graph data it can read.

Unified agent memory in any MCP client
Your team's agents are split across surfaces: the desktop assistant, the coding tool, the agents you build in-house. Each keeps its own memory or none at all. The Memory MCP Server puts them on one governed user graph, gated by your enterprise single sign-on.

Agent memory placement can cut your token bill up to 2x
Agent memory has to refresh every turn, but putting it in the system prompt breaks prompt caching and re-bills the whole conversation each turn. Here is the one-message fix, and the token savings it led to in an experiment.

Markdown is not agent memory
Some of the most capable agents in production keep their memory in plain markdown files, and for a single agent and a single user it is hard to beat. The pattern breaks in predictable places: at scale, as facts change and errors compound, and under concurrent agents.

Building Agents in Go Without a Framework
A production agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human. That shape fits Go's runtime. This post explains why, surveys the Go framework options, and shows how to build an agent without one.

Sycophancy is a design choice
Writer's "Recalling Too Well" paper says memory systems amplify sycophancy. Its own data traces the amplification to two design decisions — one in Writer's experiment itself, one in a competitor's memory product.

Smart Context Assembly: Fewer Tokens, Better Quality
Today we're announcing Smart Context Assembly, an upgrade to how Zep's default Context Block is built: higher accuracy from fewer tokens, with no code changes.

Observations: Patterns and Insights from the Context Graph
Observations are a new context type in Zep that capture patterns and insights across your Context Graphs, automatically discovered and surfaced to agents.

Zep's 5 Context Types: How to Use and Combine Each One
Zep produces five distinct types of context from a user's graph. Each captures something different. Here's when to reach for each, and how to combine them in one prompt.

Stop Letting Your Agent Decide What It Needs to Know
Your agent has the tools. It just doesn't call them — and smarter models won't fix that. Here's why the unknown unknowns problem is the hardest challenge in agent context, and what to do about it.

Building Voice Agents with Memory: Zep x LiveKit
Create personalized voice agents with long-term memory with minimal added latency

What is Context Engineering, Anyway?
From Prompt Engineering to Context Engineering: The Why's and How.

The Private Agent Memory Fallacy
AI memory wallets sound appealing but face insurmountable economic, technical, and security challenges in practice.

Stop Using RAG for Agent Memory
Here's why you shouldn't be using RAG for agent memory—and what to do instead.

Introducing Entity Types: Smarter, Structured Memory for Agents
Zep's new Entity Types let developers precisely structure and recall domain-specific information for more accurate, personalized agents.

Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA in Agent Memory?
Mem0 claims State-of-the-Art in Agent Memory, but Zep outperforms it by 24%. We unpack why.

Building a Memory Agent with the OpenAI Agents SDK and Zep
A video walkthrough demonstrating using Zep's agent memory and the new OpenAI Agents SDK to build an AI agent with long-term memory.

Zep Is The New State of the Art In Agent Memory
Setting a new standard for agent memory with up to 100% accuracy gains and 90% lower latency.

Zep: A Temporal Knowledge Graph Architecture for Agent Memory
We introduce Zep, a novel memory layer service for AI agents that outperforms the current state-of-the-art systems.

Beyond Chat Memory: Making AI Interactions More Personal
Zep now connects user conversations and business data to help AI agents understand and serve users better.

Zep for Structured Outputs from Chat History (Video Walkthrough)
An end-to-end walkthrough demonstrating how to quickly and accurately extract data from chat histories stored in Zep.

Launching Structured Outputs from Chat History
Zep’s Structured Data Extraction is a high-accuracy tool for extracting data from chat histories. It's also 10x faster than gpt-4o.

Foundations of LLM App Building in TypeScript
Learn how to build three foundational LLM apps using TypeScript, LangChain.js, and Zep.

Zep ❤️ LlamaIndex: A Vector Store Walkthrough
LlamaIndex is a simple but powerful framework for building LLM apps. It's also an excellent tool for populating and searching Zep's Vector Store. This walkthrough demonstrates using LlamaIndex's new ZepVectorStore to do just that.

Introducing the Zep Document Vector Store
With the addition of a Document Vector Store, Zep is now a single, batteries-included platform for grounding LLM apps with long-term memory.

Diagnosing and Fixing Slow Chatbots with LangSmith and Zep
Poor chatbot response times can result in frustrated users and churn. Langchain’s new LangSmith service makes it easy to diagnose the cause of latency in an LLM app. In this article, we use LangSmith to analyze a very slow Langchain app and improve performance by an order of magnitude using Zep.

Personalizing LLM Interactions: Harnessing Generative Feedback Loops
LLM Applications can be personalized using Generative Feedback Loops through advanced memory & personalization.

LangchainJS Now Supports Zep!
LangchainJS now supports Zep Memory and Retrievers, allowing developers to take advantage of Zep's long-term memory, auto-summarization, vector search, and named entity extraction.

Introducing Zep: Long-term Memory Storage and Enrichment for AI Apps
Zep allows developers to focus on developing their AI apps, rather than building memory persistence, search, and enrichment infrastructure.
Sign up for Zep’s Newsletter
Writings on Zep, LLMs, and AI ecosystem tools.