Zep blog

Defending Agent Memory Against Poisoning
One poisoned message, web page, or document can shape every later session that reads the same memory. Persistent memory makes prompt injection durable. We explain how memory poisoning works and which controls contain it, in your application and in Zep.
Latest posts

Why we built a graph database service for agent memory
Konig, Zep's graph database service, is the data plane beneath our agent memory platform. This post covers why we built it, how it works, and what we learned along the way.

Memory for every agent your team uses
A knowledge worker's agents don't share what they know: Claude and ChatGPT often can't reach company context, and the agents you build never see what Claude and ChatGPT learn. The Memory MCP Server gives each user one agent memory across all of them, governed by policy.

Coding agents can design your Zep implementation
Zep now ships one plugin for Claude Code, Codex, and Cursor. It gives your coding agent the Zep documentation MCP server and the new building-with-zep skill, which encodes how to design and evaluate an agent memory implementation around the use case you need to deliver.

Evaluating Nemotron 3 Embed for agent memory
Zep's graph retrieval starts from entrypoints selected by semantic and BM25 search, so embedding recall bounds everything downstream. We benchmarked NVIDIA's Nemotron 3 Embed 1B, released today, on 5,954 production recall queries: first place against our production incumbent and two other models.

How Zep tracks provenance in agent memory
Agent memory is synthesized: an LLM derives facts from chat histories, documents, and business data. A derived fact matches no source word-for-word, so nothing ties it back to where it came from. Zep records that lineage; debugging, source-scoped retrieval, and compliance all depend on it.

Securing Agent Memory with ABAC
Attach policies to any Zep API key to control the exact endpoints it can call and the graph data it can read.
Sign up for Zep’s Newsletter
Writings on Zep, LLMs, and AI ecosystem tools.
More posts

Unified agent memory in any MCP client
Your team's agents are split across surfaces: the desktop assistant, the coding tool, the agents you build in-house. Each keeps its own memory or none at all. The Memory MCP Server puts them on one governed user graph, gated by your enterprise single sign-on.

Agent memory placement can cut your token bill up to 2x
Agent memory has to refresh every turn, but putting it in the system prompt breaks prompt caching and re-bills the whole conversation each turn. Here is the one-message fix, and the token savings it led to in an experiment.

Markdown is not agent memory
Some of the most capable agents in production keep their memory in plain markdown files, and for a single agent and a single user it is hard to beat. The pattern breaks in predictable places: at scale, as facts change and errors compound, and under concurrent agents.

Building Agents in Go Without a Framework
A production agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human. That shape fits Go's runtime. This post explains why, surveys the Go framework options, and shows how to build an agent without one.

Sycophancy is a design choice
Writer's "Recalling Too Well" paper says memory systems amplify sycophancy. Its own data traces the amplification to two design decisions — one in Writer's experiment itself, one in a competitor's memory product.

The Batch API: Load Large Datasets into Agent Memory
Zep's Batch API loads large datasets into agent memory faster, in batches up to 50,000 items, with a progress dashboard and no impact on real-time ingestion.