Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
I Failed a Memory-layer System Design Round. So I built one!
Artificial Intelligence   Latest   Machine Learning

I Failed a Memory-layer System Design Round. So I built one!

Last Updated on September 25, 2026 by Editorial Team

Author(s): Swaraj

Originally published on Towards AI.

How studying Mem0-style architecture led me to CogLayer — a hands-on lab for the read/write memory loop behind modern AI agents.

A few weeks ago I sat in a system-design conversation with a NYC-based startup.

They asked about to build a system design for a memory layer for their sales agent: not just a basic RAG, but how an agent should remember, update, and forget facts about a user and their conversation with the system over time — without blocking the chat reply, without stuffing the entire history into every prompt, and without treating every similar embedding as the same fact.

I knew the buzzwords. I did not have a crisp mental model.

Well I did build a complete RAG with strong redis cache and cache TTL and invalidation as well, unfortunately didn’t clear that round, so I decided to fill my knowledge gap by building one and currently one of the best memory layer architectures for conversational agents out there is owned by mem0. I read how Mem0-style pipelines are described in the wild — hybrid retrieval + LLM judgment — and implemented a small but complete version: CogLayer. Here’s the research paper by them which I referred— https://arxiv.org/pdf/2504.19413

This article is the write-up I wish I’d had before that interview: the why, the architecture, and a path you can click through yourself.

In chatbots, people often mix three different things:

1. Context window — whatever fits in the current prompt

2. Conversation history — raw turns (“User said… Assistant said…”)

3. Durable memory — atomic facts you want to keep across sessions

(“Alex prefers dark roast coffee,” “User adopted a dog named Scout”)

Mem0-style systems focus on (3), while still using (2) as fuel for extraction.

The hard part is not storing text. The hard part is:

  • Finding candidate related memories cheaply
  • Deciding whether a new fact is about the candidate new, a refinement, a contradiction, or already known.
  • Updating storage without making the user wait on that judgment.

That last point is why the architecture splits into two paths.

The core idea: hybrid recall + judgment

Vector search is excellent at cheap recall:

“Given this sentence, what are the top-s similar memories for this user?”

It is terrible at semantic judgment on its own.

“Likes X” and “hates X” can sit next to each other in embedding space. Similarity will happily retrieve both. Only a model (or careful rules) can say: update, delete, or noop.

So the hybrid loop is:

Embeddings / Qdrant — Narrow the world to a few relevant memories |

LLM EXTRACT — Pull durable candidate facts from a message pair + recent context

LLM DECIDE — Pick exactly one tool: `ADD` / `UPDATE` / `DELETE` / `NOOP`

Async worker — Apply the tool to the store so chat stays snappy

Mem0 popularized this style of thinking for production agents. CogLayer is a learning implementation of that loop — not a clone of Mem0 Cloud, and not a claim to replace it. The goal is understanding you can demo and explain.

Two paths: Read (sync) and Write (async)

Read path:

Every `/chat` request:

1. User sends a query

2. Embed the query → Qdrant top-k , filtered by `user_id`

3. Load recent conversation turns

4. Build a prompt (instructions + memories + history + query)

Subscribe to the Medium newsletter

5. LLM generates a reply

6. Return the reply immediately

7. Fire a background job (Celery / RabbitMQ) — do not block the HTTP response.

Write path:

The worker receives the message pair (user query + assistant reply):

1. Save the pair to the conversation store

2. Build extraction prompt P: rolling summary S + last ~10 turns + new pair

3. LLM #1 — EXTRACT→ candidate facts

4. For each candidate: embed → top-s similar memories for that user

5. LLM #2 — DECIDE → one tool call

6. Apply ops to JSON memory store/or you own storage and Qdrant(re-embed on ADD/UPDATE)

Summary S in CogLayer is computed when the write path runs(on each new pair). Some production designs cache summary with a periodic summarizer to save LLM cost at scale. For learning and for short demos, on-write summarization is the honest default.

Why system design interviews care about this?

If you only say “we’ll use a vector database/RAG,” you miss:

  • Isolation — every search scoped by user (and often agent/session)
  • Non-blocking writes— queue the maintenance job
  • Idempotent identity — stable memory IDs, not random duplicates every turn
  • Contradiction handling — DELETE / UPDATE, not infinite ADDs
  • Dual stores— searchable vectors and a source of truth for text/metadata

Mem0-style design is a clean story for all of the above. Building even a thin version forces you to confront the failure modes: stale vectors after JSON updates, eager .delay() with no worker, EXTRACT inventing facts, DECIDE ADDing near-duplicates.

Those scars are better interview fuel than a perfect diagram.

Learning it by seeing it:

I Failed a Memory-layer System Design Round. So I built one!
CogLayer Architecture

CogLayer includes an animated lab (Excalidraw-ish whiteboard UI) that steps through Read and Write:

  • Play / Pause / Prev / Next
  • Nodes light up; arrows stay visible and highlight on the active hop
  • Each step has a short “what” and “why”.

Repo — https://github.com/Swaraj07082/CogLayer

I didn’t build CogLayer to compete with Mem0.

I built it because I couldn’t explain a memory layer under pressure — and explanation without implementation is fragile.

If you’re learning agent memory, preparing for AI system design, or teaching the difference between RAG-for-documents and memory-for-users, walk the /learn on my repo once. Then break it on purpose: kill the worker, skip Qdrant sync, force a contradictory fact. Watch what fails.

That’s how the diagram becomes intuition.

Further reading:

Peace Out! ✌️

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.