RAG from Scratch [Part 3]: Chunking — The Decision That Makes or Breaks Your Retrieval
Author(s): Sumit Vedpathak Originally published on Towards AI. TL;DR Imagine you’re studying for an exam. You have a 400-page textbook. This post explains why chunking is essential for RAG—because embedding models and LLMs have size limits and retrieval quality depends on how …
RAG from Scratch [Part 2]: Loading — The Step Everyone Skips and Everyone Regrets
Author(s): Sumit Vedpathak Originally published on Towards AI. RAG from Scratch [Part 2]: Loading — The Step Everyone Skips and Everyone Regrets Series 2 of 5 The article argues that most RAG failures begin at the ingestion/loading stage rather than later steps …
The Silent Speedup: How KV Cache Makes AI Feel Instant
Author(s): Sumit Vedpathak Originally published on Towards AI. The Silent Speedup: How KV Cache Makes AI Feel Instant Think of a chef who writes down every recipe the first time they make a dish — so the next time someone orders it, …