The Evaluation Stack: Metrics That Predict Production Quality
Author(s): Armin Norouzi, Ph.D Originally published on Towards AI. The Evaluation Stack: Metrics That Predict Production Quality The hardest engineering problem in LLM deployment is not latency or cost — it is knowing whether your model got better. Metrics that seem rigorous …
Tool Call Orchestration: Sequential, Parallel, and DAG Execution
Author(s): Armin Norouzi, Ph.D Originally published on Towards AI. Tool Call Orchestration: Sequential, Parallel, and DAG Execution A research agent calling 6 tools sequentially waits 1,832ms at P50. The same 6 tools run in parallel — ignoring all dependencies — finish in …
Production RAG API with FastAPI, pgvector, and Claude
Author(s): Armin Norouzi, Ph.D Originally published on Towards AI. Production RAG API with FastAPI, pgvector, and Claude Most RAG tutorials show you how to call an embedding API and do a similarity search. That is the easy 20%. The hard 80% is …
Fault-Tolerant Agent Pipelines: Checkpoint, Retry, and Compensate
Author(s): Armin Norouzi, Ph.D Originally published on Towards AI. Fault-Tolerant Agent Pipelines: Checkpoint, Retry, and Compensate An autonomous agent that runs for 20 minutes without any fault-tolerance mechanism is a production incident waiting to happen. The agent calls an external API, the …
Counterfactual Evaluation in Ads: IPS, SNIPS, and Doubly Robust
Author(s): Armin Norouzi, Ph.D Originally published on Towards AI. Counterfactual Evaluation in Ads: IPS, SNIPS, and Doubly Robust You have a new ranking model. You want to know if it’s better than the one in production before you ship it. The honest …