LLM Observability Tools Compared: MLflow vs. Langfuse vs. Confident AI
Author(s): allglenn Originally published on Towards AI. The 2 a.m. page that tracing can’t explain A support bot answers a billing question with total confidence and gets the refund policy wrong. Nobody notices for three days because the response looked fine: grammatically …
Before Kimi K3 Goes Open: 8 Secrets Every Developer Needs to Know
Author(s): allglenn Originally published on Towards AI. 1. What You’re Actually Deploying: 2.8T MoE With KDA and Two Attention Code Paths This guide covers eight concrete things you need to understand before you touch the download button: architecture, VRAM math, tooling dependencies, …
SGLang: the AI Tool You’ve Never Heard of Just killed Hugging Face’s Own Inference Engine
Author(s): allglenn Originally published on Towards AI. SGLang: the AI Tool You’ve Never Heard of Just killed Hugging Face’s Own Inference Engine SGLang started as a UC Berkeley research paper in 2023. By 2026 it was running on 400,000+ GPUs, generating trillions …
The Inkling Model: What OpenAI’s Former CTO Has Been Cooking
Author(s): allglenn Originally published on Towards AI. Mira Murati’s Thinking Machines Lab just shipped its first model, and led with the line “this is not the strongest model available.” Here’s what Inkling actually is, and why that admission is the interesting part. …
I Replaced My $400/Month Claude API Bill With GLM-5.2 + vLLM Here’s the Exact Playbook
Author(s): allglenn Originally published on Towards AI. GLM-5.2 + vLLM Here’s the Exact Playbook Last month, three separate engineers on my team opened Slack at 6 AM asking the same question: why did our AI spend spike from $290 to $440 in …
How to Use Claude Code with Kimi K3 (and Switch Models by Typing One Word)
Author(s): allglenn Originally published on Towards AI. How to Use Claude Code with Kimi K3 (and Switch Models by Typing One Word) A verified, step-by-step guide to running Claude Code against Moonshot AI’s Kimi K3 using Anthropic-compatible environment variables, plus how to …
7 RAG & Agent System Design Questions You Will Face in Every AI Engineer Interview (With Answers)
Author(s): allglenn Originally published on Towards AI. 7 RAG & Agent System Design Questions You Will Face in Every AI Engineer Interview (With Answers) I watched a friend walk into a senior AI engineer loop last month with a portfolio full of …
7 RAG & Agent System Design Questions You Will Face in Every AI Engineer Interview (With Answers)
Author(s): allglenn Originally published on Towards AI. 7 RAG & Agent System Design Questions You Will Face in Every AI Engineer Interview (With Answers) I watched a friend walk into a senior AI engineer loop last month with a portfolio full of …
Kimi K2.7 Code vs. GLM-5.2: which open-weight coding model to self-host on vLLM
Author(s): allglenn Originally published on Towards AI. Core concepts: what makes these models tick You’ve just finished reading the sixth “open-source model beats GPT-5.5” post this month, and you’re still no closer to an infrastructure decision. Your team needs a coding agent …
Stop Prompting Claude Code, Start Engineering Loops: Master Agentic Automation
Author(s): allglenn Originally published on Towards AI. A prompt is a request. A loop is a system In early June 2026, a tweet from Peter Steinberger, the developer behind the OpenClaw framework, hit five million views in under a day. The gist …
How to Avoid Claude Code Usage Limits: Planning, Memory, Models, and Tools
Author(s): allglenn Originally published on Towards AI. Core concepts: what “usage limits” actually measure You’re forty minutes into a debugging session. Claude just found the bug, you’re about to ask for the fix, and instead you get a message telling you to …
LookML: An Alternative Semantic Layer Approach to build a Reliable AI Analytics Agent with BigQuery
Author(s): allglenn Originally published on Towards AI. Before we talk about where to store your registry, let’s address the elephant in the room: What about LookML? If you’re already using Looker, you might be wondering whether you need to build this YAML-based …