Qwen3.8-Flash-Next on 4 GPUs: device_map="auto" Leaves GPU 0 Empty and Offloads 22 GB
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Qwen3.8-Flash-Next on 4 GPUs: device_map="auto" Leaves GPU 0 Empty and Offloads 22 GB If you load Qwen3.8-Flash-Next with transformers on four 80 GB GPUs, device_map="auto" leaves GPU 0 empty and …
Claude Code Caveman Eval: Two Sentences Took the Median Cut From 14.6% to 50.6%
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Code Caveman Eval: Two Sentences Took the Median Cut From 14.6% to 50.6% I recounted the repo’s own eval snapshots: the jump came from two added sentences, v3.1.0 still …
Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Anthropic’s SDK 0.128.0, released with Claude Opus 5.5, can hand the model a new tool mid-run without touching tools[]. It needs one beta flag the runner won’t add for you. …
Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Anthropic’s SDK 0.128.0, released with Claude Opus 5.5, can hand the model a new tool mid-run without touching tools[]. It needs one beta flag the runner won’t add for you. …
Why Does llama.cpp’s Own API Give Xiaomi’s MiMo V2.6 Flash 9x the KV Cache?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. At 128K tokens the model needs 2.96 GiB of cache. libllama’s default context settings would allocate 27.19 GiB, llama-cpp-python’s 28.31. Two flags fix it. If you load Xiaomi’s new MiMo …
Why Does llama.cpp’s Own API Give Xiaomi’s MiMo V2.6 Flash 9x the KV Cache?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. At 128K tokens the model needs 2.96 GiB of cache. libllama’s default context settings would allocate 27.19 GiB, llama-cpp-python’s 28.31. Two flags fix it. If you load Xiaomi’s new MiMo …
Google’s AX Takes a Port in Every Egress Rule. Nothing Downstream Can Enforce One.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Google's AX Takes a Port in Every Egress Rule. Nothing Downstream Can Enforce One. Google’s agent orchestrator sets a port on six egress rules in its own repo. I traced …
Google’s AX Takes a Port in Every Egress Rule. Nothing Downstream Can Enforce One.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Google's AX Takes a Port in Every Egress Rule. Nothing Downstream Can Enforce One. Google’s agent orchestrator sets a port on six egress rules in its own repo. I traced …
Claude Code Can Fine-Tune This 8.5 MB Model. Which of Its 19 Sizes Should You Ship?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Code Can Fine-Tune This 8.5 MB Model. Which of Its 19 Sizes Should You Ship? Needle 3’s depth ladder has four cliffs. One extra transformer block costs 0.58 MB …
Why Does LangChain’s New Jev Auto Mode Forget Your Request at Tool Call 16?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Why Does LangChain's New Jev Auto Mode Forget Your Request at Tool Call 16? After introducing Jev and the idea of using it as a guardrail classifier in LangChain’s …
Why Does LangChain’s New Jev Auto Mode Forget Your Request at Tool Call 16?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Why Does LangChain's New Jev Auto Mode Forget Your Request at Tool Call 16? The article explains that LangChain’s Jev-powered Auto Mode, as implemented in the current middleware, starts …
Vercel's New Coding Agent Takes Away Your MCP Tool List. It Sends the Same 15 Schemas.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. fx declares seventeen built-in tools, and a full turn with a subagent host carries fifteen function schemas. I measured a stock four-server MCP setup at 37 tools and 22,226 bytes …
Qwen Code Ditched Google 10 Months Ago. Why Do 1,110 Files Still Say “Copyright Google”?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. I checked every one of them against every blob Google’s Gemini CLI has ever committed. The header is still literally true for 58. I was reading Qwen Code’s source last …
Claude Code’s Rust Rival Calls Its Extension Gate 123/123. Its Own Evidence File Reads ‘fail’.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Code’s Rust Rival Calls Its Extension Gate 123/123. Its Own Evidence File Reads ‘fail’. There is a particular kind of document that shows up in ambitious open-source repositories, and …
Claude Code's Edit Format Beat omp's Default on 12 of My 15 Fumbled Edits
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. A coding agent that ships five different ways to change a file, and the one it turns on by default is the only one that put my code on the …