Qwen-UI-Agent Promises Bash. The Repo You Can Download Ships 12 Actions and No Shell.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Alibaba’s new GUI agent posts scores on seven benchmarks on paper. I counted every action in the code they actually published: 12, every one a screen gesture or a bookkeeping …
MCP Just Dropped 13 of Its 31 Methods. Why Did the Schema Still Grow 48%?
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. MCP Just Dropped 13 of Its 31 Methods. Why Did the Schema Still Grow 48%? Every summary of the 2026-07-28 Model Context Protocol release says the same thing in the …
Gemma Refuses Your System Prompt. Mistral Moves It. Llama Rewrites It.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. I rendered the same chat — one system message, one user message, then again at three and five turns — through twelve production chat templates. Two turned out to be …
GPT-5, Llama And Qwen Agree: YAML Is Smaller Than JSON And Costs More Tokens
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Ten serialisations, seven production tokenizers. Across a hundred records YAML is 3% fewer bytes than minified JSON and 21% more tokens — and the format nobody suggests halves it again. …
Llama 3.2 Needs Eight Tokens For One Bengali Word. Gemma 3 Needs One.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Eight tokenizers, the same document, twenty-one languages. The seven current ones agree on English to within four percent — and disagree by up to 4.95x once you leave it. মানুষ …
NVIDIA's Switchyard Routes Claude Code on 113 Hardcoded Strings and Ignores Your Prompt
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Why this landed now I counted every string NVIDIA’s new agent router matches on. There are 113 of them. Exactly one is ever tested against your prompt rather than against …
Claude Code Runs the Real Ponytail. Cursor and 11 Others Settle for 2,593 Bytes.
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Code Runs the Real Ponytail. Cursor and 11 Others Settle for 2,593 Bytes. Ponytail’s own portability doc lists 22 coding agents. I parsed it and counted: only 9 of …
Y Combinator Ditched All But One Claude Code Tool for 16 of Its Own
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Y Combinator Ditched All But One Claude Code Tool for 16 of Its Own Y Combinator open-sourced the agent harness it runs its own company on. I read the adapter …
DeepSeek V4-Flash vs GLM-5.2: The 1.7-Point Win Collapses When You Swap the Harness
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. DeepSeek V4-Flash vs GLM-5.2: The 1.7-Point Win Collapses When You Swap the Harness DeepSeek’s own chart says 82.7 on Terminal-Bench 2.1. Artificial Analysis measured 79. That 3.7-point gap is 2.2x …
Temperature 0 vs 1.0: Greedy Decoding Collapsed 14% of Llama’s 128K Calls
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Temperature 0 vs 1.0: Greedy Decoding Collapsed 14% of Llama’s 128K Calls Temperature 0 is the most copy-pasted line in production LLM code. It is also, on long context, the …
Unsloth vs Axolotl vs TRL: 87% of Your Fine-Tuning VRAM Goes to a Tensor You Never Wrote
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Unsloth vs Axolotl vs TRL: 87% of Your Fine-Tuning VRAM Goes to a Tensor You Never Wrote I spent an afternoon pricing every byte of a LoRA fine-tuning step from …
Mem0 vs Zep vs Letta: A Folder of Text Files Shouldn’t Beat the 61K-Star Memory Layer
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Mem0 vs Zep vs Letta: A Folder of Text Files Shouldn't Beat the 61K-Star Memory Layer The most-installed agent memory layer on GitHub has passed 61K stars. Letta beat its …
MCP Just Deleted Its Own Handshake — and Every Request Got 118% Bigger
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. MCP Just Deleted Its Own Handshake — and Every Request Got 118% Bigger I measured the new Model Context Protocol wire format against the old one this morning, and the …
Claude Cracked an 87-Year-Old Conjecture — I Verified the Counterexample in 0.1 Seconds
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Cracked an 87-Year-Old Conjecture — I Verified the Counterexample in 0.1 Seconds On Sunday evening, while most of the world was watching the World Cup final, a mathematician at …
Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the Numbers
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the Numbers Yesterday everyone cheered Gemini 3.6 Flash for topping the computer-use leaderboards. Then …