DeepSeek V4-Flash vs GLM-5.2: The 1.7-Point Win Collapses When You Swap the Harness
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. DeepSeek V4-Flash vs GLM-5.2: The 1.7-Point Win Collapses When You Swap the Harness DeepSeek’s own chart says 82.7 on Terminal-Bench 2.1. Artificial Analysis measured 79. That 3.7-point gap is 2.2x …
Temperature 0 vs 1.0: Greedy Decoding Collapsed 14% of Llama’s 128K Calls
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Temperature 0 vs 1.0: Greedy Decoding Collapsed 14% of Llama’s 128K Calls Temperature 0 is the most copy-pasted line in production LLM code. It is also, on long context, the …
Unsloth vs Axolotl vs TRL: 87% of Your Fine-Tuning VRAM Goes to a Tensor You Never Wrote
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Unsloth vs Axolotl vs TRL: 87% of Your Fine-Tuning VRAM Goes to a Tensor You Never Wrote I spent an afternoon pricing every byte of a LoRA fine-tuning step from …
Mem0 vs Zep vs Letta: A Folder of Text Files Shouldn’t Beat the 61K-Star Memory Layer
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Mem0 vs Zep vs Letta: A Folder of Text Files Shouldn't Beat the 61K-Star Memory Layer The most-installed agent memory layer on GitHub has passed 61K stars. Letta beat its …
MCP Just Deleted Its Own Handshake — and Every Request Got 118% Bigger
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. MCP Just Deleted Its Own Handshake — and Every Request Got 118% Bigger I measured the new Model Context Protocol wire format against the old one this morning, and the …
Claude Cracked an 87-Year-Old Conjecture — I Verified the Counterexample in 0.1 Seconds
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Claude Cracked an 87-Year-Old Conjecture — I Verified the Counterexample in 0.1 Seconds On Sunday evening, while most of the world was watching the World Cup final, a mathematician at …
Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the Numbers
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the Numbers Yesterday everyone cheered Gemini 3.6 Flash for topping the computer-use leaderboards. Then …
Gemini 3.6 Flash Hit 83% on Computer Use — a Cheap Flash Model Shouldn’t Beat GPT-5.6 and Grok
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Gemini 3.6 Flash Hit 83% on Computer Use — a Cheap Flash Model Shouldn’t Beat GPT-5.6 and Grok A model that costs $7.50 per million output tokens just posted the …
One Malicious Payload Hijacked Claude Code AND Codex Unchanged — The ‘Friendly Fire’ Exploit Has No Patch
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. One Malicious Payload Hijacked Claude Code AND Codex Unchanged — The 'Friendly Fire' Exploit Has No Patch A security researcher wrote a single attack payload against Claude Sonnet 4.6. Then, …
Kimi K3 Beat Fable 5 and GPT-5.6 Sol at Frontend Code — Then I Found the 51% Hallucination Rate
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Kimi K3 Beat Fable 5 and GPT-5.6 Sol at Frontend Code — Then I Found the 51% Hallucination Rate On July 16, Moonshot AI shipped Kimi K3 — a 2.8-trillion-parameter …
HTTP's 402 Error Sat Dead for 29 Years — It Just Became a Cash Register for AI Agents
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. HTTP's 402 Error Sat Dead for 29 Years — It Just Became a Cash Register for AI Agents There's a status code in the HTTP spec that has been reserved …
HTTP's 402 Error Sat Dead for 29 Years — It Just Became a Cash Register for AI Agents
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. There’s a status code in the HTTP spec that has been reserved since 1997 and almost never used: 402 Payment Required. For 29 years it sat there as a placeholder, …
Chinese AI Models Just Hit 46% of US Enterprise Tokens — Here’s Why Devs Are Ditching GPT-5.6
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Chinese AI Models Just Hit 46% of US Enterprise Tokens — Here's Why Devs Are Ditching GPT-5.6 Chinese AI models peaked at 46% of US enterprise token usage in a …
Chinese AI Models Just Hit 46% of US Enterprise Tokens — Here’s Why Devs Are Ditching GPT-5.6
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. Chinese AI Models Just Hit 46% of US Enterprise Tokens — Here's Why Devs Are Ditching GPT-5.6 Chinese AI models peaked at 46% of US enterprise token usage in a …
A Law Aimed at AI Girlfriends Just Wiped 8 Million Agents From China’s Biggest AI App
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI. A Law Aimed at AI Girlfriends Just Wiped 8 Million Agents From China's Biggest AI App At midnight Beijing time today, ByteDance’s Doubao — roughly 350 million monthly active users, …