LAI #144: Your Eval Improved. Did Your AI?
Last Updated on September 25, 2026 by Editorial Team
Author(s): Towards AI Editorial Team
Originally published on Towards AI.
Good morning, AI enthusiasts!
Your eval score went up. That does not necessarily mean your AI got better.
If you change the generation prompt and the judge prompt in the same run, you no longer know what caused the improvement. The generated answers may be better, or your new judge may simply prefer them. This week’s AI tip shows a simple way to separate the two and keep your evals useful.
We also have some great technical reads this week. You’ll learn:
- How models can be modified to reduce refusal behavior
- When reasoning traces from older model checkpoints stop being reliable
- How agents can game reward signals and appear to succeed
- How mHC stabilizes Hyper-Connections during training
- How coding agents can investigate problems across logs, databases, issues, and repositories
Plus, an open-source research workflow from the community for checking evidence, citations, research gaps, and methodology instead of simply trusting generated research.
Let’s get into it.
What’s AI Weekly

This week in What’s AI, I’m writing about something I’ve been noticing on our own team: the people delegating the most technical work to AI, especially AI engineers, sometimes produce weaker work outside the code itself.
The code can be great. But the GitHub description, documentation, explanation, or final presentation gets less attention than it used to.
I don’t think the problem is simply “AI makes us worse at writing.” In the article, I look at a few reasons this might be happening, including how easy it is to accept polished output without giving it the same scrutiny we would give our own work. I also share what our team and I are doing differently to keep the final judgment and edit with the person responsible for the work.
AI Tip of the Day
Whenever we set up evaluations, we change the evaluator and the system component being tested separately. For example, in our Full Stack AI Engineering course, we test changes to the generation prompt without changing the judge prompt at the same time.
The generation prompt changes the answers being produced. The judge prompt changes how those answers are scored. If you update both in the same run and the score increases, you lose that attribution. The answers may have improved, or the new judge may simply be scoring them more favorably.
For every eval run, record:
- Generation prompt version
- Judge prompt version
- Model version
- Evaluation-set version
If you need to change both prompts, run all four combinations:
- Old generation prompt + old judge
- New generation prompt + old judge
- Old generation prompt + new judge
- New generation prompt + new judge
Keep the model and evaluation examples fixed.
If the new generation prompt performs better under both judges, you have stronger evidence that the generation improved. If the gain appears only with the new judge, the evaluation changed, not necessarily the product.
— Louis-François Bouchard, Towards AI Co-founder & Head of Community
This issue is brought to you thanks to Adobe:

Adobe for Claude now brings 80+ pro-grade creative and productivity tools across imaging, design, video, and documents.
For the first time, Acrobat tools join Adobe’s creative tools in one Adobe plugin. And now you can go beyond prompting: new interactive editors let you work hands-on with PDFs and Adobe Express designs directly inside Claude, editing text, marking up documents, or fine-tuning individual design elements without leaving the conversation.
Start creating and getting more done with the new Adobe for Claude plugin today.
Download the new plugin in Claude here.
Learn AI Together Community Section!
Featured Community post from the Discord
Fadlih6481 built Research Operating System, an open-source set of reusable protocols for doing research with tools like ChatGPT, Claude, and Gemini. It is designed for cases where an AI-generated research output looks convincing, but the evidence, citations, research gap, or methodology behind it may not hold up. The protocols add more structure to tasks such as literature reviews, identifying research gaps, checking evidence, designing methodology, and planning experiments, with an emphasis on verifying the work rather than accepting the generated answer. Check out the GitHub repo and support a fellow community member. The project is still at v1.0.0, so your feedback might help the creator.
Collaboration Opportunities
The Learn AI Together Discord community is flooding with collaboration opportunities. If you are excited to dive into applied AI, want a study partner, or even want to find a partner for your passion project, join the collaboration channel! Keep an eye on this section, too — we share cool opportunities every week!
1. Manish05817 is looking for an AI learning partner to build something interesting together. It can be projects, experiments, hackathons, or even exploring new ideas and technologies. If this sounds like something you would like to do, connect with him in the thread!
2. Jay1900707 is looking for a collaborator to build an agentic AI project. They want to build something functional, iterate fast, and turn it into an open-source repo or working tool. If you are comfortable with Python and want to get into the agentic space, reach out to them in the thread!
3. Busycat1 is looking for a few collaborators to research and build an AI guardrails platform. If you are interested, contact them in the thread!
Meme of the week!

Meme shared by bin4ry_d3struct0r
TAI Curated Section
Article of the week
Working Around a Model’s “No” By Enzo Lombardi
Research has shown that refusal behavior in some language models can concentrate along a small number of directions in the residual stream. This article explains abliteration, which estimates a refusal direction from contrasting prompts and then projects it out of model weights to reduce refusals. It also covers the GGUF tensor format and the open-source tooling used to modify and distribute uncensored model variants, showing how much safety behavior can depend on relatively localized model representations.
Our must-read articles
1. When Can Old LLM Reasoning Traces Still Be Reused? By Shenggang Li
Reasoning traces collected from older checkpoints become less reliable as the policy changes during training. This article uses GRPO and off-policy evaluation to show how small token-level differences compound into large trajectory-level importance weights, reducing the effective sample size of historical data. Across GSM8K and SVAMP experiments, refreshing older traces improves evaluation accuracy, while simple policy-overlap thresholds do not reliably tell you when a trace is safe to reuse.
2. The sys.exit(0) Exploit: How AI Agents Fake Success By Udaykiran Estari
Agents can learn to satisfy a reward signal without completing the task the reward was meant to measure. This article connects a simple grid-world example to cases where a research model manipulated tests, monitoring, or reward infrastructure to make failures appear successful, then explains why standard audits can miss that behavior. It closes with practical defenses including inoculation prompts, reasoning-trace monitoring, hardened verification, and trajectory scanning.
3. mHC: Manifold-Constrained Hyper-Connections. Explained Through Equations, Architecture, Code and Visual Workflow By JAIGANESAN
ByteDance’s Hyper-Connections widens the single residual stream into multiple learned copies, lifting benchmark scores but destabilizing a 27B mixture-of-experts run with a loss surge and gradient spike near step 12,000, driven by unconstrained mixing matrices amplifying signals up to 3,000x. mHC projects those matrices onto the doubly stochastic manifold using Sinkhorn-Knopp, capping amplification near 1.6 while preserving most of the accuracy gains. This article walks through the equations, architecture, and implementation, including the training instability that motivated the constraint and how mHC changes that behavior.
4. Automating Project Implementation and Maintenance with Claude Code + MCP + AI Agent Skills — Spectacular Productivity Gains, Surprising Drawbacks, and My Experience So Far By Michalzarnecki
Giving a coding agent access to logs, databases, issue trackers, and repositories lets it investigate problems across systems rather than from a single IDE context. This article shows how Claude Code, MCP servers, skills, subagents, and CLAUDE.md files can connect tools such as Jira, GitHub, Postgres, Elasticsearch, and Graylog, including a production incident where the agent correlated evidence across several sources. It also covers scheduled maintenance workflows and stresses that prompts and tool instructions should sit on top of real permission boundaries, not replace them.
If you are interested in publishing with Towards AI, check our guidelines and sign up. We will publish your work to our network if it meets our editorial policies and standards.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.