Has taste, skips it
The expert who outsources judgment to the model. Output drifts back toward average.
The gap between AI slop and AI work has become a problem of execution, not capability. AI can be used to increase the quality of your work, but for many it is reducing it. To produce a great output, and not just a passable one, AI needs your help.
The teams pulling ahead are putting expert judgment back into the workflow at every step. Some of that work requires structured AI training. Some happens in building custom AI systems.
Two questions split the field. Does the operator have taste at the task? Do they bring that taste into the AI workflow?
The upper-right box is where the leapfrog is most likely. The expert keeps judgment in the loop, rejects the first draft, and turns the model into a tool for range, speed, and finish.
Most AI debate confuses these boxes, and most companies do too. Many rollouts accidentally move experts into the trap: the model gets the task, but the expert's judgment never reaches the work.
The expert who outsources judgment to the model. Output drifts back toward average.
Direction at every step. Work gains range, speed, and finish without losing judgment.
Below a useful baseline and stays there. The pre-AI default for non-experts at this task.
Lifted to average. A real and useful gain. This is most users today.
The model was trained on a large body of text, pictures, and code. Without direction, it tends to return the patterns it has seen most often.
It scores many possible next words and turns those scores into probabilities. In unguided use, familiar openings, safe structures, and obvious punchlines often sit near the top.
That default behaviour matters. Interesting work usually begins when a person changes the path with taste, constraints, examples, rejected drafts, and sharper standards.
Run the prompt cold and you get the median answer, polished into clean sentences.
At many specific tasks, people sit below a useful baseline. A support agent writes a clearer refund note. A founder turns messy notes into a tolerable investor update. A junior analyst gets the first SQL query working instead of staring at a blank editor.
This is real value. The model can supply structure, grammar, syntax, and a first pass at the thing the person could not yet make alone.
This is the easiest part of the AI story to defend, and the part most often dismissed. Baseline competence matters, even when the result is useful rather than distinctive.
A writing quirk can be charming when one person owns it. A familiar cadence, tidy transition, stock joke, or neat three-part structure may pass unnoticed in a single piece.
The irritation starts when a million people ship the same move. The phrase stops reading like style and starts reading like residue. When sameness is everywhere, average output becomes brand drag, review burden, and operational noise.
An expert has real taste built over years. They open the model, type a generic prompt, and ship whatever comes back. The output may look polished, but the expert is no longer present in the work.
Their edge came from judgment. The prompt removed it. What used to be their value is now a setting they did not change.
Skill that does not enter the workflow does not enter the work.
A smaller group does the harder thing. They bring decades of career experience: real customer conversations, hard-won edge cases, and industry memory that was never cleanly captured in the model's training set. Then they reject the first draft, ask for alternatives, compare them, name what is wrong, and drive it into a second shape.
AI has jagged intelligence: superhuman in some narrow moves, weak in places a good operator spots immediately. The leapfrog happens when the model supplies range and speed while the human supplies taste, context, touch, and judgment. The reader cannot always name why, but the work feels chosen rather than produced.
A non-engineer can vibe-code a working app in an afternoon. A non-designer can produce a credible landing page in an hour. A non-writer can put a long report on the page by lunch. The speed gains are real and unevenly distributed: people who used to be locked out of building the artifact at all are now shipping demos. For first-draft work, AI compressed the visible part of the job by something like an order of magnitude, and that compression is not going away.
The artifact is not the deliverable. A demoable app still has to choose its state model, handle auth and rate limits, survive a flaky network, deal with the empty state and the error state, log enough to debug an outage, and make the small interface decisions that decide whether anyone returns next week. A polished report still needs the source check, the contradicting study the model did not surface, the structural cut, the falsifiable claim, and the line that earns the close. The discount on that work is much smaller than the discount on the first draft.
That second layer still belongs to expert hours. The total project is faster than it used to be, but the savings curve is steepest at the start and flattens hard once the surface looks done. A vibe-coded prototype that took three hours can take another forty before it scales as a product. A vibe-written report that took ninety minutes can take another two days before it survives review. The ratio inverts: most of the remaining time sits in the part nobody outside the team can see.
That is where the manager trap lives. The person who built the prototype reads it as ninety percent done. The engineer or editor who can name what is missing reads it as forty, and the schedule is being set by the higher number. The same pattern shows up in design comps that fall apart on a single user test, in research output that breaks on the first peer challenge, and in marketing copy that does not survive a legal read.
Vibe coding makes everyone look like a builder. The expert is the one who knows what is still missing.
The flood of average AI work has bred suspicion of all AI work. Readers, viewers, hiring managers, editors. The reflex is to ask whether a tool was used at all.
The reflex has a real cause. Most published AI work today comes from free, weaker models. A smaller share comes from strong models used without judgment in the loop. Either way, the median output is generic, and the suspicion is earned. The trap is letting that suspicion become a tool test rather than a quality test.
Agentic coding has already broken that reflex. No serious person sneers anymore because a developer did not type every line by hand. The useful signal is whether the person understood the system, made the right tradeoffs, and owned the result.
The same standard should apply outside code. Ask whether human judgment shaped the work.
Set aside the long history of false positives. Assume AI detectors worked perfectly. They would tell you whether tokens look machine-generated. They would not tell you whether a person with judgment shaped the result.
The dimension everyone cares about, quality and direction, is the one thing the tool cannot see.
Editors, teachers, and hiring managers using detectors are sorting on the wrong axis.
You can put real insight and effort into an AI collaboration and still lose the reader at the surface. A few familiar AI signatures can make careful work look careless.
Even the strong models leak the same signatures. The tics are model-wide, not a sign of weak models. Careful work that does not get scrubbed reads like the median and gets judged that way.
If you want the work judged on its merits, strip the signals that mean nothing about the work and everything about the source.
A great output requires a package of skills; many people have top-tier talent in one area that never reaches an audience because their surrounding skills are subpar.
The weakest skill sets the ceiling. A person can be world-class at one thing and remain blocked because the surrounding skills sit below the publishable line.
AI can lift the supporting skills above that line, and the talent finally reaches an audience.
Stage zero is familiar work without AI. It is slower, but the accountability chain is clear. People know who made the call and where judgment entered the work.
Stage one is shadow AI. People grab free tools, paste work into generic interfaces, and invent private workflows. Quality can dip below the pre-AI line when experts delegate judgment instead of applying it. Risk rises too, from data leakage to brand damage to compliance gaps.
Stage two is managed AI. Sanctioned tools, paid licences, policies, and training pull quality back up. The company is using AI on purpose, but much of the work still happens inside generic interfaces.
Stage three is embedded AI. Custom workflows are built into core processes, with expert review points, domain rules, retrieval, evals, and audit trails, and the leapfrog from earlier sections is scaled across an organisation.
The goal is higher-quality output from humans and AI together. Speed comes for free at every stage. Quality only rises when human judgment lives inside the system rather than arriving as cleanup.
Enterprise adoption is real, but many companies still use the same generic interface for specialized work. That leaves expert judgment outside the system, where it arrives late as correction, cleanup, or rejection.
A generic chatbot can answer a question. A vibe-coded app can demo an idea. Neither automatically carries an engineer's taste for states, failure modes, latency, permissions, data shape, review loops, or the small interface decisions that make a solution feel inevitable.
The larger value sits in applications built around the work itself: role-specific interfaces, company knowledge, domain rules, human review points, evals, audit trails, and integrations into existing systems. What an engineering team adds that a generic chatbox or quick demo cannot:
The payoff is discipline. The leapfrog becomes repeatable where the organisation can define quality, measure it, and keep expert judgment in the path of the work.
Models will keep improving, but judgment stays human. People with talent who learn to direct the model will make work they could not have made alone.
The future worth building is one where AI raises the ceiling for the people who care about the work.