The Taste Gap: Why AI work that lands is rare
Towards AI The Taste Gap
1 The question

AI is everywhere. AI work that lands is rare. Why?

The gap between AI slop and AI work has become a problem of execution, not capability. AI can be used to increase the quality of your work, but for many it is reducing it. To produce a great output, and not just a passable one, AI needs your help.

The teams pulling ahead are putting expert judgment back into the workflow at every step. Some of that work requires structured AI training. Some happens in building custom AI systems.

2 The map

Two questions, four kinds of operator.

Two questions split the field. Does the operator have taste at the task? Do they bring that taste into the AI workflow?

The upper-right box is where the leapfrog is most likely. The expert keeps judgment in the loop, rejects the first draft, and turns the model into a tool for range, speed, and finish.

Most AI debate confuses these boxes, and most companies do too. Many rollouts accidentally move experts into the trap: the model gets the task, but the expert's judgment never reaches the work.

HAS TASTE AT TASK   →

Has taste, skips it

The expert who outsources judgment to the model. Output drifts back toward average.

THE TRAP

Has taste, uses it

Direction at every step. Work gains range, speed, and finish without losing judgment.

THE WIN

No taste, no AI

Below a useful baseline and stays there. The pre-AI default for non-experts at this task.

UNCHANGED

No taste, uses AI

Lifted to average. A real and useful gain. This is most users today.

THE LIFT
INJECTS TASTE INTO AI WORKFLOW   →
Fig 02 / The operating model behind good and bad AI work.
3 The baseline

Generic prompts tend toward generic output.

The model was trained on a large body of text, pictures, and code. Without direction, it tends to return the patterns it has seen most often.

It scores many possible next words and turns those scores into probabilities. In unguided use, familiar openings, safe structures, and obvious punchlines often sit near the top.

That default behaviour matters. Interesting work usually begins when a person changes the path with taste, constraints, examples, rejected drafts, and sharper standards.

Run the prompt cold and you get the median answer, polished into clean sentences.

LOW SKILL HIGH SKILL FREQUENCY AI BASELINE = MEAN POPULATION AVERAGE
Fig 03 / The bell curve shows how skill is distributed. The AI baseline is the vertical line at the mean.
4 The lift

For most people, average is a promotion.

At many specific tasks, people sit below a useful baseline. A support agent writes a clearer refund note. A founder turns messy notes into a tolerable investor update. A junior analyst gets the first SQL query working instead of staring at a blank editor.

This is real value. The model can supply structure, grammar, syntax, and a first pass at the thing the person could not yet make alone.

This is the easiest part of the AI story to defend, and the part most often dismissed. Baseline competence matters, even when the result is useful rather than distinctive.

LOW SKILL EXPERT NATIVE SKILL AT TASK EXCELLENT POOR OUTPUT QUALITY AI BASELINE EVERYONE BELOW GETS LIFTED
Fig 04 / Each dot is one person. Below-baseline people rise vertically to the average.
5 The flood

AI slop is repetition at scale.

A writing quirk can be charming when one person owns it. A familiar cadence, tidy transition, stock joke, or neat three-part structure may pass unnoticed in a single piece.

The irritation starts when a million people ship the same move. The phrase stops reading like style and starts reading like residue. When sameness is everywhere, average output becomes brand drag, review burden, and operational noise.

ONE HUMAN QUIRK MASS REUSE ONE WRITER a signature COMMON PATH 1,000,000 OUTPUTS OPENER RHYTHM CLOSE REPETITION BECOMES NOISE
Fig 05 / One repeated move can be harmless. At scale, the pattern becomes the problem.
6 The trap

The expert who hands the keys over.

An expert has real taste built over years. They open the model, type a generic prompt, and ship whatever comes back. The output may look polished, but the expert is no longer present in the work.

Their edge came from judgment. The prompt removed it. What used to be their value is now a setting they did not change.

Skill that does not enter the workflow does not enter the work.

LOW SKILL EXPERT NATIVE SKILL AT TASK OUTPUT QUALITY AI BASELINE EXPERT LOST EDGE Collapses to average
Fig 06 / The expert who skips the taste step ends up exactly where everyone else ends up.
7 The leapfrog

The expert who directs the loop.

A smaller group does the harder thing. They bring decades of career experience: real customer conversations, hard-won edge cases, and industry memory that was never cleanly captured in the model's training set. Then they reject the first draft, ask for alternatives, compare them, name what is wrong, and drive it into a second shape.

AI has jagged intelligence: superhuman in some narrow moves, weak in places a good operator spots immediately. The leapfrog happens when the model supplies range and speed while the human supplies taste, context, touch, and judgment. The reader cannot always name why, but the work feels chosen rather than produced.

LOW SKILL EXPERT NATIVE SKILL AT TASK OUTPUT QUALITY AI BASELINE START HOLDOUT THE LEAPFROG Expert + AI + Taste
Fig 07 / The arrow visibly arcs over the holdout. Same starting skill, very different finish.
8 The vibe gap

Vibe coding makes everyone look like a builder.

A non-engineer can vibe-code a working app in an afternoon. A non-designer can produce a credible landing page in an hour. A non-writer can put a long report on the page by lunch. The speed gains are real and unevenly distributed: people who used to be locked out of building the artifact at all are now shipping demos. For first-draft work, AI compressed the visible part of the job by something like an order of magnitude, and that compression is not going away.

The artifact is not the deliverable. A demoable app still has to choose its state model, handle auth and rate limits, survive a flaky network, deal with the empty state and the error state, log enough to debug an outage, and make the small interface decisions that decide whether anyone returns next week. A polished report still needs the source check, the contradicting study the model did not surface, the structural cut, the falsifiable claim, and the line that earns the close. The discount on that work is much smaller than the discount on the first draft.

That second layer still belongs to expert hours. The total project is faster than it used to be, but the savings curve is steepest at the start and flattens hard once the surface looks done. A vibe-coded prototype that took three hours can take another forty before it scales as a product. A vibe-written report that took ninety minutes can take another two days before it survives review. The ratio inverts: most of the remaining time sits in the part nobody outside the team can see.

That is where the manager trap lives. The person who built the prototype reads it as ninety percent done. The engineer or editor who can name what is missing reads it as forty, and the schedule is being set by the higher number. The same pattern shows up in design comps that fall apart on a single user test, in research output that breaks on the first peer challenge, and in marketing copy that does not survive a legal read.

Vibe coding makes everyone look like a builder. The expert is the one who knows what is still missing.

9 The wrong question

The audience asks the wrong question.

The flood of average AI work has bred suspicion of all AI work. Readers, viewers, hiring managers, editors. The reflex is to ask whether a tool was used at all.

The reflex has a real cause. Most published AI work today comes from free, weaker models. A smaller share comes from strong models used without judgment in the loop. Either way, the median output is generic, and the suspicion is earned. The trap is letting that suspicion become a tool test rather than a quality test.

Agentic coding has already broken that reflex. No serious person sneers anymore because a developer did not type every line by hand. The useful signal is whether the person understood the system, made the right tradeoffs, and owned the result.

The same standard should apply outside code. Ask whether human judgment shaped the work.

"Did you use AI?"
Was there human judgment in the loop?
Fig 08 / Provenance is the weaker question. Judgment is the stronger one.
10 AI detectors

AI detectors solve the wrong puzzle.

Set aside the long history of false positives. Assume AI detectors worked perfectly. They would tell you whether tokens look machine-generated. They would not tell you whether a person with judgment shaped the result.

The dimension everyone cares about, quality and direction, is the one thing the tool cannot see.

Editors, teachers, and hiring managers using detectors are sorting on the wrong axis.

DETECTOR VERDICT
AI probability92.4%
Tokens flagged147 / 612
ConclusionLIKELY AI
Not measured: Human judgment, taste, structure, originality, point of view, accuracy, purpose.
Fig 09 / Confidence on the wrong axis.
11 Practical hygiene

Even so, scrub the obvious tells.

You can put real insight and effort into an AI collaboration and still lose the reader at the surface. A few familiar AI signatures can make careful work look careless.

Even the strong models leak the same signatures. The tics are model-wide, not a sign of weak models. Careful work that does not get scrubbed reads like the median and gets judged that way.

If you want the work judged on its merits, strip the signals that mean nothing about the work and everything about the source.

Surface signals that weaken trust
  • Em dashes used as a tic
  • Forced contrast cadence
  • Over-polished stock vocabulary
  • Three bullet conclusions to every section
  • Unnecessary symmetric tricolons
  • Hedge phrases in every paragraph
  • Generic "in conclusion" wrap-ups
Fig 10 / Surface fixes that buy you a fair reading.
12 The bottleneck unlock

World-class hidden human talent can be revealed by AI.

A great output requires a package of skills; many people have top-tier talent in one area that never reaches an audience because their surrounding skills are subpar.

The weakest skill sets the ceiling. A person can be world-class at one thing and remain blocked because the surrounding skills sit below the publishable line.

AI can lift the supporting skills above that line, and the talent finally reaches an audience.

Rare talent Clears bar Below bar Lifted by AI
Songwriter · generational hooks
Bottleneck: vocals + production
Writes melodies people will hum for years. Cannot sing them. Has no idea how to mix.
Without AIDoes not ship
Melody
Lyrics
Vocals
Mix
Two failing skills keep the songs in a notebook.
With AIShips full package
Melody
Lyrics
Vocals
Mix
AI vocals and mix. The hooks finally reach listeners.
Operations lead · rare process judgment
Bottleneck: tooling + reporting
Knows where work gets stuck. Cannot turn that knowledge into a reliable internal tool.
Without AIManual process
Process
Data
Tooling
Reports
Expert time leaks into cleanup, reminders, and status chasing.
With AIWorkflow ships
Process
Data
Tooling
Reports
A custom workflow fills the gaps. Process judgment becomes operational leverage.
Fig 11 / The dashed line is the publishable bar. The weakest leg decides if the package ships.
13 The maturity curve

Companies move through four stages of AI adoption.

Stage zero is familiar work without AI. It is slower, but the accountability chain is clear. People know who made the call and where judgment entered the work.

Stage one is shadow AI. People grab free tools, paste work into generic interfaces, and invent private workflows. Quality can dip below the pre-AI line when experts delegate judgment instead of applying it. Risk rises too, from data leakage to brand damage to compliance gaps.

Stage two is managed AI. Sanctioned tools, paid licences, policies, and training pull quality back up. The company is using AI on purpose, but much of the work still happens inside generic interfaces.

Stage three is embedded AI. Custom workflows are built into core processes, with expert review points, domain rules, retrieval, evals, and audit trails, and the leapfrog from earlier sections is scaled across an organisation.

The goal is higher-quality output from humans and AI together. Speed comes for free at every stage. Quality only rises when human judgment lives inside the system rather than arriving as cleanup.

Stage 0 Stage 1 Stage 2 Stage 3 No AI Shadow AI Managed AI Embedded AI Productivity Output Quality Org Risk RISK PEAKS QUALITY DIPS Traditional methods. No automation. Free tools used unsanctioned. Official policies, training, platform. Custom workflows built into core ops.
Fig 12 / Shadow AI is where risk rises before quality systems catch up.
14 From general to custom

The next jump is from general tools to custom applications.

Enterprise adoption is real, but many companies still use the same generic interface for specialized work. That leaves expert judgment outside the system, where it arrives late as correction, cleanup, or rejection.

A generic chatbot can answer a question. A vibe-coded app can demo an idea. Neither automatically carries an engineer's taste for states, failure modes, latency, permissions, data shape, review loops, or the small interface decisions that make a solution feel inevitable.

The larger value sits in applications built around the work itself: role-specific interfaces, company knowledge, domain rules, human review points, evals, audit trails, and integrations into existing systems. What an engineering team adds that a generic chatbox or quick demo cannot:

  • Design and product taste that makes the tool genuinely usable, beyond merely functional.
  • Workflow design that bakes in the end user's expertise exactly where the model needs it.
  • Evals built from real domain work, rather than generic benchmarks.

The payoff is discipline. The leapfrog becomes repeatable where the organisation can define quality, measure it, and keep expert judgment in the path of the work.

What custom adds
01
Company context enters first Knowledge, examples, policies, customer language, and edge cases are part of the workflow before generation starts.
02
Expert review is designed in The system asks for judgment at the points where judgment changes the answer, before damage is done.
03
Quality becomes measurable Evaluation sets, audit trails, and acceptance criteria turn taste from a private preference into an operating standard.
04
The workflow ships AI moves from a generic chat box to a domain-shaped application that can be improved, governed, and reused.
Fig 13 / Custom applications put expert judgment, company context, and measurement inside the workflow.
15 The hope

AI as an amplifier of taste.

Models will keep improving, but judgment stays human. People with talent who learn to direct the model will make work they could not have made alone.

The future worth building is one where AI raises the ceiling for the people who care about the work.

The Taste Gap