Before UX was a function, engineers designed software. Not badly. Not carelessly. But the definition of “done” belonged to the people building the thing, and to them done meant the function executed without falling over.
Nobody in the building had a job that succeeded or failed on what the person at the other end actually experienced.
That is the part people get wrong when they retell it. UX did not arrive because engineers lacked empathy. It arrived because there was no accountability for the outcome, and eventually somebody costed that out and it justified a headcount line. The empathy language showed up afterward, as a way of explaining what the new function was for.
We are in the same moment again. I also think most people have the gap in the wrong place, including me until fairly recently.
The bet everyone is making, and why it expires
The common line right now is that AI cannot do judgment. It cannot infer, cannot handle abstraction, produces output without understanding. So the thinking is safe and the humans keep the interesting part.
I have been working with these systems every day for two years, as an operator rather than a commentator, and I do not believe that survives the next eighteen months.
Inference and abstraction are things current models do reasonably well. Ask for something vague and directional and you will get something defensible back. That capability is improving fast. Any career thesis built on it has an expiry date printed on it somewhere.
Betting on a capability gap is a bad trade. Capability gaps close. The question worth asking is which gaps are structural, meaning they are properties of the arrangement rather than of the model.
Four gaps that do not close
Working closely with agents on real product work, the failures cluster. None of them are really about intelligence.
It infers from priors, not from the situation. A model works from what is typical, not from what is true here. It cannot know that the third stakeholder hates that colour because of where they worked before, or that the integration everybody assumes is fine is held together by one person’s goodwill and a cron job. It will give you something reasonable for the average case. The average case does not exist.
No continuity of taste. Every session starts over and re-derives judgment from whatever context is in front of it. Human taste accumulates. It is the residue of shipping things that did not work and remembering how that felt. A system that reconstructs its preferences from scratch each time can have a style. That is not the same as taste.
Plausibility beats tension. This is the one that matters most and gets discussed least. When two good things conflict, speed against craft, breadth against depth, this customer against that segment, a model will reach for a synthesis that reads well and quietly dissolves the conflict. But design judgment is very often the refusal to dissolve it. Choosing. Then saying out loud what you gave up. Something optimized for a satisfying answer is structurally disinclined to leave you uncomfortable, and being left uncomfortable is frequently the point.
No stake. It does not live with the thing after it ships. It will not be in the room in six months when the call turns out to have been wrong. You cannot bolt accountability onto a model. It comes from having something to lose.
None of those four close with a better model. They are the same shape as the gap that produced UX, which is why a function forms around them.
It will not be the function people expect
Here is where I part ways with the version of this argument I hear most, which is that UX becomes the critical skill.
Classical UX is deterministic. You design a flow and the flow is what happens. The user takes the path you built or they fail, and if they fail you fix the path. Journey maps, wireframes, task analysis, usability testing against a defined task, all of it assumes an authored route through the product.
Agentic systems do not behave that way. They are probabilistic. You are not authoring the path, you are setting conditions, and the path arises differently every time. Sometimes in ways you did not anticipate. Occasionally in ways you would not have allowed.
What comes out of the design process is not a flow. It is a possibility space.
There is a discipline that has been doing that for forty years and it is not UX.
Game design is the craft that transfers
A level designer does not script the player’s experience. They place constraints, affordances, feedback and pacing, and the experience emerges out of the interaction between those and a person they will never meet. The designer is accountable for an experience they did not directly author.
That is the whole job, and it is exactly the problem agentic systems present.
- You design the possibility space, not the route. What can happen. What should be easy, what should be hard, what should be impossible.
- Difficulty curves. How much should the system do for someone at minute one versus month six? Every agentic product has a difficulty curve problem. Almost none of them know it.
- Feedback legibility. Games are obsessive about telling you why something happened. Agentic systems are famously terrible at it. Solved in one field, wide open in the other.
- Failure as content. Games assume you will fail and make failing informative and survivable. Most AI products treat failure as an exception to apologize for, which teaches people to stop trusting the system rather than to get better at driving it.
- Emergent misuse. Game designers expect players to break things creatively and design for it. Enterprise software files that under edge case.
I am not arguing UX practitioners cannot do this. I am arguing the methods do not transfer cleanly, and that a field which has spent decades accountable for emergent outcomes starts a long way ahead of one that has spent decades authoring paths.
I should admit I had this filed under old history. I produced games before I was a product executive and treated it as an interesting detour for about fifteen years. It stopped being a detour around eighteen months ago.
Which makes the name a problem
Call this new thing UX and it gets filed where UX currently sits in most organizations, with the people who make screens. Production work. Exactly the work being automated fastest.
That is an unpleasant irony and I do not have a fix for it. The discipline best equipped to own the judgment layer carries a name that now means the layer above it. I distrust invented job titles, so I am not going to invent one. If you are arguing for this internally, argue for the accountability instead of the title. Whoever owns whether the output was any good is the function, whatever it ends up being called.
The best argument against all of this
Worth putting up myself rather than waiting for somebody else to.
RLHF is literally encoding human taste at scale. If preference can be learned from enough examples, the judgment layer collapses into the model instead of forming into a role. That is not a weak objection. It is the mechanism by which every previous “machines will never do X” prediction has failed.
My honest read is that preference learning gets to median taste extremely well and to distinctive taste never, because distinctive taste is by definition off-distribution. It is the thing most people would not have picked. Train toward the centre of a preference distribution and you reliably get competent, unsurprising, faintly generic work, which is fine, and often enough, and is also the exact line separating products people put up with from products people love.
So the function survives. Smaller and more senior than UX became, though. Fewer people, further from production, higher up. That is a real downgrade on the optimiztic version of this argument and I would rather hold it than sell the bigger claim.
Timing, and where the opening actually is
UX took roughly fifteen years to go from recognized idea to standard function with budgets, titles and career ladders. Early nineties to the late 2000s, give or take. The capability showed up long before the org chart did.
The technology is moving faster this time. The organizational part is not, and cannot, because it runs at human speed. Somebody has to write a job description, defend a headcount, build a career path and convince a CFO. Expect the capability gap to close faster than the accountability gap.
That lag is the whole opportunity. It is not a race against the models. It is a window where the work is obviously necessary and no role exists to do it, which is precisely the window UX was invented in, mostly by people who gave themselves the title before anyone offered it.
What I would actually do
- Stop arguing AI cannot think. You will be wrong in public, on a schedule set by someone else’s release notes. Argue instead that it has no continuity, no situation, no stake, and no willingness to leave a tension unresolved.
- Instrument judgment. If nobody measures whether the output was any good, the role has nothing to stand on. Somebody needs a number that would look bad if the AI produced plausible rubbish quickly.
- Go read game design. Not for inspiration. For method. Difficulty curves, feedback legibility and designing for productive failure are directly applicable and almost entirely absent from enterprise product practice.
- Name the accountability before you name the job. “Who signs off that this was worth shipping” is a more useful question than what to call them.
The empathy gap was never really the point in 1998 either. What changed things was one person in the building being answerable for the experience. That is the gap now. It is the one worth standing in.
