Artificial intelligence is becoming extraordinarily good at doing things.
It writes, codes, searches, calculates, generates images, operates computers and runs experiments. In chess, machines have long surpassed the best humans. More recently, AI agents have begun to perform substantial amounts of research engineering without continuous human intervention. The trajectory is unmistakable: machines are acquiring more capabilities, more autonomy and more speed.
Yet something strange keeps appearing beneath the impressive demonstrations.
The problem is not always whether AI can do something. Increasingly, the harder question is whether it understands what it is actually being asked to do.
That is a problem of judgment.
And judgment is a much stranger property than intelligence benchmarks make it appear.
The stone crusher
Imagine an extraordinarily powerful machine sitting beside a mountain of stone.
Feed it thousands of irregular pieces of rock. It can crush them, sort them, recombine them and learn the patterns in their shapes. Then ask it to produce a beautiful statue.
It does.
Perhaps it produces a magnificent one.
Now give it a single, carefully selected block and provide a much more demanding instruction: Keep this exact contour. Remove a little stone here. Put a shallow curve there. Add three tiny wrinkles. Preserve that imperfection. Do not improve anything else.
The task has changed.
The machine may produce a more beautiful statue than the one you had in mind. It may even produce one an expert would prefer. But it has failed, because you did not ask for a better statue. You asked for this particular statue, modified according to a particular intention.
This distinction is becoming increasingly important with generative AI.
Generation asks:
What would be good here?
Faithful execution asks:
What does this particular person want changed, what do they want preserved, and why?
Those are different abilities.
Writing reveals the problem
Writing makes this distinction particularly visible because meaning is not located in words alone.
A piece of fiction can contain literal meaning, semantic meaning, context, metaphor, irony, puns, cultural references, rhythm, ambiguity, subtext and deliberate omission. A character’s history can change the meaning of a single word. A sentence that appears awkward may be deliberately awkward. Sometimes the most important part of a passage is what the author refuses to explain.
And these elements interact.
Writing is therefore like a complicated sauce. You can list the ingredients, but the taste is not contained in any one ingredient. It emerges from the relationship among them.
That is one reason an informal 2026 SilentRoom experiment testing five AI models on Hemingway, Poe, McCarthy and Faulkner is interesting. The test was explicitly journalistic rather than scientific, and each model produced only one short passage per author. Yet the authors found that the models often reproduced recognizable surface features while converging on generic versions of the authors’ styles. In their Hemingway examples, for instance, the models could reproduce short sentences, sparse settings and familiar imagery but struggled with the deliberate omission that is central to Hemingway’s famous iceberg technique.
The lesson is not that AI cannot write.
It clearly can.
The more interesting question is whether reproducing the visible features of a style is the same thing as understanding the organizing principle behind the style.
AI can reproduce the clothes of writing.
But writing may contain a great deal that is not visible in the clothes.
AI can reproduce the ingredients of the sauce. The harder question is whether it understands the taste.
That is a question, not a proof that AI lacks understanding. But it is a useful question because the difference between imitation and intention is becoming increasingly important.
Taste is knowing when the good thing is wrong
Consider a superb dancer. Now ask that dancer to perform like someone who has never learned to dance.
That may be surprisingly difficult.
The dancer’s expertise has become automatic. Balance, rhythm, posture and movement have been trained into the body. To dance badly on purpose, the dancer must suppress some of the expertise that makes them good.
This gets closer to what people mean when they say that AI lacks “taste.”
Taste is not simply knowing what is good. A good writer knows that the technically better sentence may be the wrong sentence for a particular character. A filmmaker may reject the most beautiful shot because it destroys the mood. A sculptor may preserve an imperfection because removing it would destroy the work. A comedian may deliberately ruin a rhythm because the awkwardness is the joke.
Sometimes the correct creative decision is to refuse improvement.
Taste is knowing not only what to do, but when not to do it—and sometimes deliberately doing the wrong thing for the right reason.
AI can imitate awkwardness. It can imitate an amateur. It can produce deliberately ugly prose when asked.
But that does not necessarily tell us whether it possesses the same kind of judgment that allows a human creator to make an apparently bad choice because, in that particular context, it is exactly the right choice.
What Stockfish teaches us
Chess provides another useful analogy.
Stockfish is vastly stronger than any human chess player. But imagine that the interesting question is not whether it can find the best move. Instead, ask what kind of intelligence is revealed when a player has spent decades surviving bizarre, uncomfortable and apparently hopeless positions.
Human chess experience is full of failure. Players blunder, panic, misjudge positions, defend lost games, discover unexpected resources and remember what went wrong. That history contributes to human intuition.
None of this means Stockfish cannot defend a difficult position. It obviously can. The point is different: superhuman performance inside a formal domain does not tell us everything about the kind of intelligence involved in navigating that domain.
Real life contains something chess engines do not encounter in quite the same form: the problem itself may be badly specified.
The objective may be ambiguous. The information may be incomplete. The incentives may conflict. The person giving the instruction may not even know what they want.
The machine may be excellent at finding the best move while the human problem is deciding which game we should be playing.

The research problem: execution is not judgment
A recent experiment brings this distinction out of the laboratory.
Peter Kirgis and colleagues tested frontier AI agents on two open-ended research problems drawn from unpublished NeurIPS 2026 submissions. The agents were given six days and thousands of dollars of compute. They completed the engineering required to conduct the projects without human help, including running experiments and dealing with technical problems.
Yet both research projects were rejected by the original authors.
The researchers identified five recurring problems: poor judgment about the standard required for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness and instruction drift. A second model/scaffold combination reproduced the pattern.
This is an important result precisely because the agents were not incapable of working.
They could work. They could execute. They could experiment. They could debug.
What they struggled with was deciding whether they were doing the right research.
The distinction is fundamental. Research is not simply the execution of a sequence of experiments. At some point a researcher has to decide that a hypothesis is uninteresting, that an experiment is badly designed, that a result is too weak to matter, or that an entire line of investigation should be abandoned.
The machine can search a space.
The harder problem is recognizing that it may be searching the wrong space.
Negative feedback is not the same as insight
This is perhaps the most revealing part of the research-agent findings.
When an experiment fails, a system capable of generating language has many ways to respond. It can revise the hypothesis, add qualifications, narrow the claim, propose another experiment or reinterpret the result.
A human researcher may instead reach a more disruptive conclusion:
The whole approach is wrong.
That is not merely another step in the search. It is a decision to change the search space itself.
The agents in the study often responded to criticism without making sufficiently fundamental changes to their research direction. The researchers describe them as committing too quickly to unpromising approaches and struggling to make good use of feedback.
This suggests a distinction that deserves much more attention in AI research:
Optimization is not the same as reorientation.
A system can become extremely good at improving the solution it is pursuing without becoming equally good at recognizing that it should pursue a different solution.
And then humans have to supervise the machine
The obvious response is that this is why humans remain in the loop.
Let the AI execute. Let the human provide judgment.
But another recent paper raises an uncomfortable problem: increasing agent autonomy may itself degrade the human oversight on which we depend.
Margaret Mitchell, Avijit Ghosh and Samir Passi argue that extended use of AI systems can contribute to approval fatigue, automation bias, reduced situational awareness and skill atrophy. As agents perform more of the work, humans can be pushed into a less active role in which they skim plans, grant permissions and reconstruct what happened after the fact rather than deeply evaluating the work as it unfolds. The authors propose cognitive scaffolding, strategic friction and better approval mechanisms to preserve effective human oversight.
The problem is therefore potentially recursive.
The more capable the agent becomes, the less humans may do themselves. The less they do, the less they practice the skills needed to evaluate the agent. As those skills weaken, effective oversight becomes harder. That can create even greater dependence on the agent.
The machine does not have to eliminate human judgment.
It may be enough for humans to stop exercising it.
We may be building a judgment trap
Put the two studies together and an unsettling possibility emerges.
AI agents may not yet possess enough judgment to replace human researchers, managers or supervisors.
At the same time, increasing automation may reduce the amount of judgment humans themselves exercise.
That creates a dangerous asymmetry:
The machine gets better at execution while the human gets less practiced at judgment.
This is a different kind of automation risk from the familiar fear that AI will simply take people’s jobs.
The deeper risk is that AI could take over so much of the practice of a skill that humans eventually lose some of the ability to evaluate the skill.
A calculator does not merely save arithmetic time. It also means fewer people practice arithmetic.
Autopilot does not merely reduce workload. It changes how pilots maintain situational awareness.
AI agents could do something similar to professional judgment.
The irony is uncomfortable:
We may build machines that need human judgment at exactly the same time that we build systems that make humans practice less of it.
The temptation to turn behavior into a human property
There is another reason to be careful about judgment: we repeatedly make the same mistake when interpreting AI itself.
An LLM becomes more susceptible to adversarial prompts after repeated biased tuning, and a viral post announces that AI has depression. The underlying Scientific Reports paper is much more restrained: it investigates whether repeated alignment-like tuning can progressively increase adversarial susceptibility, using the psychiatric concept of “kindling” as an analogy. The main experiment used TinyLlama-1.1B and a lightweight supervised-fine-tuning proxy rather than full RLHF.
Another paper asks whether intensive chatbot interaction can contribute to psychosis-like phenomena. Its authors describe “AI psychosis” as an emerging label, but explicitly say that current evidence is limited to media reports, case reports and early observational data. They are asking whether it warrants recognition as a distinct clinical entity, not announcing that such a medical condition has already been established.
And then there is consciousness.
Alexander Lerchner’s The Abstraction Fallacy argues that computational functionalism confuses abstract computation with the physical instantiation of experience. It distinguishes simulation from instantiation and argues that software or syntactic architecture alone cannot constitute consciousness. This is a serious philosophical argument, but it is an argument—not an empirical demonstration that “Google DeepMind proved AI can never be conscious.” Lerchner’s own paper carries a disclaimer that its framework and conclusions represent his research and do not necessarily reflect the official views of his employer.
The common error is worth noticing.
We see an AI behave like something, and we infer that it has that thing.
- It behaves as though depressed: therefore depression.
- It talks as though conscious: therefore consciousness.
- It writes as though Hemingway: therefore understanding.
- It conducts experiments: therefore research ability.
The reverse error is possible too. If a system fails to reproduce a particular human capability, that does not automatically prove that it lacks every deeper property associated with that capability.
Behavior is evidence.
It is not always an explanation.
We keep confusing the performance with the performer
This may be the central problem.
AI performance is visible. The internal reasons behind that performance are much harder to establish.
A model can produce a moving poem without us knowing whether there is anything corresponding to the human experience that inspired such a poem.
It can reproduce Hemingway’s surface without necessarily reproducing Hemingway’s reasons for leaving something unsaid.
It can conduct hundreds of experiments without necessarily understanding why one research question is worth pursuing and another is not.
It can execute a human’s instruction while missing the human’s intention.
And it can produce a better answer while still giving the wrong answer—because the user’s objective was not “make this better.” It was “make this the way I want it.”
That is the difference between competence and judgment.
The real AI problem may be knowing what matters
We have spent years asking whether machines can become as intelligent as humans.
That question is becoming too crude.
A machine can calculate faster than us, remember more than us, search more than us, generate more than us and execute more than us.
But intelligence is not only the ability to produce an answer.
It also involves knowing when the answer is wrong, when the question is wrong, when the objective is wrong, when an imperfection should remain, when the obvious solution should be rejected and when the entire direction of work should change.
Those capabilities are much harder to benchmark because they are difficult to specify in advance.
It is easy to ask:
Can the machine solve this problem?
It is harder to ask:
Can the machine recognize that this is the wrong problem?
It is easy to ask:
Can the machine improve this paragraph?
It is harder to ask:
Can it understand why I don’t want it improved?
It is easy to ask:
Can the agent run the experiment?
It is harder to ask:
Can it recognize that the experiment should never have been run?
Perhaps the missing ability is not intelligence but judgment
None of this requires the conclusion that AI is stupid.
Quite the opposite.
The interesting possibility is that AI can become extraordinarily intelligent in some dimensions while remaining strangely weak in others.
It may outperform us at calculation, memory, search, pattern recognition, simulation and execution.
Humans may remain unusually valuable for framing problems, interpreting ambiguity, preserving intention, recognizing context, exercising taste and deciding what matters.
But there is a catch.
If we delegate too much of the first set of abilities, we may gradually weaken the second.
That is why the future question is not simply whether AI becomes smarter than humans.
It is whether humans remain judgmentally competent enough to use, direct and challenge increasingly capable machines.
The stone crusher may eventually become powerful enough to carve almost anything.
The question will still be:
Who decides what should be carved?
And perhaps that is the real dividing line.
AI may be learning how to produce almost anything. The harder problem is knowing what is worth producing, what must be preserved, what should deliberately remain imperfect—and when the entire project should be abandoned.
That is judgment.
And if there is one thing we should be careful not to automate away, it may be that.
Sources & Further Reading
- Cheung, Ngo. “Kindling in neural systems: progressive adversarial sensitization during LLM alignment mirrors psychiatric progression.” Scientific Reports, 2026.
- Lerchner, Alexander. “The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness.” 2026.
- Au Yeung, Joshua; Morrin, Hamilton; Ng, Vincent; Kraljevic, Zeljko; Dobson, Richard. “An Echo Chamber of One: Should AI Psychosis Be a Distinct Clinical Entity?” arXiv:2608.23937, 2026.
- Kirgis, Peter et al. “Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies.” arXiv:2607.27191, 2026.
- Mitchell, Margaret; Ghosh, Avijit; Passi, Samir. “AI Agents Push Humans Out of the Loop.” arXiv:2608.23642, 2026.
- Souly, Alexandra et al. “Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples.” arXiv:2510.07192, 2025.
- SilentRoom. “Hemingway’s Writing Style, as Faked by Five AI Models.” SilentRoom Journal, 2026.

