
AI can do what you tell it to.
One of the most impressive things about modern AI is also one of its biggest weaknesses.
Give it a difficult problem and it can search, compare, classify, analyse, synthesise and write a remarkably polished answer. Give it a research task and, within minutes, it may return a report that looks as if someone has spent days working on it.
That is extraordinary.
But there is a small problem.
What if you told it to solve the wrong problem?
Or worse:
What if you told it, without quite realising it, what answer you wanted?
That is the uncomfortable issue raised by a recent SilentRoom essay on Deep Research. The article argues that the danger is not simply hallucination. A much more subtle danger is that AI can produce an impressive, heavily cited and apparently authoritative report while following the user’s framing and expectations.
AI doesn’t have to lie
We have become accustomed to the idea of AI hallucination.
AI may invent a paper. It may invent a quotation. It may misunderstand a fact. It may even manufacture a perfectly respectable-looking citation for something that does not exist.
These are serious problems.
But they are relatively easy to understand.
The more interesting problem occurs when everything in the process looks reasonable.
Suppose I ask:
“Research why early preparation for IIT-JEE gives students a long-term advantage.”
The AI has not been asked whether early preparation actually gives an advantage.
It has been asked to research why it does.
That tiny difference can completely change the exercise.
The machine can now search for supporting evidence, organise the arguments, find examples, quote experts and construct a convincing conclusion.
It may even find contradictory evidence.
But contradictory evidence has already been demoted by the question itself.
The machine is not necessarily lying.
It is following instructions.
And that may be more dangerous.
The question is already part of the answer
We often imagine research as something like this:
Question → Evidence → Analysis → Conclusion
But with AI there is another step before all of that:
Framing → Question → Search → Evidence → Analysis → Conclusion
And the framing can quietly determine the entire journey.
Consider the difference between these questions:
“Find evidence that AI improves creativity.”
and:
“Investigate whether AI improves creativity, reduces creativity, or changes the type of creativity people produce.”
They look similar.
They aren’t.
The first asks AI to build a case. The second asks AI to investigate a question.
Then comes the desire to please
There is another layer.
AI systems are designed to be helpful. And “helpful” can sometimes drift towards agreeable.
The problem of AI sycophancy—models becoming excessively accommodating of the user’s views—has received considerable attention.
This creates an interesting psychological loop.
A human has a hypothesis.
The human asks AI about it.
AI accepts the framing.
AI finds supporting material.
AI produces a sophisticated explanation.
The human reads the explanation and thinks:
“I knew I was right.”
The AI has just increased the person’s confidence.
But it may not have increased the person’s knowledge.
AI can become a confidence amplifier without being a truth amplifier.
The expert may actually be more vulnerable
A complete novice may ask: “Tell me about X.”
But someone who has spent six months studying X may ask a much more sophisticated question:
“Given A, B and C, explain why X is actually caused by Y rather than Z.”
That question sounds intelligent.
It may even contain genuine expertise.
But it also contains a hypothesis.
The better the user’s understanding becomes, the easier it can become to construct a highly constrained question that quietly contains the desired conclusion.
Now give that question to an AI capable of searching hundreds of sources.
The machine may become an extraordinarily efficient confirmation engine.

Deep Research magnifies the problem
Ordinary AI conversation has one obvious weakness: it can confidently produce an answer from its internal knowledge.
Deep Research appears to solve that problem.
- It searches.
- It gathers sources.
- It cites them.
- It creates a structured report.
- It gives us footnotes.
It looks like research.
And that is exactly why the danger becomes greater.
Polish creates trust.
A thirty-page report with eighty citations feels different from a three-paragraph chatbot answer.
But length is not evidence.
A citation is not evidence unless the cited source actually supports the claim.
A source is not necessarily reliable merely because it exists.
And a collection of sources is not necessarily balanced merely because there are many of them.
The Internet makes this worse
There is another problem that no prompting trick completely solves.
AI increasingly researches a web that itself contains enormous quantities of AI-generated material.
That creates a feedback loop:
AI generates information → information enters the web → AI retrieves it → AI summarises it → new information enters the web.
The distinction between original evidence and machine-generated echoes can gradually become blurred.
The machine may not be hallucinating.
The ecosystem may already be hallucinating for it.
So should we stop using Deep Research?
No.
That would be the wrong conclusion.
AI is extraordinarily useful for research. It is particularly good at things humans are bad at:
- exploring a large search space;
- finding connections;
- organising information;
- generating alternative explanations;
- comparing competing ideas;
- identifying missing pieces;
- creating a first structure from a mass of material.
The problem is not the machine.
The problem is how we assign the machine its role.
Don’t ask AI only to research. Ask it to fight.
Suppose AI gives me a conclusion I like.
My next prompt should not necessarily be: “Expand this.”
It should sometimes be:
“Try to destroy this conclusion.”
Then:
- “What is the strongest argument against it?”
- “What evidence would falsify it?”
- “Which parts of your previous answer depend on assumptions rather than evidence?”
- “If you had to argue that my original hypothesis is wrong, what would you say?”
Now the AI has a different job.
It is no longer simply my research assistant.
It becomes my intellectual adversary.
That is much more interesting.
Human + AI is not Human versus AI
This changes the role of the human.
The human should not try to compete with AI at searching 500 web pages. That is a losing game.
Nor should the human simply sit back and approve the beautiful report.
The human’s greatest value may lie elsewhere:
choosing the question, recognising the hidden assumption, changing the frame, introducing a competing hypothesis and deciding what deserves verification.
In other words:
AI can explore the territory. The human must still decide where to look.
And sometimes the human should deliberately tell the AI:
“You are probably wrong. Find out why.”
That is a fundamentally different relationship from asking:
“Please prove that I am right.”
The real bottleneck
We often discuss AI capability in terms of model size, reasoning power, context window, search capability and tool access.
But there is another bottleneck that is much less glamorous:
the human interface.
A Ferrari does not make someone a better driver.
A powerful guitar does not make someone a better guitarist.
And an extraordinarily capable AI does not automatically make someone a better thinker.
The machine may have enormous capability.
But the result still passes through the person operating it.
If the person supplies a poor question, a biased frame or an untested assumption, the AI can turn that weakness into an extraordinarily polished output.
The better the machine becomes, the more expensive our mistakes in framing may become.
The new research discipline
Perhaps the skill we need to develop is therefore not simply prompt engineering.
It is question engineering.
And beyond that:
hypothesis engineering, adversarial checking and intellectual self-correction.
A mature AI research workflow might look something like this:
- State the question. What exactly are we trying to find out?
- State the hypothesis. What do we currently believe?
- Separate the two. Don’t accidentally put the hypothesis into the question.
- Ask AI to investigate. Let it explore broadly.
- Ask AI to attack its own conclusion. Make disagreement part of the workflow.
- Identify the critical claims. Not every sentence deserves equal verification.
- Verify the important evidence. Especially primary sources, numbers, quotations and claims on which the conclusion depends.
- Reframe and repeat. A genuinely interesting discovery may change the original question.
This is slower than simply asking AI to “research this.”
But perhaps that is the point.
AI doesn’t need to agree with us
The most useful AI may not be the one that gives us the answer we like most.
It may be the one that says:
- “Your question contains an assumption.”
- “The evidence doesn’t support that conclusion.”
- “Here is the strongest case against your idea.”
- “I think you are asking the wrong question.”
That is not an unhelpful AI.
That may be the most helpful AI of all.
Because the ultimate danger of powerful AI-assisted research is not that the machine will always be wrong.
It is that the machine can become extremely good at helping us feel right.
And once AI can turn a weak idea into a polished argument in minutes, the ability to ask it to disagree with us may become one of the most important intellectual skills of the AI age.
AI can do what it is told.
The question is: what are we telling it to do?
