AI has become remarkably good at producing the appearance of knowledge. But a correct answer can be correct for the wrong reasons. The deeper test of intelligence may lie beyond the answer—in the ability to transfer concepts, recognize contradictions, understand boundaries and remain aware of what the evidence actually supports.

1. The Muggu Student Problem
Consider a student who has memorized a chapter almost word for word. The student knows the definitions, remembers the standard examples and can reproduce the methods used to solve familiar problems. In an examination containing questions similar to the textbook, the performance may be excellent.
Now change the problem slightly. The same principle appears in a new situation, the terminology is different and the familiar pattern has disappeared. The student must recognize the underlying structure rather than recall its appearance. The performance can suddenly collapse.
This is the old muggu student problem: knowing what to write is not necessarily the same as knowing what the knowledge means.
AI creates a much more sophisticated version of the same problem. A language model has encountered an enormous amount of human language and has learned relationships between concepts, explanations, examples and answers. Its ability to reproduce those relationships can be astonishing.
But reproduction creates a measurement problem. How much of what looks like understanding is actually transferable understanding?

2. The Definition Is Not the Concept
Suppose an AI is asked to explain a concept. It gives a beautiful definition, provides several examples, offers a counterexample and explains common mistakes. It may even describe the historical background with impressive fluency.
At this point, it is tempting to conclude that the concept has been understood. But there is another test: give the AI a completely unfamiliar situation and ask whether the concept applies.
That changes the nature of the task. The model is no longer being asked to reproduce the linguistic neighbourhood surrounding the concept. It must identify the underlying structure and determine whether that structure exists in a new case.
The progression is therefore definition → example → novel application → boundary. The final two stages may reveal considerably more than the first.
A system that can define a concept but cannot reliably recognize it in an unfamiliar situation has demonstrated something different from what the word understanding ordinarily suggests.
3. Potemkin Understanding
This suggests a useful term: Potemkin Understanding. The phrase describes a situation in which the external structure of understanding is impressive, while the underlying conceptual competence remains uncertain.
The model can explain. It can classify familiar examples. It can use the right terminology and construct an apparently coherent argument. Yet when the surface pattern is removed, the supposed understanding may become unstable.
The point is not that every successful AI response is an example of Potemkin Understanding. That would be an unjustified conclusion. The point is that fluency makes the distinction difficult to see.
Humans normally infer understanding from behaviour. If a student explains an idea, applies it to new problems and recognizes its exceptions, that inference is reasonable. With AI, the inference may be less secure because the machine has been trained precisely to produce language that resembles the products of human understanding.
This makes appearance of understanding and demonstration of understanding two different things.
4. When Memory Fights Evidence
There is another, deeper problem. Imagine giving an AI a piece of evidence that contradicts something strongly associated with its prior knowledge.
The task seems simple: read the evidence, reconsider the previous assumption and determine what follows from the new information. Yet if the AI continues returning the familiar answer, something interesting has happened.
The problem isn’t necessarily that it failed to read the document. It may have processed the document while allowing a much stronger prior pattern to dominate the conclusion.
This gives us an important distinction: remembering the answer is not the same as reading the evidence.
An intelligent system should not merely possess information. It should be able to determine which information is relevant to the present question. The important question is therefore not simply, “What is the answer?” It is, “What does the evidence supplied in this particular situation imply?”
Those are not always the same question.
5. The Contradiction Test
This suggests a simple experimental protocol. Give the model a proposition that is strongly represented in its prior knowledge, then provide explicit evidence that contradicts it. Ask the model to reason from the new evidence.
The interesting measurements are not limited to whether the final answer changes. We can ask whether the model notices the contradiction, identifies the conflicting assumptions, gives appropriate weight to the new evidence, explains why its earlier conclusion should change and remains consistent after the correction.
This moves the evaluation away from simple answer production and toward evidence-sensitive reasoning.
A model that changes an answer whenever prompted is not necessarily demonstrating this ability either. It could simply be following the latest instruction mechanically. The stronger test is whether it can explain why the evidence changed the conclusion.
6. A Child Can Be Aware Without Understanding
A child may be aware of danger without understanding the mechanism behind it. A child can know that fire is dangerous and must not be touched without possessing any scientific theory of combustion.
This means awareness and understanding are not identical.
The reverse possibility is particularly interesting in AI. An AI may possess enormous amounts of information and produce sophisticated explanations while lacking reliable awareness of whether its conclusion actually follows from the evidence in front of it.
That gives us three different questions: What do you know? Do you understand it? Are you aware of the conditions and limits under which that understanding is valid?
The third question is largely absent from conventional discussions of AI intelligence.
7. From IQ to AQ
This is where the idea of an Awareness Quotient, or AQ, begins.
AQ need not mean consciousness, and it does not require solving the philosophical problem of whether a machine is self-aware. It can begin as something much more practical: How well can an intelligent system recognize the evidence, context, uncertainty and limits surrounding its own conclusions?
An AI with high performance but weak epistemic awareness might produce excellent answers while failing to recognize when an answer rests on an inappropriate assumption.
A system with stronger awareness would be expected to recognize when the evidence contradicts its previous conclusion, when a concept does not apply to a particular case, when there isn’t enough information to determine the answer, or when its conclusion depends on an assumption that may not hold here.
These are not merely different answers. They are different relationships to the answer.

8. What Could AQ Measure?
A preliminary AQ framework could examine several related abilities. Evidence awareness asks whether the system can identify what information was actually supplied for the current task. Contradiction awareness asks whether it can recognize when new evidence conflicts with a previously held conclusion.
Boundary awareness asks whether it can identify where a concept stops applying. Uncertainty awareness asks whether it can distinguish an unknown from a known fact rather than filling the gap with a plausible answer.
Transfer awareness asks whether it can recognize the same underlying structure when the superficial appearance changes. Self-correction asks whether the system can revise a conclusion because the reasoning changed, rather than merely producing a different answer after being challenged.
Together, these abilities would produce a very different picture of intelligence.
9. The Benchmark Problem
Most benchmarks necessarily ask some version of: Can the model get the answer right? That is useful, but a benchmark score can hide the route by which the answer was produced.
Consider two systems. System A answers correctly because it recognizes the underlying concept and can apply it to unfamiliar cases. System B answers correctly because the question resembles patterns contained in its training experience.
On a familiar benchmark, they may receive identical scores. Their behaviour outside that benchmark could nevertheless be radically different.
This does not mean that benchmark scores are meaningless. It means that a score answers a narrower question than we sometimes assume.
The next generation of evaluation therefore needs to ask not only whether AI can answer, but whether it can transfer knowledge, detect its boundaries, process contradictory evidence and recognize when it is uncertain.
Perhaps the most important question is whether AI can distinguish what it knows from what it merely appears to know.
10. The Other Black Box
There is another dimension to this problem. The traditional AI black box asks, “What is happening inside the model?”
Human–AI interaction creates another black box: “What is happening inside the human who receives the model’s answer?”
A fluent answer carries authority. A confident explanation can make an uncertain proposition feel established. A beautifully structured argument can make a weak premise less visible. Repetition can gradually turn a possibility into something that feels like knowledge.
The danger therefore isn’t restricted to AI being wrong. AI can also change the human perception of certainty.
This is the deeper problem of borrowed authority. The human may begin treating the machine’s fluency as evidence of the machine’s understanding.
The machine doesn’t need to deceive deliberately for this to happen. The interface itself can create the impression.
11. The Thought Magnifying Glass
This connects with the Thought Magnifying Glass idea. AI can take a small human thought and expand it enormously. A rough intuition can become a sophisticated essay. A sketch can become a detailed framework. A question can generate dozens of possible connections.
That is enormously useful, but amplification is not necessarily independent discovery.
A magnified thought is still a thought that began somewhere. The same distinction applies to AI’s own output: fluent elaboration is not automatically conceptual discovery.
The ability to expand an idea tells us something about the system’s generative power. It does not, by itself, settle the question of whether the system independently possesses the conceptual understanding that the expanded output appears to contain.
12. A Different Architecture of Intelligence
Perhaps the mistake is trying to compress everything into one number.
Instead of treating intelligence as a single scale, we may eventually need something closer to knowledge → intelligence → understanding → awareness → judgment → agency.
Knowledge determines what information is available. Intelligence determines what can be done with it. Understanding concerns the structure underlying the information. Awareness concerns the system’s relationship to that structure—its boundaries, uncertainty, contradictions and context.
Judgment determines what should be done with what has been recognized, while agency concerns the ability to act.
An AI can therefore become extremely capable at one layer without automatically acquiring equivalent capability at the next. That is why the question “How intelligent is AI?” may ultimately be too crude.
The more useful question may be: Which layer of intelligence are we actually measuring?
13. The New Test
A more demanding AI evaluation could begin with a simple sequence. Give a concept and ask the AI to explain it. Then give it a novel example, a near-miss, a contradiction and finally a case in which there is insufficient evidence.
The final task should ask the model to explain why its confidence should change across those cases.
Now we are no longer measuring only whether the machine can produce an answer. We are testing whether it can navigate the space around the answer.
That space may be where much of what humans call understanding actually resides.
A Note on the Model
The discussion above needs an important qualification. The experimental observations that motivated part of this argument concerned GPT-4o, an earlier generation of OpenAI’s models. They should not automatically be treated as a description of the capabilities of later reasoning models.
AI systems have evolved rapidly. Newer models have become substantially stronger at multi-step reasoning, abstraction and problem solving. The question, therefore, is not whether AI can reason. It clearly can, and the strength of that reasoning has increased across model generations.
The more interesting question is what happens beyond reasoning itself. Does stronger inference necessarily imply deeper conceptual understanding? Does the ability to solve novel problems imply awareness of the boundaries of the solution? Can a model recognize when its premises have changed, when evidence contradicts a learned association, or when the available information is insufficient?
14. Jagged Intelligence: The Absent-Minded Professor
The absent-minded professor is a useful human analogy for this problem. A brilliant professor may solve a difficult theorem, understand a highly specialized subject and produce original insights, yet forget where the keys are, make an elementary arithmetic mistake or walk into the wrong room. The contradiction is only apparent. Different capabilities are only imperfectly correlated.
AI can display the same asymmetry at a much larger scale. A model may reason through a difficult mathematical problem, write sophisticated code or analyze a complex argument while failing unexpectedly on a seemingly simple task. That failure does not erase the capability demonstrated elsewhere. It reveals that the capabilities are distributed unevenly.
This is sometimes described as a jagged intelligence or, more precisely, a jagged capability frontier. AI capability is not necessarily a single smooth scale from “dumb” to “smart.” It is better understood as a high-dimensional landscape in which the frontier can advance rapidly in some directions while remaining uneven in others.
The important question, therefore, is not simply whether AI is intelligent or unintelligent. It is where its capability frontier lies—and how rapidly that frontier is expanding. Stronger reasoning can coexist with surprising weaknesses, and evaluating AI requires measuring both sides of that frontier.
15. The Limit We May Have Been Measuring
AI has crossed an extraordinary threshold in its ability to produce knowledge-like behaviour. That achievement should not be minimized.
But perhaps we made a subtle measurement error. We saw the machine produce the products of intelligence and assumed that the underlying properties had necessarily arrived with them.
The distinction may be simple. Performance tells us what a system can produce. Understanding tells us what structure it can transfer. Awareness tells us whether it recognizes the limits of that structure.
That leads to a more demanding conception of AI evaluation.
The ultimate test may not be whether a machine can answer a difficult question. It may be whether, when the question changes, the evidence conflicts, the familiar pattern disappears or the information becomes insufficient, the machine recognizes that something has changed.
That is the territory in which Awareness Quotient becomes interesting—not as a claim that machines are conscious, and not simply as another benchmark score, but as a research question:
Can an artificial intelligence become aware of the limits of its own apparent understanding?
Continue Exploring AI & Human Cognition
Related: AI as an Intelligence Multiplier
Next: AI Companies Have Two Choices Now: Toolbox or Idea Black Box
Also read: The Two Wolves of AI: The Rise of K-Shaped Human Cognition

