This distinction presents an increasingly important problem for the design of AI-powered knowledge systems. In discussions of artificial intelligence, explainability is generally treated as an unquestionable virtue. When an AI system recommends a document, connects two concepts, ranks a search result, or suggests a node in a knowledge graph, users naturally want to know why. Designers therefore introduce features such as a “Why was this recommended?” button, allowing a large language model to produce a natural-language explanation of the system’s decision. At first glance, such a feature appears to increase transparency. However, the presence of an explanation does not necessarily mean that the system has revealed the actual basis of its decision.
The central problem is that the explanation itself may also be generated content.
Consider an AI knowledge system that connects two nodes, A and B. Internally, the recommendation may have occurred because the vector representations of the two nodes were sufficiently close according to an embedding similarity metric. Perhaps their cosine similarity exceeded a predefined threshold, or perhaps one node appeared among the top results produced by a retrieval model. These computational signals constitute part of the actual mechanism through which the recommendation was produced.
Yet if a large language model is subsequently asked, “Why are these concepts related?”, it may respond with a statement such as: “Both concepts concern long-term knowledge preservation and the organization of institutional memory.” The explanation may be coherent, semantically appropriate, and persuasive. It may even be substantively reasonable. Nevertheless, unless that semantic relationship was itself part of the decision process, the explanation is not a description of the mechanism that produced the recommendation. It is a plausible interpretation generated after the recommendation has already been made.
This is the critical distinction:
an explanation of the mechanism is not the same as a plausible explanation generated after the fact.
The difference resembles the distinction between causation and narrative reconstruction. A system can produce an outcome for one computational reason and later generate a convincing verbal account based on another. In such cases, the language model is not necessarily reporting the system’s internal reasoning. It may instead be rationalizing the outcome. The danger arises precisely because the rationalization is often linguistically better than the underlying evidence.
A visibly incorrect recommendation may invite skepticism. A fluent explanation, by contrast, can suppress it.
This creates a particularly subtle form of epistemic risk. Users tend to interpret explanations as evidence of transparency. When an interface presents a confident natural-language rationale, the user may assume that the text reflects the actual process by which the system reached its conclusion. The explanation therefore acquires an authority that ordinary generated text may not possess. What appears to be an interpretability feature can consequently become another hallucination surface.
Explainability is not useful if the explanation is another hallucination surface.
This problem is especially significant in AI knowledge tools because their recommendations are often not produced by a single reasoning process. A suggested connection may depend on embedding similarity, keyword overlap, retrieval ranking, graph structure, user behavior, metadata rules, recency signals, or manually configured heuristics. In more complex systems, several of these signals may operate simultaneously. A language model that is asked to explain the recommendation afterward may have access only to the final objects being compared, not to the exact evidence that caused the recommendation to appear.
Suppose, for example, that a system recommends a research note because a specific sentence in that note closely matches the semantic representation of the user’s query. A genuinely transparent explanation would expose that source fragment and identify its role in retrieval. It might state that the note was returned because a particular passage produced a high similarity score with the query. By contrast, asking an LLM to summarize why the two documents “seem related” could produce a much broader conceptual story — perhaps one involving organizational learning, collective memory, or knowledge continuity. Such a story may be intellectually convincing while remaining disconnected from the actual retrieval event.
The same problem applies to graph-based knowledge systems. Imagine that a platform suggests creating a link between two nodes because users frequently opened them during the same session. If the explanatory interface later claims that the nodes are connected because they share a conceptual theme, the system has substituted an interpretation for a causal account. The relationship may still be meaningful, but the explanation has changed categories: it no longer answers, “Why did the system make this recommendation?” It answers a different question: “What plausible relationship can be described between these two items?”
That distinction should be made explicit in AI interface design.
A robust explainability system should therefore treat explanations not primarily as generated prose but as presentations of evidence. When users ask “Why?”, the interface should first reconstruct the decision path that existed at the time of recommendation. Depending on the system, this may include retrieval signals, source fragments, similarity scores, graph edges, ranking factors, explicit rules, metadata matches, or previous user actions. Natural language can still play an important role, but it should summarize or contextualize these underlying signals rather than invent an independent justification.
In other words, the evidence should constrain the explanation.
For instance, instead of presenting only the sentence, “These concepts are related because they both concern long-term knowledge preservation,” an AI knowledge system might display:
“Recommended because three passages in Node A were semantically similar to Node B. The strongest match was the phrase concerning the preservation of institutional knowledge. Similarity score: 0.87.”
The system could then allow the language model to provide an interpretation:
“This suggests that the two nodes may be conceptually related through the broader theme of long-term knowledge preservation.”
The difference between these two statements is fundamental. The first reports evidence from the decision process. The second interprets that evidence. Combining them can be useful. Confusing them is dangerous.
This approach also changes the design objective of explainability. The goal should not be to generate the most convincing explanation. It should be to preserve the provenance of a decision. In trustworthy AI systems, explanation quality should therefore be evaluated not only according to fluency, clarity, or user satisfaction but also according to faithfulness: does the explanation correspond to the information that actually influenced the model or system?
This principle becomes even more important as AI systems grow more agentic and personalized. Future knowledge tools may continuously recommend documents, reorganize information, create graph relationships, summarize user activity, and infer possible conceptual connections. If every automated action is followed by a polished but unverifiable narrative, users may gradually lose the ability to distinguish between system evidence and machine-generated interpretation. The interface may appear increasingly transparent while becoming epistemically opaque.
The design challenge, therefore, is not merely to make AI explain itself. It is to ensure that what the AI calls an explanation remains connected to the causal or evidential path that produced the decision.
For systems such as AI-powered knowledge graphs, this suggests a practical design principle: the “Why?” button should not simply prompt a language model to write a reason. It should retrieve the evidence that existed within the recommendation pipeline and make that evidence inspectable. The interface might show which source fragments were retrieved, which graph relationships contributed to the suggestion, which rule was activated, what similarity signal was observed, or which previous user action influenced the ranking. The LLM may then translate those signals into human-readable language, but it should not be allowed to replace them.
Such a distinction may seem technical, but it has broader implications for trust in artificial intelligence. Users do not merely need AI systems that can tell persuasive stories about their outputs. They need systems that can distinguish between what actually caused an output and what merely makes the output sound reasonable afterward.
The future of explainable AI may therefore depend less on improving the eloquence of explanations and more on improving their evidential grounding.
When an AI explains why it made a recommendation, the most important question may not be whether the explanation sounds good.
It is whether the explanation is true to the process that actually produced the recommendation.
When an AI explains why it made a recommendation, do you want a good explanation — or the actual reason?
