Infinite Graph
Journal

JOURNAL

The Architecture of Software Is Not the Architecture of Understanding

The architecture of a codebase and the architecture you need to understand it are not necessarily the same graph.

Most code analysis tools begin from a reasonable assumption: if we can reconstruct the relationships that exist inside a repository, we can make the repository easier to understand.

So they map dependencies.

A imports B.
B calls C.
C depends on D.

At sufficient scale, these relationships become a graph. Nodes represent files, modules, classes, functions, or services; edges represent imports, calls, inheritance, data flow, or other forms of dependency. The resulting visualization can be technically precise. It may tell us, with remarkable fidelity, how the software is connected.

But technical precision is not the same thing as explanatory usefulness.

A dependency graph describes the software. It does not automatically explain the software.

The Problem With Treating Structure as Explanation

Consider a large production repository.

A small utility module might be imported by hundreds of files. A logging package, configuration helper, shared type definition, or serialization function may therefore become one of the largest hubs in the dependency graph.

From a graph-theoretic perspective, this is important information. The module has high connectivity. Removing or modifying it may affect a substantial portion of the system.

But imagine onboarding a new engineer.

Would you begin by saying:

“Start here. This utility module has the highest degree centrality in the repository.”

Probably not.

The module may be structurally central while being conceptually peripheral.

What the engineer actually needs to know is likely something closer to:

What does the product do?
→ Which subsystem owns this behavior?
→ Where does the relevant state live?
→ How does data move through the system?
→ Which modules implement that behavior?

Only after establishing this conceptual framework do low-level dependencies become meaningful.

This reveals an important distinction: the order in which software executes is not necessarily the order in which software should be explained.

Two Graphs Exist Inside Every Codebase

We can think of a sufficiently complex software system as containing at least two different graphs.

The first is the execution graph.

This graph represents the relationships required for the software to operate. It captures imports, function calls, service dependencies, event propagation, data flows, package relationships, and runtime interactions. Its purpose is descriptive: it tells us what depends on what.

The second is the explanation graph.

This graph represents the relationships a human needs in order to construct a useful mental model of the system. Its edges are not necessarily imports or calls. They may instead represent relationships such as:

feature → subsystem

subsystem → responsibility

responsibility → state

state → lifecycle

concept → implementation

behavior → relevant modules

The execution graph answers:

“How is this software connected?”

The explanation graph answers:

“What should I understand first so that the next thing makes sense?”

Those are fundamentally different questions.

Why the Two Graphs Diverge

Software architecture is optimized primarily for machines and engineering constraints.

Modules are separated to control coupling. Shared utilities are extracted to reduce duplication. Services communicate through interfaces. Libraries are introduced to encapsulate reusable behavior. Build systems impose their own dependency boundaries. Performance requirements may introduce caches, queues, indexes, workers, and asynchronous execution paths.

Human understanding follows different constraints.

A developer does not build a mental model by traversing every import edge. Understanding is hierarchical and selective. We begin with high-level concepts, identify responsibilities, establish boundaries, and progressively attach implementation details to those concepts.

In other words, comprehension depends on semantic relevance, not merely structural connectivity.

A database adapter may sit deep inside the dependency graph while being essential to understanding where persistence occurs. A product-level concept such as “workspace permissions” may not exist as a single module at all; its implementation may be distributed across API handlers, policy functions, database models, middleware, and frontend state.

The concept is coherent to a human even when it is fragmented in the code.

This means that a useful explanation graph may contain nodes that do not literally exist in the repository.

“Authentication” may be a node.

“Document lifecycle” may be a node.

“Billing state” may be a node.

“Realtime collaboration” may be a node.

These are semantic abstractions rather than files, but they are often the abstractions through which engineers actually reason about the system.

Accuracy Can Produce an Unreadable Graph

This creates a subtle problem for code intelligence tools.

The more faithfully a tool reproduces the raw dependency structure of a large repository, the more visually and cognitively complex the result can become.

Thousands of nodes and tens of thousands of edges may constitute an extremely accurate representation of the codebase while providing very little guidance about where a human should begin.

The failure is not one of correctness.

It is one of representation.

A map containing every road, pipe, electrical cable, property boundary, elevation contour, and sewer connection in a city might be extraordinarily accurate. It would still be a poor subway map.

The subway map succeeds precisely because it does not attempt to preserve every property of physical reality. It preserves the relationships necessary for a particular cognitive task: understanding how to move through the transit system.

Code maps may require the same distinction.

The goal should not always be to ask:

“How can we visualize the repository more completely?”

A more useful question may be:

“Which representation allows a developer to construct the correct mental model with the least unnecessary cognitive work?”

From Code Graphs to Knowledge Graphs

This distinction becomes particularly important as code analysis systems evolve from static visualization tools into code knowledge systems.

A useful code knowledge graph should arguably preserve both representations rather than forcing one to serve both purposes.

The raw dependency view remains valuable. Engineers need to inspect actual imports, callers, downstream dependencies, package boundaries, and execution paths. These relationships are indispensable when debugging, refactoring, estimating blast radius, or investigating implementation details.

But a conceptual view serves another function.

It can organize the same repository around behaviors, responsibilities, domains, state ownership, and architectural concepts. Instead of asking the engineer to infer the product architecture from thousands of low-level edges, the system can provide intermediate semantic layers through which the underlying code becomes navigable.

The result is not a replacement for the dependency graph.

It is an additional graph built for a different reader.

One graph optimizes for structural fidelity.

The other optimizes for cognitive orientation.

A mature code intelligence system may need both.

The Real Unit of Code Understanding

This distinction also suggests that the file may be the wrong fundamental unit for explaining software.

Repositories are stored as files because files are useful implementation artifacts. Developers, however, frequently reason in terms of behaviors and responsibilities.

When investigating a problem, an engineer rarely thinks:

“I need to understand files 47 through 63.”

They think:

“Where is authorization decided?”

“Who owns this state?”

“What happens after this event is emitted?”

“Which component is responsible for synchronizing this data?”

“Where does this user-visible behavior actually come from?”

These questions cut across directory boundaries and dependency hierarchies.

A code understanding system therefore becomes substantially more useful when it can move between two levels of representation: from concept to implementation, and from implementation back to concept.

That bidirectional relationship may be more important than simply generating a larger dependency graph.

Designing for Understanding, Not Just Inspection

The broader implication is that code intelligence should not treat visualization as the final objective.

The objective is comprehension.

That requires recognizing that software possesses multiple valid structures depending on the question being asked. Runtime behavior has one structure. Package dependencies have another. Organizational ownership may have another. Product concepts may have another still.

The best representation is therefore not necessarily the graph that most faithfully mirrors the repository.

It is the graph that preserves the relationships relevant to the task.

For debugging, that may be an execution path.

For refactoring, it may be a dependency graph.

For incident response, it may be a service and data-flow graph.

For onboarding, it may be a conceptual hierarchy that begins with product behavior and progressively reveals the implementation beneath it.

The important design principle is not to confuse these representations.

The architecture that makes software executable and the architecture that makes software understandable solve different problems.

And once we recognize that distinction, a different question emerges for anyone building tools for code comprehension:

If you were onboarding a new engineer tomorrow, would you really show them the raw dependency graph first?