At first, knowledge graphs appear to be systems about relationships.
Which concept is connected to which claim?
Which person works for which organization?
Which software module depends on which library?
Which paper supports which conclusion?
Because edges are visually salient, graph design is often framed primarily as a problem of relation modeling.
But there is a more fundamental decision that comes earlier.
Before we can ask what should be connected, we have to decide what is allowed to exist as a node.
Suppose we take a single document and attempt to transform it into a graph.
What should become a node?
The document?
A section?
A paragraph?
A sentence?
A claim?
A concept?
An entity?
A decision?
Every one of these choices is plausible.
And every one produces a different graph.
That is the uncomfortable part.
The hardest problem in knowledge graph design may not be determining which edges to create.
It may be deciding what counts as a unit of knowledge in the first place.
The hardest part of a knowledge graph may happen before the first edge exists.
A Graph Cannot Escape Its Choice of Atoms
Every graph begins with a discretization problem.
Continuous, contextual, overlapping information has to be divided into objects that can be assigned identity.
Only after those objects exist can the system connect them.
This act of division is easy to underestimate because it often happens inside preprocessing pipelines.
A document is segmented.
Chunks are produced.
Entities are extracted.
Nodes are instantiated.
Then graph construction begins.
But graph construction has already begun.
The decision about where one unit ends and another begins determines the graph's topology before a single semantic edge is added.
If the units are too large, the graph becomes semantically coarse.
Imagine representing an entire twelve-page technical report as one node.
You can connect it to another report with:
Document A related_to Document B
That relation may be correct.
It is also weak.
Perhaps only one paragraph in Document A contradicts one claim in Document B.
Perhaps both documents share a methodology but disagree on conclusions.
Perhaps one merely cites the other.
At document-level granularity, these distinctions collapse into a vague connection between two large containers.
The graph knows that a relationship exists, but not where its meaning lives.
Now move to the opposite extreme.
Suppose every sentence becomes an independent node.
A few hundred documents can become tens of thousands of fragments.
Many of those fragments depend on context:
“This approach solves the problem.”
“However, this assumption does not hold.”
“It therefore performs better.”
As isolated nodes, these sentences may be grammatically complete while remaining semantically incomplete.
The graph gains resolution but loses coherence.
Push atomization further and the graph begins to resemble an inverted index with edges.
The system contains enormous numbers of addressable fragments, yet few of them are meaningful knowledge objects.
This produces a fundamental tradeoff.
Too coarse, and relations become vague.
Too fine, and identity becomes meaningless.
Bad graph structure often starts with bad atomization, not bad linking.
Granularity Is Not a Cosmetic Choice
Node granularity affects almost every downstream capability.
It changes retrieval.
It changes clustering.
It changes summarization.
It changes provenance.
It changes deduplication.
It changes synchronization.
It changes how conflicts are detected.
It changes which nodes users can sensibly edit.
It changes which relations are worth storing.
It changes what an AI system can reason over.
Consider two representations of the same material.
In one system, each paragraph is a node.
In another, each independently verifiable claim is a node.
These systems may contain the same underlying text, yet behave very differently.
A paragraph-oriented graph preserves source structure well.
It is easy to map back to the original document.
But one paragraph may contain three claims, an example, and a qualification.
A claim-oriented graph can connect propositions much more precisely.
One claim can support another.
Two claims can contradict each other.
A claim can be traced to multiple sources.
But extracting claims requires interpretation, and different models may disagree about where one claim ends and another begins.
The choice is therefore not between a correct representation and an incorrect one.
It is between representations optimized for different operations.
Granularity is an architectural decision because it determines what kinds of meaning the graph can make explicit.
Chunking Is Not the Same as Atomization
This becomes particularly important in RAG systems.
Retrieval pipelines routinely divide documents into chunks.
A typical chunk may contain a few hundred tokens.
The purpose is pragmatic.
The chunk must be large enough to preserve useful context and small enough to retrieve selectively.
But a retrieval chunk is not necessarily a knowledge object.
That distinction matters.
A 500-token segment may exist because it fits well into an embedding pipeline.
Its boundaries may be determined by token count, paragraph breaks, or overlap heuristics.
Nothing about those boundaries guarantees that the chunk corresponds to one coherent idea.
It is a retrieval unit.
A node in a persistent knowledge graph carries a stronger implication.
Once something becomes a node, it may receive:
a persistent identifier,
edges,
metadata,
provenance,
annotations,
history,
permissions,
synchronization state,
embeddings,
references from other objects.
The system is effectively asserting:
This thing has enough internal coherence and stability that other knowledge may refer to it as one object.
That is a much stronger claim than:
This span of text is convenient to retrieve.
Confusing these two concepts is dangerous.
A retrieval chunk is optimized for model context.
A knowledge atom is optimized for identity.
Those objectives overlap, but they are not equivalent.
A GraphRAG system that simply converts every retrieval chunk into a permanent graph node may inherit arbitrary chunk boundaries as ontology.
At that point, a computational convenience has silently become a theory of knowledge.
Node Identity Is an Ontological Commitment
The deeper issue is identity.
When we create a node, we are not merely storing content.
We are saying that something exists as a distinguishable object.
That creates questions that ordinary text storage can avoid.
If two paragraphs express the same concept, are they one node or two?
If the same claim appears in three documents, is it one claim with three sources or three source-bound claims?
If a claim is revised, does the node change?
Or should a new version become a new node?
If two concepts are later discovered to be equivalent, should they be merged?
If a node is split into two more precise ideas, what happens to all incoming edges?
These questions are not implementation trivia.
They determine how the graph understands persistence.
An edge is meaningful only if the identities at its endpoints are meaningful.
A perfectly classified relationship between badly defined objects is still a bad graph.
This is why knowledge graph design is fundamentally a problem of ontology before it is a problem of topology.
Topology asks how objects connect.
Ontology asks what objects there are.
You cannot solve the first cleanly without answering the second.
The Correct Atom Depends on the Domain
The atomization problem becomes even harder because there is no obvious universal unit.
Different domains organize knowledge around different objects.
In legal analysis, the meaningful unit may be:
a clause,
an obligation,
an exception,
a precedent,
a holding,
a legal test.
In software engineering, it may be:
a symbol,
a function,
a class,
a module,
an interface,
an architectural decision,
an invariant.
In scientific research, it may be:
a hypothesis,
a method,
a measurement,
a dataset,
a result,
a claim.
In product work, it may be:
a user problem,
a requirement,
a decision,
an experiment,
an assumption,
an observation.
In personal knowledge management, it could be almost anything.
A note.
A thought.
A quotation.
A task.
A question.
A concept.
A decision.
A source.
A project.
The graph cannot treat these as interchangeable without losing important semantics.
A legal clause has boundaries imposed by textual structure.
A software symbol has boundaries imposed by the language.
A product decision may be distributed across multiple meetings and documents.
A concept may have no obvious textual boundary at all.
This means that “What should a node be?” does not have one globally correct answer.
The answer depends on what type of object the domain itself treats as meaningful.
The Correct Atom Also Depends on the Question
Even within the same domain, there may be no single optimal granularity.
Consider a software repository.
For one query:
Which subsystems depend on authentication?
Module-level nodes may be ideal.
For another:
Which function mutates the session token?
Module-level nodes are too coarse.
Symbol-level representation becomes necessary.
Now consider:
Why was the refresh-token flow changed?
Code symbols may no longer be enough.
The relevant atom might be an architectural decision or a discussion thread.
The same corpus therefore supports multiple legitimate knowledge resolutions.
This produces an important conclusion:
The optimal knowledge atom may be query-relative.
That is a serious problem for systems that insist on one permanent granularity.
A graph optimized for broad exploration may be poor at precise factual retrieval.
A graph optimized for sentence-level verification may be unreadable to humans.
A graph optimized for source fidelity may be poor at conceptual synthesis.
No single node size necessarily dominates.
That makes the idea of a universal knowledge atom increasingly questionable.
The Universal Knowledge Atom May Be the Wrong Goal
A great deal of knowledge-system design implicitly searches for the perfect unit.
Should every node be atomic?
Should one node contain one idea?
Should every claim become independent?
Should concepts be canonicalized?
These questions assume that there is one ideal level at which knowledge naturally decomposes.
There may not be.
Knowledge is often hierarchical.
A document contains sections.
Sections contain arguments.
Arguments contain claims.
Claims mention concepts.
Concepts recur across documents.
A single statement may simultaneously belong to multiple structures.
Trying to force all of this into one node type can flatten useful distinctions.
Instead of asking:
What is the universal knowledge atom?
A better question may be:
At which resolutions should this knowledge be addressable?
That changes the architecture.
The graph no longer needs to choose one level.
It can support multiple levels explicitly.
Multi-Resolution Graphs
A more flexible representation might look like this:
Document
→ containsSection
→ containsPassage
Then:
Passage
→ expressesClaim
And:
Claim
→ referencesConcept
Meanwhile:
Claim
→ supported_bySource
Now the graph preserves several different forms of identity.
The document remains important as a source object.
The passage preserves context.
The claim becomes a proposition that can be compared or contradicted.
The concept becomes a reusable semantic object.
Each node type exists because it supports a different operation.
This allows the graph to change resolution depending on the task.
For broad navigation, it can remain at document or section level.
For semantic reasoning, it can descend into claims.
For conceptual exploration, it can move across canonical concepts.
For provenance, it can trace any synthesized object back to its source passage.
This is fundamentally different from merely storing everything at maximum granularity.
The graph is not just fine-grained.
It is multi-resolution.
Granularity Should Be Navigable
Once multiple resolutions exist, the problem changes.
Instead of asking how to find one perfect atom, we can ask how to move between resolutions without losing identity.
This introduces useful operations.
A system might expand:
Architecture Decision
into:
Claim AClaim BClaim C
Or collapse those claims back into a single higher-order object.
A concept node might aggregate dozens of source-bound mentions.
A debugging incident might expand into individual observations, actions, and decisions.
The user's mental model can remain coarse until more detail is required.
This is especially valuable in large knowledge graphs.
A graph with 100,000 low-level nodes may be semantically rich but visually useless.
A multi-resolution system can expose perhaps fifty meaningful high-level nodes, while preserving thousands of lower-level objects underneath.
Granularity becomes a property of navigation rather than a one-time preprocessing choice.
That is a much stronger architecture.
AI Makes the Boundary Problem Harder
Human-authored documents rarely announce their semantic atoms explicitly.
A paragraph boundary is visible.
A conceptual boundary often is not.
One argument may span four paragraphs.
A single paragraph may contain several independent claims.
A concept may appear under different terminology.
A design decision may emerge across a meeting, an issue thread, and a pull request.
When AI converts this material into graph structure, it is performing segmentation at the semantic level.
It must decide:
whether two statements express one claim,
whether a new node is necessary,
whether an existing concept should be reused,
whether one passage contains multiple ideas,
whether two ideas should be merged,
whether a distinction is meaningful enough to preserve.
This means AI-generated graph construction is not merely extraction.
It is ontology formation.
That is a substantially more difficult task.
The model is deciding what kinds of things exist inside the system.
If those decisions are unstable, the graph becomes unstable.
One import may create one node.
Another import may create three.
A later model version may merge them.
Edges, provenance, embeddings, and references now depend on inconsistent segmentation.
The failure is not primarily bad retrieval.
It is identity drift.
Atomization Is Also a Versioning Problem
Once nodes persist, changes in granularity become expensive.
Suppose a system initially treats the following paragraph as one node.
Later, it determines that the paragraph contains three independent claims.
What should happen?
The original node may already have:
incoming edges,
outgoing edges,
comments,
embeddings,
user annotations,
references,
synchronization history.
Splitting the node is now a graph migration.
The inverse problem exists as well.
If three nodes later turn out to represent the same semantic object, merging them requires conflict handling.
This shows why granularity should not be treated as a disposable preprocessing decision.
Atomization affects long-term schema stability.
Once identity exists, changing identity has consequences.
Granularity Controls Provenance
The issue also matters for trust.
Suppose a synthesized concept node says:
Method X reduces inference latency.
Where did that statement come from?
If the graph stores only document-level provenance, the answer might be:
“From this 40-page paper.”
That is weak provenance.
If the system stores claim-level objects linked to exact passages, the evidence can be much more precise.
But extremely fine granularity creates its own problems.
A sentence may appear to support a claim only when interpreted together with the preceding paragraph.
The graph therefore needs not merely small units, but context-preserving relationships between them.
Good provenance requires both precision and recoverable context.
Again, atomization cannot be solved simply by making nodes smaller.
Granularity Controls Retrieval
Retrieval quality is also deeply affected by graph resolution.
Suppose the user asks:
Why did we abandon architecture A?
A document-level graph may retrieve the correct design document but return too much irrelevant material.
A sentence-level graph may retrieve isolated remarks but miss the structure of the decision.
A decision-level node might be ideal:
Architecture A rejected because X, Y, and Z.
But that node may not exist naturally in the source.
It may need to be synthesized from multiple passages.
This suggests an important distinction between source atoms and semantic atoms.
Source atoms correspond to where information physically appears.
Semantic atoms correspond to what the information means.
These two structures do not have to align.
A robust PKM system may need both.
Source Structure and Knowledge Structure Are Different
Documents organize information for reading.
Knowledge graphs organize information for reference and reasoning.
Those goals are not identical.
A section heading exists because it helps a reader navigate a document.
A claim node may exist because it helps a system compare propositions.
A paragraph may group sentences rhetorically.
A concept node may group ideas semantically.
The source's layout should therefore not automatically become the graph's ontology.
This is especially important when importing information.
Naively converting every heading into a parent node and every paragraph into a child node preserves document structure.
But it may produce a graph of the document rather than a graph of its knowledge.
Those are different objects.
The first models where information was written.
The second models what information exists.
A sophisticated system needs to preserve the relationship between them without conflating them.
Visual Simplicity Can Hide Ontological Complexity
One reason atomization receives less attention is that graph interfaces encourage users to think visually.
Nodes are circles or cards.
Edges are lines.
The design question appears to be:
How should this graph be displayed?
But visual clarity is downstream of semantic structure.
A beautifully rendered graph with badly chosen nodes remains a bad graph.
No amount of layout optimization fixes an ontology where entire documents are treated as indivisible units when the task requires claim-level reasoning.
Conversely, a technically precise graph with one node per sentence may be impossible for a human to navigate.
The challenge is not merely visual density.
It is choosing which identities deserve representation at each level.
Graph UX therefore cannot be separated from graph ontology.
The interface ultimately exposes assumptions about what counts as one thing.
A Better Design Principle: Identity Before Connectivity
This suggests a useful order of operations for knowledge graph design.
Do not begin with:
Which relationships should we support?
Begin with:
Which objects deserve persistent identity?
For each candidate node type, ask:
Is this object semantically coherent?
Can it be referenced independently?
Does it remain meaningful outside its original location?
Can it change without becoming a different object?
Does the domain naturally recognize this as a unit?
Will users want to search for it?
Will AI systems reason over it?
Can provenance be attached to it precisely?
What happens if it is split or merged?
At what resolution is it useful?
Only then should edge design begin.
This produces a more stable ontology because relationships are built on top of identities that have been deliberately chosen.
The Graph Is a Theory of What Exists
Every knowledge graph contains an implicit worldview.
If documents are nodes, the system treats documents as primary objects.
If claims are nodes, it treats propositions as primary objects.
If concepts are nodes, it treats semantic abstractions as primary objects.
If tasks and decisions are nodes, the graph becomes operational.
The ontology encodes what the system believes deserves durable identity.
That is why atomization is not merely a parsing problem.
It is a modeling decision.
And modeling decisions always exclude alternatives.
A graph cannot represent everything equally at every resolution without becoming unusable.
It must choose what to foreground.
The important question is whether those choices are explicit or accidental.
Toward Query-Relative Knowledge Representation
The strongest alternative may be to stop assuming that one representation should serve every query.
A future knowledge system might maintain multiple coordinated layers.
A structural layer preserves documents, sections, and passages.
A semantic layer contains concepts, claims, entities, and decisions.
An operational layer contains tasks, procedures, and state transitions.
A retrieval system can select the appropriate layer depending on the question.
For:
Where was this mentioned?
Use the source layer.
For:
What arguments support this conclusion?
Use the claim layer.
For:
What concepts connect these two research areas?
Use the semantic layer.
For:
What should I do next?
Use the procedural layer.
This turns granularity from a global constant into a query-dependent choice.
The system does not search for one perfect knowledge atom.
It maintains interoperable forms of identity at different scales.
The Real Design Problem Comes Before the Edge
Knowledge graphs are often introduced through relationships because edges make graphs distinctive.
But edges are only meaningful after identity has been established.
Before we can say:
A supports B
we must know what A and B are.
Before we can say:
A contradicts B
we must know whether each proposition has been separated correctly.
Before we can say:
A depends_on B
we must decide whether A is a repository, module, class, function, or architectural capability.
The quality of the edge depends on the quality of the atoms.
That is why poor graph structure often begins before linking.
It begins during segmentation.
During abstraction.
During canonicalization.
During the decision to say:
This is one thing.
And perhaps that is the more fundamental challenge for PKM, RAG, and GraphRAG systems.
Not merely how to connect knowledge.
But how to decide what knowledge is allowed to possess identity.
Because there may be no universal knowledge atom.
The right unit may depend on the domain, the task, the query, and the level of reasoning required.
The best graph may therefore not be the one that discovers the perfect node size.
It may be the one that allows knowledge to exist at multiple meaningful resolutions without losing provenance, context, or identity.
The hardest part of a knowledge graph may happen before the first edge exists.
So before asking which nodes your PKM should connect, ask something more fundamental:
What is the smallest unit of knowledge in your system—and why does that unit deserve to exist as one thing?
