Abstract
The historical extent of human knowledge is still organized primarily around printed and digital documents. Research papers, journals, books, databases, citations, conferences, repositories and institutional archives preserve enormous amounts of information, but they poorly represent the actual structure by which knowledge evolves.
A scientific claim is rarely an isolated document. It depends on prior propositions, evidence, assumptions, definitions, models, experiments, objections, replications, failures, revisions and competing interpretations. Yet these relationships remain fragmented across publications and institutions.
The Emergent Epistemic Ledger (EEL) proposes an open-source infrastructure for representing knowledge as a continuously evolving, independently verifiable epistemic graph. Its central object is not the paper. It is the claim and its history.
Each claim can be linked to its dependencies, supporting evidence, contradictions, replications, simulations, revisions, alternative formulations and downstream consequences. A signed, append-only ledger preserves the provenance of those epistemic state transitions without declaring any claim permanently true. Researchers, institutions, computational systems and eventually AI agents may contribute to a shared ontology while retaining full attribution and preserving minority branches.
The ledger itself carries no currency. It records research requests, the work done in answer to them, and who validated that work, so that funding and resource exchange can be built on top of a public record rather than inside it. Credit for validated work is non-transferable, and it is earned for carrying out a test correctly, whatever the result.
The long-term objective is not to construct an unquestionable universal worldview. It is to create a globally collaborative process through which an increasingly coherent ontology can emerge through variation, constraint, testing, selection, retention and revision. The system therefore attempts to make the process of knowledge formation itself open source.
1. The problem
Scientific knowledge is distributed across millions of documents. This architecture creates several limitations.
- Research conclusions are easier to locate than the complete reasoning histories that produced them.
- Negative results and failed replications are systematically underrepresented.
- Citation counts poorly distinguish foundational contribution, replication, falsification, criticism, formalization or downstream integration.
- Scientific disagreement is usually encoded rhetorically in papers rather than explicitly in machine-readable dependency and contradiction graphs.
- Knowledge provenance is vulnerable to institutional fragmentation, link rot, revisions, withdrawn artifacts and changing access policies.
- Research resources are often allocated through indirect institutional mechanisms rather than toward the questions with the greatest unresolved epistemic value.
Most importantly, the scientific record does not function as an integrated model of what humanity currently believes and why. EEL treats this as an infrastructure problem.
2. From documents to epistemic objects
The basic unit of EEL is an addressable epistemic object: a proposition, definition, piece of evidence, observation, experiment, dataset, objection, counterexample, replication, mathematical model, simulation, method, prediction or revision.
A proposition such as “Stable effective spatial dimensionality can emerge from lower-order relational dynamics” receives a persistent identifier. Its epistemic neighborhood could include:
CLAIM
├── DEPENDS_ON → lower-order persistence
├── DEPENDS_ON → recursive composability
├── SUPPORTED_BY → simulation A
├── CHALLENGED_BY → objection B
├── REPLICATED_BY → experiment C
├── ALTERNATIVE_TO → model D
└── REVISED_INTO → claim version 4
The ontology is therefore not a collection of declarations. It is a graph of relationships among contestable propositions.
The relationships are contestable too. That a piece of evidence supports a claim is itself a judgment, so every edge should be a first-class object with an author, a history, and objections of its own. A graph that hid its most important judgments inside its structure would undermine the transparency it exists to provide.
3. The ledger does not record truth
A foundational principle is:
Ledger consensus ≠ truth consensus.
A distributed, cryptographically signed ledger can establish whether a record existed, when it existed, who signed it, and whether it has since been altered. It cannot determine whether nature agrees with the claim. EEL therefore separates two forms of verification.
Cryptographic verification answers:
- Was this contribution submitted?
- Who signed it?
- Which prior objects did it reference?
- Has its content changed?
- Was this artifact available in this form at this time?
Epistemic verification asks:
- Does the evidence support the claim?
- Was the experiment correctly designed?
- Can the result be independently reproduced?
- Are there counterexamples?
- Is there a simpler model?
- Does the claim survive adversarial testing?
The ledger secures the first. The research community performs the second.
4. An evolutionary ontology
The system should not presume that a single ontology can be correctly specified in advance. Instead, ontologies branch and compete.
main
├── relational-emergence
├── informational-primitive
├── structural-realist
└── alternative-model
Branches preserve common ancestry while making disagreements explicit. Canonicalization therefore means the current best-supported shared ontology under explicitly stated epistemic procedures. It does not mean final truth.
A minority branch can later become dominant if it explains more evidence, requires fewer unsupported assumptions, generates better predictions or survives stronger tests. Some branches, however, may be empirically equivalent: no available evidence separates them. In that case the shared, canonical layer is what all of them agree on, namely the observed regularities, and the branches remain parallel above it until an experiment is found whose predicted outcome differs between them.
The ontology evolves through:
variation → constraint → testing → selection → retention → revision
This architecture deliberately parallels the broader emergence framework from which the project originated. That framework is the motivating case, not a privileged branch: it enters the ledger on the same terms as any other.
5. Epistemic provenance
Every contribution retains ancestry. A claim may evolve through:
P₁ → O₁ → E₁ → R₁⁻ → P₂ → R₂⁺ → P₃
where P is a proposition, O an objection, E an experiment, R⁻ a failed replication and R⁺ a successful replication.
The historical record is never destroyed. A correction is a new event that points to what it corrects. The ledger therefore becomes a form of collective epistemic memory: retained distinctions across transformations in collective belief.
6. Trusted genesis nodes
A decentralized epistemic system faces an unavoidable bootstrap problem. At launch, no contributor has network-derived reputation, because no contribution history yet exists. EEL therefore begins with genesis validators: explicitly designated, first-order trusted nodes whose initial authority derives from the founding constitution. Their powers are constrained.
- Domain-bounded authority. A validator trusted in theoretical physics does not automatically receive authority in oncology, economics or linguistics.
- Multisignature procedures. No genesis validator can canonicalize a consequential claim alone.
- Public provenance. Validation activity is attributable and inspectable.
- Sunset mechanics. Genesis privileges diminish as network-derived reputation becomes sufficient.
- Declared interests. Genesis validators publish their own stake in the claims they validate, and do not validate claims they authored.
The intended transition is:
constituted trust → earned trust
Initial trust is granted. Enduring trust must be earned through independently verifiable epistemic performance.
7. Validation is not canonicalization
A validator may confirm: this experiment was executed according to the declared protocol and produced the reported result. That is not equivalent to: this proposition is true.
Canonicalization should require accumulated evidence across several dimensions, for example:
evidence + independent replication + objection resolution + methodological review → eligibility for canonicalization
The ontology therefore distinguishes procedural validity from epistemic standing.
8. Ontology-driven research requests
The ontology itself can identify missing knowledge. Suppose proposition P₂₀ depends on P₄ and P₉, but P₉ has weak evidence, and P₉ is also required by P₃₁, P₄₄ and P₇₁. The system can identify P₉ as an epistemic bottleneck and generate a research request:
TARGET: P9
NEEDS: 2 independent replications
1 alternative-model comparison
1 new dataset
EXPECTED IMPACT: resolves dependencies affecting 4 downstream propositions
A research request is a ledger object like any other. It names the claim, the kind of work wanted (an experiment, a replication, a simulation, a formal or literature review, data collection, a search for counter-models) and the conditions under which the work counts as done. Supporters can attach commitments to it, researchers submit results, and validators check them. Attention and effort can then move toward the unresolved regions of the knowledge graph.
9. Rewarding epistemic work
The infrastructure should reward improvements in collective knowledge rather than attention. It should not reward views or popularity. It should recognize replication, falsification, high-quality review, prediction, data production, integration, formalization, methodological correction and useful computation.
A failed experiment can therefore be valuable. If it removes a large unsupported region of possibility space, it has created knowledge.
The most important incentive rule is that a researcher is credited for carrying out a test validly, whatever its outcome. If credit depended on a hypothesis surviving, researchers would be paid to confirm it. In EEL, a result is credited once enough validators confirm that it was produced according to the declared protocol, whether it supports the claim, contradicts it, is null or is inconclusive.
This credit is a non-transferable record of validated work. It cannot be bought or sold, and it does not by itself confer reputation or authority.
10. Separating the roles around a test
A test of a claim involves several distinct roles, and EEL keeps them apart:
- Supporters back a research request because they want the question investigated.
- Executors perform the initial experiment or simulation.
- Replicators independently reproduce the result.
- Validators check that each step followed its declared protocol.
Commitments attached to a request record which kind of work they are meant for: execution, replication, or general support for a line of inquiry. Separating these roles reduces conflicts of interest. Resources fund investigation. Evidence determines epistemic standing.
11. Independent confirmation
If a result is independently reproduced, the network has acquired stronger evidence. But replication counts must measure genuine independence: ten accounts controlled by one laboratory must not equal ten replications.
Independence has several dimensions: separate identities, institutions, data, methods and implementations. Most of these cannot be computed by a protocol. They are judgments. EEL therefore treats a statement of independence as an attested claim in the graph, made by named parties and open to objection like any other claim, rather than as a score the system calculates. How such attestations should be weighed is one of the project’s main open design problems.
12. Information gain
The infrastructure should value uncertainty reduction rather than confirmation alone. If H_before is the uncertainty among competing models before a test and H_after the uncertainty afterwards, then
IG = H_before − H_after
is a conceptual measure of information gain. A decisive falsification may be worth more than another confirmation of an already robust claim, so research requests should favor experiments with high expected discriminatory power.
Two conditions keep this measure honest. First, information gain is only defined relative to a registered comparison set: a declared list of candidate models, each stating what outcome it predicts for the proposed test. That register is itself a contestable object in the graph, which stops anyone from inflating the measure by inventing weak rivals for a test to eliminate. Second, what should be valued is the discriminatory power declared before the test, not surprise after it. Rewarding dramatic results after the fact would recreate the very incentive the system is designed to remove. A test can discriminate between two models only where they predict different outcomes for it.
13. Computation as evidence
Research questions can become computational jobs. For example: search for relational update rules that produce persistent higher-order organization without encoding geometry beforehand. Workers search different parameter regions, promising structures become artifacts, independent nodes reproduce them, and validated results attach to the relevant claims. This produces a recursive loop:
ontology → research question → computation → evidence → ontology revision
For this to be trustworthy, simulations must be deterministic, so that any validator can replay a run from its declared inputs and obtain the same result byte for byte. Initial verification relies on deterministic replay and redundant independent execution. Later work may add random partial challenges or verifiable-computation proofs.
The status of such evidence must be stated carefully. A simulation that finds a rule producing stable three-dimensional structure shows that the mechanism can produce it. It does not show that nature does. The graph should record simulation results as demonstrations of possibility, distinct from observations, so that model vocabulary is never silently promoted into an asserted ontology.
Computation of this kind is never used to secure the ledger itself. Useful scientific searches lack the properties a consensus mechanism needs (tunable difficulty, unpredictable instances, no advantage from choosing the problem), so the ledger’s integrity rests on signatures and validator procedures alone.
14. Reputation
Epistemic reputation must not be purchasable. Economic stake, epistemic reputation and governance authority remain separate:
stake ≠ reputation ≠ governance
Reputation should be multidimensional and domain-specific. Possible dimensions include replication reliability, formal reasoning, experimental design, simulation verification, objection quality, prediction calibration, data quality and ontology integration. A researcher may be highly trusted in one domain without gaining automatic authority elsewhere.
15. Attribution
The ledger distinguishes different intellectual contributions: origination, formalization, validation, falsification, replication, integration, curation, methodological improvement, implementation, data generation and maintenance. This creates a richer intellectual genealogy than citation counts alone.
16. Long-term value
The value of knowledge often becomes visible slowly. A contribution C can therefore have a time-dependent profile of epistemic impact, V_C(t), with dimensions such as validated downstream descendants, replication history, reuse across domains, prediction success, persistence and dependency centrality.
This must not be treated as a single truth score. It describes downstream epistemic usefulness. And any metric that carries reward will be gamed: claims split to multiply nodes, spurious dependencies, allies citing allies, objections raised only to be resolved. Every such profile therefore needs an explicit threat model, and its components should remain open to objection.
17. Governance
The network will need procedures for protocol upgrades, infrastructure, validator rules, research-request design, moderation and repository standards. Whatever form that governance takes, it must not vote scientific propositions into truth. Its constitutional principle is:
Governance controls the epistemic process, not reality.
18. Canonical ontology
The project’s eventual ambition is a globally collaborative, evolutionary ontology. “Canonical” means the current best-supported shared state under transparent rules. It does not mean compulsory, infallible, permanent or unanimous. Alternative branches remain recoverable, and a superior minority branch can eventually displace the incumbent.
19. Interoperability
EEL should work with existing scientific infrastructure rather than replace it: DOIs, ORCID, Git, institutional repositories, preprint servers, research knowledge graphs, decentralized storage, high-performance computing systems and scientific databases.
Large artifacts stay off the ledger. The ledger stores cryptographic commitments (content hashes) and provenance relationships, so anyone can check that an artifact is the one that was registered.
20. Architectural layers
A mature system has these layers:
ONTOLOGY What does the shared model currently contain?
LEDGER How did that model change?
REQUESTS Which unresolved questions need work, and what would settle them?
COMPUTE What simulations or analyses can constrain those questions?
REPUTATION Who has demonstrated reliability in which domains?
GOVERNANCE How are shared procedures and infrastructure maintained?
Together these form an open research operating system.
21. The DATA XCHANGE: a separate layer
Research also needs resources: datasets, computing time, laboratory work, review and replication labor, and funding. Coordinating and paying for these is a different problem from recording knowledge, and it carries different risks, including financial regulation. It therefore belongs to a separate project, the DATA XCHANGE, which is not part of this ledger.
The boundary is simple. The ledger publishes a public record: research requests and their acceptance conditions, the commitments attached to them, the results submitted, whether validators confirmed them, and which replications succeeded. An exchange, or any other funding body, can read that record and decide whom to pay. Nothing in the ledger depends on it, and no amount of money changes what the ledger records about a claim.
22. Philosophical architecture
The infrastructure follows a progression:
distinction → information → evaluation → knowledge → shared meaning → collective inquiry
The network then becomes an externalized collective memory. Humanity produces knowledge. The system remembers how that knowledge changed. That history informs later inquiry. The process becomes recursive:
knowledge → knowledge about knowledge → improved knowledge production
23. Research constitution
The founding constitution includes at least these principles:
- Ledger consensus is not truth consensus.
- Economic stake cannot purchase epistemic authority.
- Evidence can overturn canonical claims.
- Minority branches remain recoverable.
- Falsification is a valuable contribution.
- Replication must be independently attributable.
- Provenance cannot be silently rewritten.
- Scientific claims remain revisable.
- Governance controls procedure, not reality.
- Initial trust may be constituted; enduring trust must be earned.
24. The prototype
An open-source prototype of the ledger exists and can be run locally. Written in Rust and released under the Apache-2.0 licence, it implements signed events and identities, the append-only ledger with Merkle-rooted blocks, the projection of a shared ontology from the event history, branches, domain-bounded and multisignature validation, research requests with outcome-independent credit, independent replications, and deterministic simulation with replay by independent validators. Every node that replays the same events arrives at the same ontology.
The prototype contains no currency, tokens, wallets or payments. Every place where the design was underspecified is recorded, with the choice made, in the repository’s decision log.
25. Open problems
The hardest questions remain open, and the project should be judged partly on whether it makes progress on them:
- Independence. How should attested independence between replications be weighed, and by whom?
- Comparison sets. Who registers the candidate models against which information gain is measured, and how is a register challenged?
- Empirically equivalent branches. How should the system present branches that no current evidence can separate?
- Gaming. Which mechanisms stop contributors from inflating graph metrics?
- Governance. How does authority move from genesis validators to earned reputation without capture?
- Usefulness. What evidence would show that a claim-level ledger actually helps researchers, and what would show that it does not?
26. Long-term vision
The system could grow from an open repository into an epistemic graph, a distributed ledger, a network for distributed computation, and eventually a global knowledge network, with resource exchange built alongside it as a separate layer.
Its purpose would not be to create an authority able to announce what humanity must believe. It would be to make the process by which humanity earns belief inspectable, reproducible, contestable, attributable and globally collaborative.
The ledger provides memory. The ontology provides structure. Research requests direct attention. Distributed computation provides experimentation. Researchers provide judgment. The community provides revision. Together they form a self-extending system through which collective understanding can evolve.
27. Foundational statement
The Emergent Epistemic Ledger is an open-source infrastructure for preserving, testing, exchanging and evolving humanity’s collective models of existence. It records not truth itself, but the independently verifiable history through which claims are proposed, challenged, tested, retained, revised or abandoned.
Its canonical ontology is therefore not humanity’s final answer. It is humanity’s continuously revisable record of what has survived the search so far.