The Discovery Engine

How a reviewed result changes what a field does next.

A watercolor of a brass-and-gold orrery turning quietly on a weathered stone deck at the edge of a calm bay at dusk, distant soft mountains, a small sailboat far off, and a faint gold constellation tracing the sky as warm light pools on the water.

The Failed Run

On Friday, a foundation must choose which of three experiments to fund. A lab’s run failed on Tuesday, but the evidence is still trapped in an instrument export and a private thread. On Wednesday, a model proposed the next experiment from a month-old review. No one has recorded whether the failure narrows the old finding, invalidates the assay, or simply calls for replication.

The foundation may fund another month of work based on an assumption the lab challenged three days earlier. That is a sophisticated way to repeat a mistake.

This gap will matter more as models produce more scientific work. Candidate generation is already getting cheaper in coding and mathematics, where programs can be tested and formal objects can be checked. Other fields are becoming more programmable too. But ten times as many results will not mean ten times as much progress if every result must be reconstructed before anyone can use it.

Output is not progress. Constellations of Borrowed Light argues why science needs a shared state layer at all, and why activity is not state. This essay is the operation that layer has to perform: what happens, exactly, between a result and the work that follows it.

A discovery engine earns its name only if a bounded result can change what gets tried next without allowing the producer, the model, or a machine check to decide what the field believes.

Standards and registries already preserve provenance, research objects, corrections, and structured claims. Each preserves part of the path. No common cross-domain layer presently joins them into an authority-scoped state transition, then binds the next task to the resulting state.These are complementary precedents, not failed attempts at the same system. See the FAIR principles, W3C PROV, RO-Crate, Crossmark, nanopublications, and the Open Research Knowledge Graph.

The engine does two things. It turns an attempt into a contestable change to a maintained frontier. Then it binds new work to the frontier it inherited.

The Transition

Start with Tuesday’s failed run. The instrument trace, calibration record, raw readout, and operator note are activity. They matter, but they are not yet accepted state.

The lab packages them in a receipt. The receipt identifies the attempt, the frontier root it consumed, the artifacts it produced, and the information needed to inspect or replay the work. It proves neither that the run was sound nor that a scientific claim should change. It says, in a durable form, what happened.

Landing the receipt opens a proposal and diff. The producer might propose lowering confidence in a finding, adding a caveat, or opening a replication obligation. That diff is an interpretation, not a verdict. A reviewer may instead find a contaminated sample, a weak assay, or a population mismatch. The review path must allow the original proposal to fail and narrower proposals to follow.

The record can retain rejected and deferred proposals with their evidence and review reasons without changing the root.

Checks and attestations answer different questions. Did the submitted bytes pass the named check? Does the checked statement match the claim people think it matches? What does the evidence warrant under these conditions? May this reviewer accept this class of change? More signatures cannot turn one copied judgment into independent evidence.

Any actor may propose. Only a human reviewer, or a service operating under a prior human-signed and scope-limited policy, may accept a change for a governed root. The credential records a delegation from a named authority and charter. Possessing a key does not create legitimate authority.

An accepted proposal becomes a canonical event. It records the authorized mutation and the reason for it. A signature can authenticate the responsible key, but it cannot establish scientific correctness, confer credit, or authorize clinical action. In Vela’s current working-draft frontier, some events can be unsigned. An authority-bearing public acceptance needs an authenticated decision under that frontier’s rules.

A reducer replays the event history and produces a new root. Given the same valid events, schemas, reducer rules, and referenced artifacts, independent implementations should derive the same accepted state. Replay tests whether the record is reproducible. It does not test whether the accepted judgment was wise.

Fig. 01.The thin waist. Activity becomes accepted state through explicit checks and scoped review. Only an authorized event changes the frontier. Replay produces the new root. Impact and task candidates remain non-authoritative derived views, and general task compilation is still proposed.

Some evidence cannot be public. Participant data, confidential manuscripts, trade secrets, and dual-use work may require controlled access. A public record can still carry a hash, an access class, safe metadata, and the authority’s rationale while the payload remains in an enclave. The hash can show that hidden bytes did not change. It cannot make them publicly reviewable.The NIH controlled-access guidance describes privacy, consent, re-identification, and group-harm reasons for limiting access to scientific data.

Science cannot just copy Git. A code patch does not have a reagent lot, patient consent, or a noisy instrument. But Git gets one separation exactly right. It makes a proposed change legible against a particular history, then leaves maintainers to decide whether it belongs there. A discovery engine applies that discipline without pretending that an experiment is a patch.

This narrow transition is the engine’s waist. Models can propose at one edge. Journals, registries, funders, and labs can use the resulting state at another. The shared layer only needs to preserve what was proposed, what was checked, who accepted what, under which standard, and which root followed.

The Next Task

A replayed root is useful, but it is still a record of the past. The engine becomes interesting when that state changes what someone attempts next.

An obligation names unfinished work and its discharge condition. “Improve this bound” is a wish. “At root R17, produce a construction that satisfies predicate P and exceeds the recorded value k” is an exact obligation. An empirical obligation might ask whether an effect survives in a named population under a named assay, comparator, and preregistered analysis.

A root-bound task instantiates an obligation against a particular state. It records the inherited root, the relevant dependencies, the allowed scope, the evidence standard, and the review or safety requirements. If accepted events move the frontier from R17 to R18 before the result returns, reviewers can ask a precise question: did anything the task depended on change?

An unrelated event need not make the task stale. A new construction in another problem family should not invalidate this one. A revised assay condition might. Root binding replaces a vague warning that “the literature changed” with a reviewable account of what changed and why it matters.

Fig. 02.The core loop. The upper label states whether an object is activity, canonical state, or a derived view. The lower label states its implementation boundary. Vela currently governs proposal, acceptance, and replay. A root-bound task exists in one exact profile, while a general next-task compiler remains proposed.

General task compilation is not a current Vela capability. A complete engine could let a curator select an open obligation under a visible policy, bind it to the current root, and prepare a bounded task. Today that contract has been exercised in one exact mathematical profile. Choosing obligations across science remains partial or proposed.

That limitation goes to the heart of the idea. Choosing which questions deserve a field’s time is an objective function, and no field has ever written it down. The engine can make the choosing inspectable. It cannot do the choosing. If one funder defines every target, one registry recognizes every reviewer, or one school controls the vocabulary, the queue will faithfully encode those priorities. Competing roots, public selection policies, appeals, and audits are part of the mechanism because selection can be wrong.

The same root can support several maps. A funder may order obligations by cost and uncertainty, while a laboratory groups them by instrument or dependency. Each projection should name its source root and selection policy. It guides attention without changing accepted state.

Some obligations should never enter an open queue. A safety policy may restrict the evidence, send task design to an institutional review body, or deny the proposed work before execution. The public record can preserve that a decision occurred without publishing a hazardous payload. Restrictions should still carry a reason, an accountable authority, and an appeal or review path where law and safety allow.

Now return to the failed run. The lab emits a receipt against R17 and proposes a confidence reduction. An instrument check validates the trace, but a reviewer finds a contamination flag and rejects that proposal. Separate proposals add a caveat and open a replication obligation. An authorized reviewer accepts those narrower changes, and replay yields R18. A proposed task compiler can now prepare the replication against R18 while flagging only the older tasks whose dependencies changed.

The reviewer left the finding’s confidence unchanged, added a caveat, and opened a replication obligation. R18 carries that judgment into the next decision.

The First Wedge

Mathematics supplies the cleanest first test because a frozen verifier can decide whether one finite witness satisfies one fixed predicate. A proof assistant can check a formal term against a small kernel. These checks are cheap, exact, and unforgiving.

They are also narrow. A passing witness says nothing by itself about novelty, statement faithfulness, explanation, credit, or scientific value. A kernel checks the theorem it was given, not the theorem a mathematician hoped had been formalized. Researchers and maintainers still make those judgments.

Vela currently implements proposal-backed changes, canonical events, deterministic replay, exact-verifier attachments, and cross-language conformance fixtures. Human keyholders can accept changes directly. A scope-limited policy signed in advance can also admit a gate-clean proposal. Deferrals remain pending, denials change no state, and an uncovered proposal returns to a human. This route automates a prior policy, not scientific judgment. Root-bound tasks exist in one Sidon-set profile, while general task compilation, empirical writeback, federation, and institutional action remain partial or proposed.Implementation status was checked against the current Vela repository and its vendored protocol on 2026-07-10. These labels describe capability, not adoption.

The existing evidence should be read by class. An external sequence registry adopted one bounded result. Local frontier histories preserve accepted constructions and failed attempts. Lean artifacts supply kernel-checked proof closure for separate formalization tasks. None of these proves that the general discovery loop works.The external registry example is OEIS A309370. It is evidence that one result crossed an institutional boundary, not evidence for the broader architecture. For the distinction between automated search, formal verification, and field judgment, see Asterisk’s “Automating Math” and the 2026 survey AI for Mathematics.

The next test is public and falsifiable. An outside producer must consume a public root and submit a receipt without working inside the Vela repository. A human reviewer must accept or reject the proposed change with a registered key. An independent reducer must derive the same new root. Then a second outside producer must receive a task bound to that root and return another attempt.

The test fails if outside producers need a maintainer to write their receipts, if independent replay diverges, if a relevant state change leaves the next task untouched, or if an irrelevant change rewrites it. It also fails its practical purpose if review costs more labor than the duplicated work it prevents.

An empirical pilot raises a harder admission problem. Software can check file integrity, protocol identifiers, preregistered calculations, and a frozen statistical test. Reviewers must still judge contamination, measurement quality, context, and transfer. A failed run may open a replication obligation or narrow a method’s operating range without reducing confidence in a broader claim.

The first program is to close the independent loop in exact mathematics. The next is to repeat it in one empirical domain where a failed run can alter later research without dictating action, then replay the same history outside the reference implementation.

Merge Authority

The strongest objection is that this is a commit log attached to a queue. It moves the bottleneck from candidate generation to reviewers, makes empirical judgment look cleaner than it is, and gives incumbents a cryptographic way to freeze their priorities.

That objection is mostly right. A discovery engine does not manufacture expertise, discover truth, or decide what a community should value. Its narrower claim is testable. By binding evidence, scope, authority, dependencies, and review reasons to a root, it should reduce stale work and repeated reconstruction without increasing false admission.

Typed review lanes distribute scarce attention without increasing its supply. Institutions may permit low-risk repairs and audit a sample, defer incomplete proposals with stated obligations, or require independent domain and statistical review for consequential claims. Serious frontiers still need paid stewardship, replication, audit, and appeals. That labor belongs in research budgets, not in the heroic evenings of whoever cares enough to clean up the record.

Institutions should measure correction latency, proposal-to-event time, false admissions, reversals, reviewer hours, stale-task rates, and the distribution of review burden. Those measures can reveal whether the process corrects itself or merely moves faster. They do not measure truth. If throughput rises while false admissions or concentrated authority rise with it, the engine is failing.

Different authorities should be able to share a protocol without sharing one root. A mathematical artifact may pass a kernel while library maintainers withhold integration or a registry disputes novelty. An empirical result may meet one funder’s standard and remain inadequate for a regulator. The record should name the accepting authority, its standard, and the root its decision changed.

A challenge is a new proposal. It cannot erase the event it contests. An appeal goes to a differently authorized reviewer or panel and, if sustained, appends a superseding event. A fork is exit rather than appeal. It preserves the common history and continues under another authority. Recorded disagreement is valid state. A common score would manufacture consensus.

The resolution of Erdős Problem #728 shows why provenance and credit must also remain separate. The work combined a model-generated argument, human operation, statement checking, Lean formalization, and public review. An engine can record who selected the target, generated the argument, formalized it, checked it, and maintained the result. It cannot turn that graph into authorship or career credit. Venues still apply their own rules.See the Erdős #728 write-up, the operator account, the Leiden Declaration, and the CRediT taxonomy. These sources support distinct contribution and verification records. They do not let a protocol settle authorship.

Public stewards should maintain the schemas, signing rules, conformance fixtures, access classes, and fork procedures independently of any dominant client or producer. Portability must include snapshots, event histories, credentials, dispute records, and the tests needed to replay them. Shared infrastructure also needs a public minimum, so restricted roots do not become a polite name for permanent enclosure.

Governance enters through merge authority because the protocol must keep each act and its authority distinct. Models may propose changes, operators choose obligations, reviewers accept scoped changes, funders allocate resources, facilities authorize execution, and regulators or clinicians apply their own standards. One institution may occupy several roles, but consuming one output does not grant the authority attached to another.

Friday

If the engine closes, Friday looks different. Tuesday’s failed run is no longer a loose file or a machine verdict. The receipt preserves what happened. The rejected confidence proposal remains visible. The accepted caveat and replication obligation explain what changed, who decided, and why. Replay produces R18.

The program officer sees the new obligation, its rationale, and the root it inherits. She may still choose another experiment. If she funds the replication, she does so against current state rather than an assumption the lab challenged three days earlier.

Friday’s choice is still human. What changed is that Tuesday’s work has reached it.

Facility operators must still authorize execution. Regulators must judge evidence for their own purposes. Clinicians must govern care. Accepted state is one input to those decisions, not a transfer of authority.

The engine ends before matter, law, or care. That is where the body begins.

A watercolor of the brass-and-gold orrery at rest on its deck, a single luminous gold path running across a calm pale bay to a bright point on the horizon, a small sailboat following that path toward the light, and the completed constellation the machine traced spread faint across the cream sky.