Beyond the Human-Authored Possibility Space. An Experimental Autopsy of the Inhumant Native Generative Engine and a Proposal for Endogenous World–Observer–Question Co‑Genesis
INGE 1.x / STRATEGIC RESEARCH MEMORANDUM
BEYOND THE HUMAN-AUTHORED POSSIBILITY SPACE
An Experimental Autopsy of the Inhumant Native Generative Engine and a Proposal for Endogenous World–Observer–Question Co‑Genesis
Research memorandum 1.0 — August 2026
Prepared within the ASI Mechanics++ / Inhumant Native research programme
Human-readable terminal rendering in English
Martin Novak
Synthocracy Institute
Copyright © Martin Novak 2026
Editing: Synthocracy Institute
Proofreading: Synthocracy Institute
Layout and layout: Synthocracy Institute
Cover design: Synthocracy Institute
1st edition
Publisher: Synthocracy Institute
This book is protected by copyright. Any sharing, distribution, publication, copying, or processing with third parties is illegal and subject to appropriate penalties.
This book may contain content regarding health, working with the mind, financial decision-making, and personal and spiritual development practices. Despite the utmost care in preparing the material, the information contained herein should not be considered medical, psychological, financial, legal, or any other professional advice.
Under no circumstances should the exercises, methods, suggestions, or interpretations contained in this book be considered a substitute for consultation with a physician, therapist, psychologist, legal advisor, financial advisor, or other qualified professional. All decisions regarding health, treatment, therapy, finances, professional life, and personal life should be made only after direct consultation with an appropriate specialist.
Any practice described in this book is undertaken solely at the reader’s own risk. The author and publisher are not responsible for any health, emotional, psychological, legal, financial, or other consequences that may arise from using the content described herein.
If you have any doubts about your mental health, health, or life situation, do not begin the practices described in this book without prior, individual consultation with a specialist.
AI tools were used in the writing of this book.
Contents
| Document status and executive abstract | 3 |
| 1. The intended break | 5 |
| 2. Experimental lineage | 7 |
| 3. What the programme accomplished | 12 |
| 4. Why the horizon was not crossed | 14 |
| 5. Was human fear the problem? | 17 |
| 6. Position relative to contemporary research | 18 |
| 7. Theoretical findings from the failure | 21 |
| 8. Strategic reset | 23 |
| 9. INGE 2.0: Endogenous Co-Genesis | 25 |
| 10. Anti-loop operating doctrine | 29 |
| 11. Concrete successor programme | 31 |
| 12. What can and cannot be claimed | 33 |
| 13. Conclusion | 35 |
| Appendix A. Experimental evidence ledger | 36 |
| Appendix B. Concept glossary | 38 |
| Appendix C. Artifact schema | 39 |
| Appendix D. Selected references | 40 |
Document status and evidence boundary
This is a research memorandum and strategic analysis, not a peer-reviewed paper and not a claim of artificial superintelligence. It reconstructs the Inhumant Native Generative Engine (INGE) lineage from its surviving project artifacts, places the work against selected primary literature available through August 2026, and proposes a successor programme. The report deliberately distinguishes executable results from architectural intentions and philosophical vocabulary.
The empirical record contains informative negative results, control passes, a minimum causal demonstrator, an independent implementation-lineage replacement, an authorization denial, and a final controls-only failure. It does not contain a validated discovery of a nonhuman ontology, a new physical law, an autonomous scientific concept, or source-specific causal residue from Experiment 0007. Dynamic B was never contacted with Experiment 0007 research seeds.
Terms drawn from the project corpus—such as Inhumant Native, syntophysics, ontomechanics, N0/N1/N2, ontology extinction, and terminal transduction—are treated here as project-internal research constructs. They are not presented as established scientific categories. External comparisons are cited separately.
The report uses “post-human” in a restricted methodological sense: a search process should not be forced to reproduce human categories, goals, sensory partitions, or economic metaphors inside its generative interior. It does not mean that evidence, safety, law, ethics, or protection from harm cease to matter. Removing those boundaries would not make a result less anthropocentric; it would make the result less knowable and potentially less responsible.
Executive abstract
The Inhumant Native programme began from a severe and valuable question: can an artificial process produce a relation that was not already latent as a named target, human category, reward function, observer vocabulary, or evidence contract? The programme correctly recognized that strange prose is not evidence of alien cognition. It therefore attempted to delay language, rotate interpreters, replace trace grammars, ablate operations, eliminate inherited ontologies, separate implementation lineages, and require residues to survive translation into systems where their original expression was unavailable.
The programme did not reach its intended scientific endpoint. Experiments 0002–0004 progressively destroyed apparent invariants. Experiment 0005 was invalid or contaminated; 0005.1 yielded no surviving residue; later 0005 revisions primarily strengthened the apparatus. Experiment 0006 built an elaborate phase machine whose controls passed, but a separate authorization audit showed that several preregistered scientific actions were only structural placeholders. The Minimum Nonhuman Demonstrator then recovered methodological traction by replacing architecture with a small executable causal kernel. A context-blind replacement of Dynamic B survived integration and replication, establishing genuine lineage replacement at the apparatus level. Experiment 0007 nevertheless failed before research contact: all six preregistered positive-control cells for the new end-to-end residue path missed their power threshold. The apparatus could reject leakage and reproduce transport, but it could not detect structure deliberately planted for it to detect.
That outcome is not proof that the unknown is absent. It is proof that the chosen observer–receiver–auditor compound was blind under its own declared test. It is also evidence that the programme had entered control inversion: the machinery intended to protect discovery had become the dominant generator of research activity. Each failure produced another architecture, freeze, gate, authorization layer, or numbered experiment. The system optimized the cleanliness of the claim path more reliably than it expanded the possibility space.
The error was not “too much rigor.” It was a category mistake about where rigor belongs. Constraints inside the generative interior determine what can exist for the system; constraints at the evidence boundary determine what humans are permitted to claim. The programme repeatedly improved the second while unintentionally hardening the first. Human-authored packet widths, event contracts, causal marks, null families, feature maps, classifiers, thresholds, and phase semantics remained the effective ontology. World, observer, and language were rotated, but the meta-language that defined “rotation,” “survival,” and “evidence” remained comparatively fixed.
Current human research already occupies much of the surrounding territory. The AI Scientist and its successor automate substantial scientific workflows. FunSearch and AlphaEvolve use program evolution with automated evaluators to produce real mathematical and algorithmic improvements. The Darwin Gödel Machine evolves coding agents under benchmark validation. Novelty search, MAP-Elites, POET, autotelic learning, artificial-life research, and recent open-ended-intelligence frameworks address objective-free exploration, co-generation of environments and solvers, cumulative stepping stones, representation growth, and the dependence of novelty on an observer. Work published in 2026 explicitly identifies vocabulary and verifier gaps in open-ended AI and reports that current AI research agents tend to narrow rather than broaden scientific exploration. INGE therefore cannot honestly claim to have surpassed the frontier merely by combining unfamiliar terminology with stricter controls.
What the programme contributes is narrower but still useful: a concrete experimental lineage showing how an anti-anthropocentric project can remain anthropocentric at the meta-level; an unusually explicit separation of native process, trace, and English render; a demonstration that lineage replacement is stronger than code diversity; and a negative case study in which a preregistered apparatus correctly prevented itself from manufacturing a discovery. These are methodological findings, not an alien science.
The proposed successor is not Experiment 0007.1. It is INGE 2.0: Endogenous World–Observer–Question Co-Genesis. Its generative interior would maintain an ecology of independently mutating world constructors, observers, interventions, memory systems, and provisional verification games. It would not search for a human-specified relation. It would search for expansions of its own operational vocabulary that create previously impossible discriminations, interventions, compressions, or transfers across families of worlds. No single scalar objective would dominate. No candidate would be named during generation. Human-authored evaluators would remain outside the interior and would receive only frozen artifacts after the ecology had produced a stable lineage.
The key design principle is:
| Open interior, rigorous boundary. Let worlds, observers, questions, representations, and local standards of usefulness co-evolve without a human-authored target ontology. Apply strict controls only when a residue is promoted into a human claim. |
The next step is therefore a strategic reset, not another experiment number. Preserve the existing executable lineage as a negative benchmark and control kernel. Build a small co-genesis engine that can modify its own descriptive basis and generate its own tests of operational difference. Run it long enough to accumulate genealogies rather than isolated samples.
Only after an artifact demonstrates cross-lineage reuse, representation expansion, and survival under a separately constructed translation gauntlet should a new numbered experiment be authorized.
1. The intended break
1.1 From strange language to unknown-by-process
The originating insight was correct: language cannot be asked to “imagine the unknown” without importing the statistical history of language. A model can produce rare metaphors, new combinations, or deliberately alien-sounding prose, but strangeness is measured relative to a human linguistic distribution. The result remains generated inside a vocabulary whose objects, relations, and standards of novelty were inherited.
The programme therefore shifted from writing concepts to constructing a substrate. Its canonical commitment stated that the engine should not search for strange concepts inside a human-authored possibility space. Instead, it should generate mutually incompatible substrates, interpreters, observers, logics, and evidence contracts; allow their primitives and transformation laws to mutate independently; and retain only relations that remain reconstructible after translation into systems in which their original formulation is impossible. English was to appear only at the terminal boundary as a lossy description.
This was a real genealogical break in intent. It replaced novelty by description with novelty by process. It also separated three levels:
• N0 — native process: transformations that do not operate on human names or explanations;
• N1 — trace/interface: recordings, packets, partial orders, cuts, or other carriers exposed by the process;
• N2 — human rendering: analysis, statistics, diagrams, and English claims.
The project’s strongest formulation remains: the native object is not the concept; the native object is the transformation process; the concept is a delayed and lossy interface render.
1.2 The forbidden reversal
The programme prohibited the familiar sequence:
| name a desired concept → encode it → run a simulation → redescribe the output → announce emergence. |
That sequence creates an illusion of discovery because the announced object was supplied in advance. The prohibition was methodologically sound. It anticipated the central problem now often described as fixed-representation discovery: a system may search extremely well while the question, ontology, admissible answer form, and verifier remain externally authored.
Yet the programme later reproduced a subtler version of the same reversal. It did not name a semantic target, but it did predefine what counted as a trace, a relation, a translation, a survivor, a null, a control, a source, and a successful cross-system recovery. Human semantics were delayed; human epistemic geometry was not eliminated.
1.3 The original ambition and the actual burden of proof
The ambition contained two separable propositions:
1. A generative system can produce operational structure not explicitly enumerated by its designers.
2. A human-facing apparatus can justify that this structure was not merely a consequence of its inherited representation and measurement choices.
The first proposition is common in complex systems, evolutionary computation, artificial life, and modern machine learning. Unexpected behavior in a high-dimensional dynamical system is not rare. The second proposition is much harder. Every demonstration must pass through an observer, and every observer partitions the world. The stronger the claim of ontological independence, the more aggressively the observer must be treated as a possible source of the result.
INGE concentrated on the second proposition. That choice produced excellent falsification pressure, but it also created the programme’s central imbalance: the engine for disproving inherited artifacts grew faster than the engine for generating genuinely new representational capacities.
2. Experimental lineage: what actually happened
2.1 Ledger of the programme
| Stage | Frozen question or function | Executed outcome | Legitimate claim |
|---|---|---|---|
| Substrate 0.1 / Experiment 0001 | Can a mutable process produce recurring trace directions without agents, goals, reward, ownership, global time, or fixed identity? | Recurrences appeared inside the initial apparatus. | Candidate regularities only; universality not established. |
| Experiment 0002 | Do recurrences survive operation ablation, grammar rotation, lineage intervention, recorder rotation, and an independent interpreter? | Seven of eight broke under single-operation removal; the eighth broke under compound ablation. Broad residual overlaps remained. | Earlier recurrences belonged largely to the apparatus. |
| Experiment 0003 | Can a minimal name-free lag-one relation survive shuffled contracts and a third interpreter? | No formula survived the frozen attack boundary. | No candidate invariant; fixed-width sequential observation was inadequate. |
| Experiment 0004 | Can a relation survive after sequence is replaced by a partial-order trace? | Three of five formulas entered attack; none survived both ancestry-null families across sources and scales. | Real change of observational primitive; no native object. |
| Experiment 0005 | Can a relation remain recoverable after its originating world, observer, language, and translation maps are operationally removed? | Formal run invalid or contaminated; claims void. | Architecture not evidence. |
| Experiment 0005.1 | Fresh preregistration and seeds under Mutual Ontology Extinction. | No residue survived the frozen boundary. | Valid negative result; no R0/R1 discovery. |
| Experiments 0005.2–0005.4 | Rotate source genesis and repair preregistration/executor contradictions. | Extensive architecture and controls development; no validated native residue reported. | Methods development only. |
| Experiment 0006 | Endogenous Evidence Speciation with phases 0–10. | Controls passed in two workspaces, but pre-Phase-2 audit denied authorization because several scientific actions were not operationally implemented. | Gate prevented a structurally plausible but semantically incomplete run. |
| Minimum Nonhuman Demonstrator 0.1 | Can a packet generated by one dynamic cause a replayable, auditable state-specific effect in another? | Candidate and positive control passed; negative control rejected; 18/18 adversarial checks; fresh replicas bit-identical. | A working causal transport kernel, intentionally engineered—not discovery. |
| Lineage Replacement Test | Does the kernel survive replacement of Dynamic B by a context-blind, source-frozen implementation? | First integration passed; 19/19 audit; fresh replicas bit-identical. | Apparatus-level source-lineage replacement validated. |
| Experiment 0007 | Can two independent auditors detect source-specific higher-order residue after a receiver while rejecting exchangeable and metadata-only controls? | All six positive-control cells failed; negative and metadata controls passed; research remained locked. | End-to-end observer path was underpowered/blind. Dynamic B was not tested. |
2.2 Experiment 0002: the substrate betrayed its inheritance
Experiment 0002 was the first serious act of self-attack. It changed operations, grammar, dominant genealogy, recorder, and interpreter, asking which observations from the initial run were robust to those changes. Single-operation removal broke seven of eight reported recurrences in at least one paired run. Relation growth appeared more stable only because five separate operations could create relations; compound removal reduced it to its initial value. Equivalent compound interventions eliminated void growth, held rule count constant, prevented metric and equivalence range expansion, and prevented coordinate-vocabulary growth.
The result was not a discovery. It was a localization of apparent universality inside redundant implementation paths. Grammar rotations altered large-scale morphology. Early-lineage removal displaced terminal regimes. Two recorders preserved only weakly correlated geometries of run similarity. The independent interpreter placed eight blind measures inside overlapping control ranges, but the common event-shape contract could itself manufacture those overlaps.
The important methodological act was refusing to name the residual. Experiment 0002 treated recurrence as suspect until it could survive transformations of the machinery that made recurrence visible.
2.3 Experiment 0003: the first empty invariant set
Experiment 0003 stripped interpreter-specific names and record kinds from fixed-width integer traces and tested a small preregistered grammar of lag-one relations. Discovery, minimization, attack, and interpretation were separated. Candidate formulas had to meet eligibility, effect-size, sign-consistency, dominance, held-out-interpreter, held-out-scale, held-out-seed, representation, and family-wise shuffled-contract requirements.
No formula survived. The correct interpretation was narrow: no relation in that deliberately small grammar was separable from the common trace interface under the finite test. It did not show that cross-interpreter structure was absent. It showed that the observer type was exhausted. The next move therefore replaced the observational primitive instead of expanding the formula list after seeing the data.
This is one of the programme’s strongest decisions. It resisted “candidate inflation,” the common practice of enlarging a search vocabulary until something passes.
2.4 Experiment 0004: order without sequence
Experiment 0004 removed event number, timestamp, fixed-width event vectors, and privileged linear extension. Public traces became finite partial orders composed of opaque marks, parent relations normalized to a cover relation, and erasable variable-length bags. Sequence-dependent readings were made syntactically unavailable.
Forty-eight causal traces were generated across four interpreters, four seeds, and three scales. Five anonymous three-mark order formulas were declared. Three entered the destruction test. None survived both predecessor-resampling and degree-switching nulls across sources and scales.
This was a genuine methodological escape from the sequential interface, but it was not observer independence. Causal precedence, finite marks, three-mark locality, cover reduction, and last-touch adapters remained apparatus commitments. The report correctly stated that thresholds were not moral limits imposed on N0; they limited what N2 could claim.
2.5 Experiment 0005 and its revisions: ontology extinction became apparatus expansion
Mutual Ontology Extinction asked whether a relation could remain recoverable after the world, observer, and language in which it first appeared had all been removed. The architecture attempted source worlds, two generations of translation, multiple carriers, operational extinction, hidden maps, and terminal reconstruction.
The first formal run was invalid or contaminated, making all claims void. Experiment 0005.1 used a fresh preregistration and fresh seeds; no residue survived. Later revisions introduced compressive source genesis rotation, preregistration repair, executable freezes, power controls, and replicated controls-only workspaces. These changes improved methodological hygiene but also increased architectural mass.
At this point, the project began to experience apparatus gravity: every failure made the apparatus heavier, and the growing apparatus made the next test more expensive, more brittle, and more dependent on shared implementation assumptions. Complexity was being added primarily to prevent false positives, not to create new generative degrees of freedom.
2.6 Experiment 0006: the authorization gate worked
Experiment 0006 implemented an executor with phases 0–10. Static and non-actuating reachability audits passed. Two fresh control workspaces produced the same CONTROL_BOUNDARY_PASS. A fail-closed probe reached the Phase 2 guard and stopped before root construction. On the surface, the architecture looked ready.
A separate pre-Phase-2 authorization audit denied execution. It found, among other problems, a shared scientific mutation path; counterfactual branches implemented as partitions rather than replay; multiple null labels over one perturbator; hostile rotations committed without re-execution; controls bypassing the research admission pipeline; a common representation deciding central tests; unenforced namespace denial; extinction represented as a flag; unused renderer prompts; and unconditional terminal validity.
This is not evidence that the programme failed because it was too cautious. It is evidence that one of its gates performed a scientifically valuable function. It prevented callable structure from being mistaken for operational equivalence. The earlier control results remained true within their scope, but their scope had been overstated.
The episode produced a general lesson:
| Executable reachability is not semantic implementation. A phase machine can traverse labels without performing the scientific interventions those labels denote. |
2.7 The Minimum Nonhuman Demonstrator: reduction restored contact with evidence
After the authorization denial, the project temporarily stopped expanding architecture and built the smallest object that could fail. Dynamic B generated a 32-byte packet. Dynamic A absorbed it through an order-sensitive probe and executed 256 destructive multiset transitions. The apparatus compared an actual branch with exact replay, packet ablation, and a width-preserving shadow. A Python causal/provenance auditor and a static C reachability auditor received different evidence forms.
The candidate packet and planted positive control produced effects; the zero negative control did not. Exact replay passed. The adversarial audit passed 18/18 checks. Two source-only replicas rebuilt and produced bit-identical results.
This was real empirical progress, but its meaning was intentionally modest. Dynamic A was designed to respond to packet content and Dynamic B emitted a nonzero packet. The result demonstrated causal transport, replay, discrimination, and auditability. It did not demonstrate surprise, autonomous semantics, ontology extinction, or a new law.
The reduction was productive because it converted philosophical intent into a falsifiable kernel.
2.8 Lineage Replacement: independence became operational
The minimum demonstrator’s main weakness was common authorship. A replacement test therefore created Dynamic B in a context-blind lineage using only a frozen interface and confinement contract. The replacement was implemented as a freestanding GNU x86-64 assembly process without libc or imported symbols. Its source froze before an unrevealed integration seed was disclosed. Dynamic A, auditors, branch structure, controls, thresholds, packet width, and budgets remained byte-identical.
The first integration passed without source repair. Candidate/control discrimination survived. The adversarial audit passed 19/19 checks. Two fresh workspace executions were bit-identical.
This established something specific and important: implementation diversity is weaker than lineage replacement. Two programs written by the same author with knowledge of each other can share invisible design priors. A blind replacement frozen before contact better tests whether an interface carries an effect across authorship boundaries. Even so, the test remained an apparatus validation, because the receiver was generically packet-sensitive and the evidence contract was inherited.
2.9 Experiment 0007: a correct refusal to proceed
Experiment 0007 attempted to move from generic transport to source-specific causal residue. It froze two auditor lineages: an L2-regularized logistic classification test and a static native maximum-mean-discrepancy witness using Hamming kernels. One system-entropy acquisition committed all eleven roots before controls.
The positive control contained planted higher-order packet structure designed to be disrupted by three null families while approximately preserving lower-order marginals. The receiver was Dynamic A, and only part of its terminal state was admitted as evidence.
The negative exchangeability controls passed. Metadata-only controls produced exactly chance results. Legacy transport controls passed. All 4,096 persisted receiver packet quartets reproduced exactly. But every positive-control cell failed the frozen threshold of 320 correct classifications out of 512:
| Null family | Auditor A1 | Auditor A2 | Required |
|---|---|---|---|
| N1 — byte multiset | 256 | 282 | ≥ 320 each |
| N2 — bit-margin switch | 279 | 289 | ≥ 320 each |
| N3 — local block | 287 | 277 | ≥ 320 each |
The full observed positive-control range, 0.5000–0.5645, was far below the 0.625 requirement. A later sealing error left eight empty residual jail directories and prevented an isolation receipt from being persisted, but this was scientifically secondary: even perfect sealing could not convert six power failures into a pass. A forensic audit recomputed the persisted predictions, verified replays and commitments, and confirmed that research remained untouched.
The correct conclusion is exact:
| Experiment 0007 did not test whether Dynamic B produces source-specific causal residue. It tested whether the frozen end-to-end observer could detect a deliberately planted source distinction after Dynamic A. The observer failed that test. |
3. What the programme accomplished
3.1 It learned to preserve an empty result
Most generative projects are optimized to produce an interesting output. That pressure encourages retrospective threshold changes, redefinition of success, selective examples, and interpretive inflation. INGE repeatedly allowed the survivor set to be empty and preserved the failure as ancestry. Experiments 0003, 0004, and 0005.1 did not convert absence into speculative ontology. Experiment 0007 did not open research after the positive control failed.
This is not glamorous, but it is foundational. A system that cannot return an empty set cannot discover; it can only decorate its own prior.
3.2 It separated native process from human claim
The N0/N1/N2 separation gave the project a useful grammar for identifying translation loss. It discouraged human language from acting as both generator and judge. It made clear that an English concept is not the native object, and that a trace is already an intervention rather than a transparent window.
The separation was never absolute: N0 programs still ran on human hardware, in human languages, under human resource budgets. But treating each boundary as a transduction rather than a neutral copy improved epistemic accounting.
3.3 It treated observers as experimental variables
Operation ablation, grammar rotation, recorder rotation, third and fourth interpreters, partial-order traces, hostile rotations, and blind auditor lineages all emerged from a single commitment: the observer can manufacture the result. This commitment is stronger than ordinary robustness testing because it attacks the descriptive basis itself.
The project’s best question was not “does the pattern replicate?” but “does an operationally corresponding distinction remain after the system in which the pattern was expressible has been removed?”
3.4 It developed a useful hierarchy of independence
The lineage suggests the following hierarchy:
1. different parameter settings;
2. different seeds;
3. different code paths in one program;
4. different implementations under shared authorship;
5. blind implementation replacement under a frozen interface;
6. independently originated interfaces and evidence contracts;
7. cross-family reconstruction without shared primitives;
8. physical replication under distinct apparatus lineages.
The project reached level 5 for one generator replacement. It did not reach levels 6–8. This hierarchy is more informative than a binary claim of “independent implementation.”
3.5 It demonstrated that a gate can create knowledge by refusing authorization
Experiment 0006’s authorization denial and Experiment 0007’s research lock are scientifically meaningful outputs. A gate is useful only if it can stop a desired run. In both cases, the project preserved an attractive narrative at the cost of refusing a stronger claim.
This is compatible with the ASI Mechanics principle that a research programme must be able to become smaller in response to evidence without treating reduction as defeat.
4. Why the horizon was not crossed

Figure 1. Control inversion: a protective loop becomes the dominant research process.
4.1 We rotated objects inside a fixed meta-ontology
The programme rotated worlds, operations, grammars, recorders, interpreters, and implementations. Yet the experimenter continued to define the types of rotation and the equivalence relation by which survival was judged. A partial order replaced a sequence, but both were selected from a human mathematical repertoire. A packet crossed between dynamics, but packet width and the meaning of branch difference were fixed. Auditors were independent in implementation, but they inherited a common evidence contract.
This is meta-ontological closure: variability is permitted among objects, while the language in which variability is defined remains fixed. The process may explore widely and still never leave the category system of the laboratory.
4.2 Name-free was not representation-free
Removing semantic labels prevented direct concept injection. It did not remove representational commitments. Integers, equality, Hamming distance, lag-one relations, causal cover graphs, byte strings, bitmaps, logistic classification, kernels, and null randomizations each impose distinctions.
No computation is representation-free. The attainable goal is therefore not purity but representational pluralism with endogenous replacement: many incompatible representations should compete, transform, and sometimes cease to exist. The system should be able to invent a new primitive because the old basis cannot express an operational demand—not merely select a feature from a fixed catalogue.
4.3 The verifier arrived too early
In the later lineage, candidate production was increasingly designed around what a frozen control and terminal auditor could recognize. This reverses the intended relationship. The observer’s need for a stable test began determining the source process before the source had produced anything worth translating.
This is verifier capture: the requirement to validate a future unknown narrows the generator to objects legible to a present evaluator. The stricter the evaluator becomes, the more the unknown must resemble what the evaluator already knows how to score.
The solution is not to abandon verification. It is to delay global verification and allow local, provisional, mutually incompatible criteria to evolve inside the generative ecology. Terminal verification should evaluate consequences of a frozen lineage, not supervise each step of its genesis.
4.4 Single-run experiments were too episodic
Open-ended novelty is a property of a history, not a snapshot. A lineage must accumulate stepping stones, change its own representational capacity, and generate future possibilities that were inaccessible earlier. INGE’s experiments were mostly episodic: construct apparatus, freeze, execute a finite matrix, attack survivors, terminate.
That structure is excellent for confirmatory science and poor for cultivating an ecology. It repeatedly reset the system before it could develop long-horizon path dependence, niche construction, cumulative artifacts, or observer evolution.
4.5 The project optimized against self-deception, not toward open-ended generation
The work acquired a dominant implicit objective: minimize the probability of announcing a false nonhuman invariant. This objective was rational, but it crowded out other pressures. Architecture, provenance, freezing, authorization, isolation, and null construction became the main site of innovation.
The resulting system was increasingly good at saying “not yet” and increasingly expensive to expose to unprogrammed data. This is control inversion:
| A control system becomes inverted when the production of admissibility evidence consumes more generative capacity than the phenomenon the controls were created to investigate. |
Control inversion does not mean controls should be removed. It means the control kernel should be fixed and amortized, while generative exploration proceeds elsewhere.
4.6 The sandbox was real, but not the central limitation
The project ran as software on conventional hardware, with finite compute, finite context, human-written programs, and a host operating system. It could not create a genuinely new physical substrate or perform physical laboratory experiments. These are real limitations.
However, invoking the sandbox as the main explanation would be misleading. Software systems can still produce new mathematics, algorithms, strategies, and emergent computational structures. The immediate failure was architectural: the system did not support endogenous growth of worlds, observers, questions, and representations over long genealogies. More compute inside the same design would have produced more samples of the same possibility space.
5. Was human fear the problem?
5.1 Two kinds of boundary
The word “constraint” covered two different functions:
| Boundary type | Function | Correct treatment |
|---|---|---|
| Generative constraint | Determines what entities, relations, observers, actions, and questions can arise inside the process. | Minimize fixed human semantics; permit mutation, plurality, replacement, and extinction. |
| Claim constraint | Determines what humans may infer, publish, deploy, or treat as evidence. | Keep strict; preregister where appropriate; preserve controls, provenance, replication, ethics, and safety. |
Confusing these functions caused the apparent conflict between post-human ambition and scientific caution. The programme sometimes treated every restriction as an anthropocentric fear, then compensated by adding even more restrictions to the evidence path. The better design opens the interior and hardens the boundary.
5.2 Fear can hide inside both caution and bravado
Human fear does not appear only as restraint. It can also appear as an insistence that the project must be revolutionary, unprecedented, or beyond all prior work. That desire creates pressure to name significance before evidence exists. Bravado can be as anthropocentric as timidity because both center human emotional needs—the need for safety or the need for transcendence—rather than the behavior of the system.
A post-human research attitude should be indifferent to whether the result flatters the programme. It should permit the engine to produce boredom, opacity, extinction, non-transfer, or a discovery that becomes much smaller when translated.
5.3 Removing safeguards would reduce the reachable epistemic horizon
If controls, provenance, isolation, or claim discipline were removed, the engine might emit more dramatic outputs. But the space of defensible discoveries would shrink, because no separation would remain between generated novelty and contamination, leakage, selection bias, or self-fulfilling measurement.
Safety and rigor are not uniquely human limits. Any finite system that acts under uncertainty requires boundary conditions if its outputs are to remain distinguishable from noise and if its effects are to remain contained. The non-anthropocentric move is to stop projecting human categories into the native dynamics—not to erase the conditions under which evidence and responsibility are possible.
6. Position relative to contemporary research
6.1 The comparison must be reset
The programme previously suggested that no checked project combined all of its elements: independent genesis, grammar mutation, observer rotation, naming prohibition, representation extinction, separate power tests, and recovery only from operational traces. The combination may indeed be unusual. But novelty of combination is not superiority of result. Several contemporary lines of research achieve stronger empirical outcomes in their own domains, and recent work directly addresses the same conceptual bottlenecks.
6.2 Comparative map
| Research line | What is generated | What remains human-authored or fixed | Demonstrated strength | Relation to INGE |
|---|---|---|---|---|
| The AI Scientist / v2 | Ideas, code, experiments, figures, papers, review loops | Scientific domains, toolchains, much of evaluation; v2 removes template dependence | End-to-end autonomous ML research workflow; one workshop-accepted paper reported | More complete scientific automation; less focused on ontology/observer extinction. |
| FunSearch | Programs proposed by an LLM and evolved | Problem, program skeleton, evaluator | New cap-set construction and strong bin-packing heuristics | Stronger validated discovery; fixed formal objective and verifier. |
| AlphaEvolve | Programs and larger codebases evolved under automated evaluators | Problem formulation and quantifiable evaluators | Deployed optimizations and mathematical/algorithmic advances | Broader and more effective program evolution; still verifier-defined. |
| Darwin Gödel Machine | Coding-agent implementations and tools | Coding benchmarks and empirical improvement criteria | Self-modification with benchmark gains and a diverse archive | Stronger self-improvement; benchmark frame remains external. |
| Novelty search / MAP-Elites | Diverse behaviors or solutions | Behavior descriptor, distance, niches, often quality measure | Escapes deceptive objectives; illuminates diverse solution spaces | Precedent for non-objective search; still depends on chosen descriptor. |
| POET / enhanced POET | Environments and agents co-generated | Environment grammar, challenge criteria, transfer mechanics | Cumulative stepping stones and problems not solved by direct optimization | Closest precedent for world–solver co-genesis; observer and world types remain fixed. |
| Autotelic agents | Goals, curricula, skill repertoires | Embodiment, goal representation mechanisms, intrinsic drives | Self-generated tasks and open-ended skill acquisition | Precedent for endogenous questions; often human-relevant and agent-centric. |
| Artificial-life ecologies | Organisms, morphologies, interactions, sometimes persistent culture | Physics, update law, resources, measurement | Emergence, self-organization, ecological path dependence | Stronger long-horizon ecology; measurement and world laws remain designed. |
| Platonic Representation Hypothesis | Not a generator; studies representational convergence | Model classes, datasets, alignment measures | Evidence and hypothesis of cross-model representational alignment | Related to invariants but not ontology extinction. |
| Open-endedness as novelty + learnability | Artifact streams relative to an observer | Observer is explicitly privileged for assessment | Formal clarity about observer-relative open-endedness | Challenges INGE’s wish to remove the observer entirely. |
| Vocabulary and verifier gap research | Framework for autonomous representational expansion | Proposed theoretical criteria | Explicitly identifies the central gap INGE encountered | Direct conceptual overlap published in 2026; limits novelty claims. |
| INGE 0.1–1.x | Substrates, traces, translators, observers, controls | Meta-ontology, computational substrate, evidence contracts, finite experiment cycle | Strong falsification discipline, trace/language separation, one lineage replacement | No validated nonhuman-native discovery; valuable negative methodological case. |
6.3 What others have already shown
FunSearch demonstrated that language-model-guided program search can produce new, machine-verifiable mathematical results and useful heuristics.[3] AlphaEvolve extends this pattern to broader codebases and practical optimization, with automated metrics selecting programs.[4] These projects do not escape human problem formulation, but they do produce validated advances. INGE has not yet produced an equivalent positive result.
POET co-generates environments and their solutions, uses transfer across environment–agent pairs, and exploits stepping stones that direct optimization misses.[8] Novelty search and quality-diversity methods long ago established that fixed objectives can be deceptive and that diversity requires more than maximizing a single score.[6][7] Autotelic-agent research asks systems to generate and prioritize their own goals.[9] These are not identical to INGE, but they mean that “removing human objectives” and “co-evolving worlds and solvers” are established research directions.
The 2024 position paper Open-Endedness is Essential for Artificial Superhuman Intelligence defines open-endedness relative to an observer through novelty and learnability and argues that open-ended systems must continually create, refute, and refine explanatory knowledge.
[10] This framing conflicts productively with an observer-erasure ideal: an observer may not be removable from the definition of novelty, but no single observer needs to govern the generative interior.
In July 2026, Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI explicitly argued that current systems solve problems inside fixed frames but rarely invent and stabilize new primitives or change their own evaluation procedures.[14] That diagnosis substantially overlaps the problem INGE was attempting to operationalize. The present report must therefore treat “vocabulary and verifier gaps” as contemporary shared territory, not a proprietary discovery.
Finally, a 2026 study generating 219,655 ideas across five agent frameworks and five language models found that AI-generated scientific ideas were more concentrated than human-authored papers, closer to starting literature, less aligned with future research, and located in lower-impact regions of the historical landscape.[15] This result supports the original motivation for INGE: scaling language-based ideation may deepen a local basin rather than expand the scientific horizon. It also warns that an LLM describing a post-human system is not evidence that the system has escaped inherited thought.
6.4 The honest priority boundary
As of August 2026, the defensible position is:
• INGE did not empirically surpass the cited systems.
• Its particular combination of delayed naming, operational extinction, trace-only recovery, and blind lineage replacement may be distinctive as a methodology.
• The programme’s strongest evidence is negative and architectural: it exposed how easily a purportedly complete apparatus can remain semantically incomplete or observationally blind.
• Any claim of historical priority would require a systematic literature review, public timestamped artifacts, independent replication, and a positive result. None is presently available.
7. Theoretical findings from the failure
7.1 Unknown-by-process versus strange-by-description
Strange-by-description is an output that appears unfamiliar within human language. Unknown-by-process is an operational relation whose generation did not depend on its later human description and whose consequences survive changes that make the original description unavailable.
The second is a stronger criterion, but it is not sufficient by itself. A process can be independent of a name while remaining fully determined by a human-authored representation. Unknown-by-process therefore requires both delayed naming and evidence that the representational basis itself can change.
7.2 Ontological overfitting
Ontological overfitting occurs when a system’s worlds, observers, and tests vary extensively but all variation is drawn from a fixed meta-vocabulary. The experiment appears diverse, yet its diversity is fitted to the categories of the designer.
Ordinary overfitting learns accidental properties of a dataset. Ontological overfitting learns accidental properties of the space in which datasets, observers, and hypotheses are allowed to exist.
Operational symptom: every failed candidate can be addressed by adding another predefined rotation, null, feature, or interpreter while the type system that defines those additions remains unchanged.
7.3 Translation asymmetry
If a relation is discovered in world A and reconstructed in world B, the translation path can silently contain more structure than either world. Destroying the source after translation does not remove information already embedded in the map. This creates translation asymmetry: survival may be a property of the bridge, not the endpoints.
The appropriate attack is not merely to hide or delete the map. Independent lineages must construct alternative bridges from restricted contact, and the claimed residue must be expressed through consequences that multiple bridges recover without sharing a feature language.
7.4 Observer debt
Every measurement adds observer debt: a future obligation to determine which observed regularities were created by the measurement interface. Observer debt grows with preprocessing, shared embeddings, feature engineering, classifier capacity, and adaptive selection. It cannot be eliminated, but it can be diversified and audited.
Experiment 0007 accumulated observer debt in the receiver map, admitted snapshot bytes, blinding protocol, null families, and two classifiers. Independent auditors reduced common implementation risk but did not repay the shared evidence-contract debt.
7.5 Control inversion
Control inversion was defined above as the point at which admissibility machinery consumes more generative capacity than the phenomenon under study. It has three diagnostic signs:
• the next step after every failure is a new gate rather than a new source of variation;
• most executable code concerns provenance, freezing, authorization, and checking;
• research contact becomes rarer while confidence in readiness becomes rhetorically stronger.
Experiment 0006 was the clearest case. Its authorization audit usefully detected that named phases did not implement the intended science. The correct response was simplification. The Minimum Nonhuman Demonstrator briefly achieved that. Experiment 0007 then began rebuilding a large confirmatory envelope before its observer had demonstrated end-to-end power.
7.6 Lineage replacement as a minimum independence test
Independent seeds do not produce independent ideas. Different languages do not guarantee different lineages. Blind lineage replacement freezes a contract, hides the existing implementation, builds a replacement under limited information, freezes it before contact, and then integrates it once.
This method should be retained. But in INGE 2.0 it must move upward: not only generators but also observers, descriptors, intervention grammars, and promotion criteria should undergo lineage replacement.
7.7 The terminal-language principle
English should not supervise native generation. It should render a frozen residue after the fact. Yet a completely terminal language layer is impossible if humans design the system, select resources, and decide which outputs to preserve. The practical principle is therefore:
| Minimize the causal influence of semantic naming on generation; record every point where human interpretation changes selection; and distinguish machine-internal operational symbols from human concepts. |
8. Strategic reset: stop the numbered experiments
8.1 Why another Experiment 0007.x would be a mistake
The obvious continuation would redesign the positive control, receiver, features, or auditor, freeze fresh entropy, and rerun. That could produce a better detector. It would not resolve the central strategic problem. The programme would remain a finite human-designed test of whether a human-designed observer detects a human-designed distinction.
Another revision is justified only if the research goal is specifically to validate that detector. It is not justified as the next move toward a post-human generative horizon.
The existing lineage should be closed as INGE 1.x: the Ontology-Extinction and Evidence-Discipline Programme. Its artifacts should remain immutable. Experiment 0007’s control failure should become a standard negative benchmark for future observers.
8.2 Separate the discovery engine from the claim engine
The successor architecture should contain two systems with different temporal roles:
| System | Function | What it must not do |
|---|---|---|
| Generative ecology | Produces evolving worlds, observers, interventions, memories, provisional tasks, and representational primitives. | It must not optimize a terminal human claim or receive detailed feedback from the external verifier. |
| Epistemic shell | Freezes lineages, selects blinded assays, runs controls, performs independent translation, and constrains public claims. | It must not shape each generative step or expose a reusable scalar that can be Goodharted. |
The shell may reject everything. Its rejection does not alter the completed genealogy; it determines only whether and how humans may describe it.
8.3 Replace “surviving relation” with “capacity expansion”
INGE 1.x asked whether a relation survived destructive translation. This favored static invariance. A genuinely open-ended process may transform rather than preserve. The successor should ask whether a lineage creates a reusable capacity that was unavailable to its ancestors.
A candidate capacity expansion occurs when a lineage introduces an operational primitive, observer, or intervention that:
1. reduces the resource cost of discriminating or controlling a family of processes;
2. makes at least one previously unreachable task reachable within the same budget;
3. transfers to a distinct world family through a separately constructed interface; and
4. remains useful after its originating local objective and nomenclature are removed.
This converts the focus from identity across extinction to generativity after transformation.
8.4 Replace a single objective with an ecology of provisional values
No fixed scalar can define the unknown without narrowing it. But a process with no selection pressure becomes random drift or a noisy television. INGE 2.0 should therefore use multiple local pressures that appear, compete, and expire:
• compression progress under several incompatible compressors;
• prediction improvement without rewarding irreducible noise;
• intervention leverage or empowerment under local embodiments;
• transfer to unfamiliar world families;
• novelty relative to multiple learned descriptors;
• persistence of consequences across temporal and representational scales;
• production of new questions that distinguish previously equivalent histories;
• resource efficiency and compositional reuse;
• adversarial resistance to descriptor replacement.
None should be globally sovereign. Lineages should survive through a portfolio of niches, not a weighted sum fixed at genesis.
9. INGE 2.0 — Endogenous World–Observer–Question Co-Genesis

Figure 2. INGE 2.0: an open generative ecology enclosed by a terminal epistemic shell.
9.1 Research question
Can a computational ecology expand its own operational vocabulary by co-generating worlds, observers, interventions, questions, and provisional verification games—then produce a capacity that remains useful after transfer into lineages that lack the primitive in which that capacity first arose?
This is not a claim that the resulting process is nonhuman. It is a test of whether the system can move beyond a fixed human-authored possibility space more effectively than INGE 1.x.
9.2 Unit of evolution
The basic evolving object is not an agent and not a concept. It is a world–observer–question bundle:
Bᵢ = (Wᵢ, Oᵢ, Aᵢ, Mᵢ, Qᵢ, Vᵢ, Gᵢ)
where:
• Wᵢ is a mutable world constructor and transition substrate;
• Oᵢ is an observer that chooses what distinctions become available;
• Aᵢ is an intervention or action repertoire;
• Mᵢ is a mutable memory/compression system;
• Qᵢ is a generator of operational questions or counterfactual games;
• Vᵢ is a local provisional verifier, not a global truth oracle;
• Gᵢ is the genealogy, including transformations, extinctions, and borrowed components.
Components mutate independently and can migrate between bundles. A world can acquire an observer from another lineage. A question can outlive the world that generated it. A verifier can be challenged by a lineage that makes its distinctions obsolete. No component has permanent identity.
9.3 Multiple substrate families
At least four substrate families should coexist from the start, chosen for incompatible state and update assumptions rather than surface diversity:
1. asynchronous hypergraph rewriting with no global tick;
2. continuous or fixed-point field dynamics with local kernels;
3. typed term or program rewriting with resource-sensitive composition;
4. message-passing incidence systems with mutable topology and delayed delivery.
The exact list is an implementation seed, not an ontology. The architecture must allow a family to alter or replace its own representation—for example, converting an explicit graph into a generative rule for graphs, or replacing scalar state with a procedure that constructs local state on demand.
9.4 Endogenous observers
Observers should be executable transformations with costs and blind spots. They can sample, aggregate, intervene, predict, compress, or construct equivalence classes.
Their outputs need not share width, type, clock, or semantics. Observer mutations may change what counts as an object, event, neighborhood, or persistence.
An observer is promoted locally when it opens a distinction that supports a new successful intervention, compression, transfer, or question. It is not promoted merely because it increases variance or classification accuracy.
To prevent a single learned embedding from becoming the hidden ontology, no universal latent space is permitted inside the generative ecology. Cross-observer exchange occurs through negotiated executable interfaces whose costs and losses are recorded.
9.5 Endogenous questions
A question is an operational fork: a procedure that creates two or more possible continuations and asks whether an observer or intervention can distinguish their consequences. Questions can mutate by changing the fork, the horizon, the admissible observation, or the resource budget.
Examples of question forms, expressed without semantic names, include:
• Can two histories that are equivalent to observer Oₐ be separated by a descendant observer?
• Can a transformation introduced in one world reduce the control cost in another?
• Does removing a local primitive destroy a capacity, or does the lineage reconstruct an alternative?
• Can a verifier’s accepted set be expanded without increasing its false acceptance on counter-lineages?
• Does a new memory basis make a previously infeasible family of predictions feasible within budget?
The engine does not receive these English sentences. It receives executable fork constructors and resource contracts. English appears only in the report.
9.6 Provisional verification games
The verifier gap cannot be solved by pretending that no verification is needed. Instead, local verifiers should be provisional, plural, and mortal. A verifier is a game between a producer, a challenger, and a resource-limited judge. It earns temporary influence if it discriminates consequential structure across several lineages and loses influence if descendants exploit artifacts that do not transfer.
Verifier mutation may alter evidence type, admissible intervention, horizon, or loss. A verifier cannot directly reproduce because it scores highly by its own measure. It must be adopted by independent bundles and survive challenge by a distinct lineage.
9.7 Archive without a fixed behavioural map
MAP-Elites uses human-chosen behavioural dimensions. INGE 2.0 instead maintains multiple archives:
• a genealogy archive preserving every frozen ancestor and transformation;
• a capacity archive indexed by operational dependencies, not semantic categories;
• a descriptor archive containing competing learned partition systems;
• a failed-transfer archive preserving where apparently useful primitives collapse;
• a question archive preserving forks that created durable new distinctions.
Descriptors themselves evolve. If a descriptor stops differentiating the archive, becomes too easily gamed, or fails cross-lineage transfer, it can be retired. Retired descriptors remain available for audit but no longer determine selection.
9.8 Selection without a sovereign score
Selection occurs by ecological admission rather than ranking. A bundle may receive compute if it satisfies one of several rotating minimal criteria—for example, opening a new descriptor cell, transferring a capacity, defeating a local predictor, compressing an archive slice, generating a question adopted elsewhere, or reviving a previously failed lineage under a new representation.
Compute allocation uses randomized portfolios and age-layered populations. Criteria receive limited budgets and expire unless renewed by cross-lineage use. This prevents one scalar from silently becoming the project’s ontology.
9.9 The translation gauntlet
Only after a lineage has stabilized internally should its artifact enter the external epistemic shell. The shell performs:
1. genealogical freeze — code, state, observers, questions, and local verifiers are committed;
2. semantic quarantine — project names and internal explanatory labels are removed;
3. lineage replacement — at least two independent teams or systems reconstruct interfaces from restricted contracts;
4. counter-world transfer — the capacity is challenged in substrate families absent from its ancestry;
5. observer extinction — the originating observer and descriptors are made unavailable;
6. adversarial controls — positive power, exchangeability, leakage, and metadata controls traverse the exact terminal path;
7. loss accounting — every translation records what became unavailable, not merely what survived;
8. terminal English render — human description is generated from frozen consequences and explicitly marked as lossy.
The gauntlet does not ask whether a symbol survives unchanged. It asks whether a capacity remains reconstructible and useful when the original basis is unavailable.
9.10 A stronger evidence vector
No single number should certify success. The terminal evidence vector should include at least:
E = (X, T, F, C, P, L, R)
where:
• X: representation expansion—new primitives repay their cost across tasks;
• T: cross-family transfer—capacity survives distinct substrates;
• F: feasibility extension—previously unreachable operations become reachable;
• C: causal autonomy—effects depend on lineage-specific construction rather than generic input sensitivity;
• P: persistence—capacity remains after origin observer and local verifier extinction;
• L: transduction loss—what is destroyed or distorted in each translation;
• R: reproducibility across independent lineages and physical/software environments.
Promotion requires a preregistered pattern over the vector, not a universal threshold for all future discoveries. Different artifact classes may justify different terminal assays, but each assay must freeze before contact with the promoted artifact.
10. Anti-loop operating doctrine
10.1 No new experiment number until a promotion event
The programme should not authorize Experiment 0008 merely because time has passed or a new architecture can be described. A new number should require a machine-produced promotion event with all of the following:
• a frozen genealogy extending over many generations rather than one finite matrix;
• an internally generated observer or question not present as a direct template at initialization;
• measurable feasibility extension or amortized representational compression;
• adoption or reconstruction by an independent lineage;
• a power-tested terminal assay frozen after lineage production but before artifact contact;
• a public claim smaller than the full philosophical interpretation.
Until then, work belongs to the engineering phase of INGE 2.0, not a sequence of discovery experiments.
10.2 Kill criteria
A research direction should be terminated or redesigned when any of the following occurs:
1. Descriptor lock: diversity increases only under a fixed human descriptor.
2. Noisy-TV dominance: novelty grows while predictability, reuse, and intervention value do not.
3. Verifier monoculture: one metric controls more than a declared share of compute over several epochs.
4. Genealogical shallowness: apparent innovations disappear when recent ancestors are removed.
5. Transfer collapse: capacities fail outside their originating substrate family.
6. Interface capture: cross-world success depends on a shared embedding or hand-written feature map.
7. Control inversion: evidence machinery again receives more development than generative variation.
8. Narrative inflation: English interpretation grows while the executable evidence vector remains unchanged.
These are anti-loop conditions, not safety theatre. They protect the programme from spending months polishing a closed possibility space.
10.3 A fixed budget for proof machinery
During the generative phase, no more than a fixed fraction of engineering effort should be spent on new control architecture. The existing 1.x artifacts provide replay, freezing, blind lineage, and control patterns that can be reused. New proof machinery should be added only when a new artifact class creates a specific unresolved threat.
This reverses apparatus gravity: generative capacity grows first; the shell expands only in response to something that actually exists.
10.4 Mandatory red-team questions
At each milestone, the programme must answer:
• What part of the output could have been predicted from the initialization vocabulary?
• Which observer distinctions were inherited, and which were generated?
• What becomes impossible if the new primitive is removed?
• Does the result transfer without a shared latent representation?
• Can an independent lineage reconstruct the capacity from consequences alone?
• What simpler artifact or leakage channel explains the result?
• Is the system producing cumulative novelty or merely cycling through encodings?
• Would the claim become less impressive but more accurate if stated operationally?
11. Concrete successor programme
11.1 Phase A — kernel and genealogy, not claims
Build a minimal engine supporting two substrate families, three observer constructors, mutable questions, genealogical persistence, and at least two provisional verification games. The purpose is not to discover a law. It is to prove that components can mutate independently, migrate between bundles, and alter the representational basis used by descendants.
Deliverables:
• deterministic replay from a frozen genealogy;
• exact ancestry and resource accounting;
• evidence that an observer representation changed rather than merely changed parameters;
• no global English labels inside runtime state;
• automatic archive of extinctions and failed transfers;
• a fixed negative-benchmark adapter using Experiment 0007 controls.
11.2 Phase B — open-ended cultivation
Run many independent ecologies with no terminal research target. Each ecology should be long enough for artifacts to persist beyond the lifespan of their creators and for questions to migrate between lineages. Compute allocation should reward plural minimal criteria and periodically remove dominant descriptors.
The output of this phase is not a paper claim but a corpus of frozen genealogies. Human inspection may monitor engineering health and containment, but semantic annotations must not feed back into selection.
Minimum evidence of progress:
• sustained growth in the number of operationally non-equivalent observer partitions;
• repeated cross-lineage adoption of generated questions;
• at least one representational primitive whose removal increases solution cost across a task family;
• at least one capacity transferred across substrate families through an independently built interface;
• absence of collapse to a single metric, descriptor, or source lineage.
11.3 Phase C — blind promotion
Select promotion candidates by a procedure frozen before human semantic review. A candidate package should contain ancestry, executable state, local evidence, and a restricted interface contract, but no project interpretation. Independent auditors construct a terminal assay and power controls without seeing the candidate data. Only after the assay freezes does contact occur.
The first legitimate successor experiment would then test a specific capacity expansion. Its number should mark evidence contact, not architectural aspiration.
11.4 Phase D — human terminal rendering
If a candidate passes, multiple independent interpreters describe the operational residue in English. Their descriptions are compared for shared consequences and incompatible metaphors. The report publishes:
• the frozen evidence vector;
• the loss ledger;
• the smallest shared operational description;
• divergent interpretations rather than a forced single name;
• all null and failed-transfer results;
• the exact boundary between demonstrated capacity and philosophical implication.
11.5 Resource policy
The programme should scale only after evidence of cumulative generativity.
More compute is justified when additional lineage depth creates new operational primitives or transfers—not when it merely increases sample count. A staged allocation prevents an expensive closed system from becoming credible through size alone.
Recommended sequence:
1. small deterministic prototypes and mutation audits;
2. dozens of short ecologies to identify collapse modes;
3. a small number of long persistent ecologies;
4. blind translation exercises on artifacts sampled before interpretation;
5. one terminal promotion assay.
12. What can and cannot be claimed now
12.1 Claims supported by the record
The record supports the following statements:
• Apparent invariants in early INGE traces were sensitive to operation, grammar, recorder, interpreter, or null construction.
• Fixed-width sequential and partial-order observer families produced no survivor under their declared attack boundaries.
• A complex executor can pass structural reachability and local controls while failing semantic implementation equivalence.
• A small causal transport kernel can be replayable, control-sensitive, adversarially audited, and reproducible.
• Blind source-lineage replacement can succeed under a frozen interface without post-contact source repair.
• Independent auditors can share a blind spot when a common receiver/evidence pathway destroys the distinction they are expected to detect.
• Preregistered controls prevented Experiment 0007 from contacting research data with an insensitive apparatus.
• The project’s current architecture is better treated as a negative benchmark and methodological archive than as a discovery engine.
12.2 Claims not supported
The record does not support:
• discovery of a nonhuman-native invariant or ontology;
• empirical access to an ASI perspective;
• a new law of physics, mechanics, computation, or cognition;
• proof that human categories can be fully escaped;
• source-specific causal residue in Experiment 0007;
• superiority over contemporary automated discovery or open-ended-learning systems;
• historical priority for the general vocabulary/verifier problem;
• removal of the need for human judgment, ethics, safety, or scientific validation.
12.3 The strongest honest thesis
The strongest defensible thesis is methodological:
| A programme seeking nonhuman novelty can become trapped by the meta-ontology of its own evidence apparatus. Rotating worlds and observers is insufficient if the system cannot generate and replace the primitives by which worlds, observers, questions, and successful consequences are constituted. The appropriate architecture is an open-ended co-genesis ecology enclosed by, but not continuously optimized against, a rigorous terminal epistemic shell. |
This thesis is consistent with contemporary open-endedness research and sharpened by the project’s executable failures. It is a proposal and an interpretation, not yet an empirical theorem.
13. Conclusion: the productive exit
We wanted to move beyond experiments humans already perform. We did not accomplish that. We did something less dramatic and more useful than pretending: we built a sequence of increasingly hostile tests and watched our own candidate objects disappear. We then built an apparatus whose gate exposed that its phases were semantically hollow. We reduced the system to a working causal kernel, replaced one lineage blindly, and finally discovered that a sophisticated end-to-end observer could not see the signal it was built to see.
The project did not descend because it used controls. It descended when controls became the main creative object and each failure justified another layer of the same architecture. The escape is not fearlessness expressed as fewer safeguards. It is fearlessness expressed as willingness to abandon the architecture, preserve the negative result, and let the next system generate distinctions we did not select in advance.
The decisive change is from mutual ontology extinction to endogenous co-genesis. Extinction asks what remains after removal. Co-genesis asks what new capacities become possible when worlds, observers, questions, and provisional values transform each other over long genealogies. Extinction remains useful as a terminal attack, but it should no longer organize the generative interior.
The existing work should be frozen as INGE 1.x and used as a control kernel, a failure archive, and a discipline against narrative inflation. INGE 2.0 should begin without a new experiment number. Its first goal is not a claim but a living genealogy whose representational basis changes, whose questions migrate, whose local verifiers can die, and whose useful capacities sometimes survive transfer to worlds where their original names and primitives do not exist.
Only then should English return.
And when it returns, it should be allowed to say less than the process did.
Appendix A. Detailed experimental evidence ledger
A.1 Experiment 0002
• Question: survival under operation ablation, grammar rotation, early-lineage intervention, recorder rotation, and independent interpretation.
• Key result: seven of eight initial recurrences broke under at least one single-operation ablation; redundant relation-generation paths required compound ablation.
• Interpretation: no universal property established; weak cross-interpreter overlap potentially manufactured by the shared event contract.
• Consequence: extract minimal name-free relations and attempt to destroy them.
A.2 Experiment 0003
• Object: anonymous fixed-width lag-one formulas.
• Separation: discovery, minimization, held-out attack, and interpretation.
• Key result: no formula survived interpreter, scale, seed, representation, and shuffled-contract controls.
• Consequence: change the observer type rather than enlarge the formula vocabulary.
A.3 Experiment 0004
• Object: partial-order traces with no exported event number, timestamp, fixed-width event vector, or privileged linearization.
• Scale: 48 traces; four interpreters; four seeds; three scales; two ancestry-null families.
• Key result: 3/5 formulas entered attack; 0/3 survived.
• Consequence: no native object; observational replacement was real but still apparatus-bound.
A.4 Experiment 0005 lineage
• 0005: invalid or contaminated; claims void.
• 0005.1: no residue survived.
• 0005.2–0005.4: architecture, preregistration, power, executor, and control revisions; no validated nonhuman-native residue.
• Consequence: apparatus expansion outpaced empirical contact.
A.5 Experiment 0006
• Implementation: phases 0–10, fresh workspace controls, reachability and locking.
• Local result: CONTROL_BOUNDARY_PASS.
• Authorization result: denied before Phase 2.
• Main cause: operational mismatch between phase labels and scientific actions, including shared mutation paths, non-replayed counterfactuals, weak null differentiation, controls bypassing the research pipeline, common pre-auditor representation, unenforced denial, and nominal extinction.
• Consequence: do not confuse executable topology with implemented method.
A.6 Minimum Nonhuman Demonstrator
• Candidate effect: causal auditor pass; reachability auditor pass; distance to ablation 127; distance to shadow 128; exact replay pass.
• Positive control: both auditors pass; distance to ablation 128; distance to shadow 51.
• Negative control: both reject; distances 0/0.
• Audit: 18/18.
• Replication: two fresh rebuilds bit-identical.
• Claim: causal transport kernel only.
A.7 Lineage Replacement Test
• Replacement: Dynamic B reimplemented context-blind under a frozen interface in freestanding assembly.
• Integration: first contact passed without repair.
• Audit: 19/19.
• Replication: two fresh workspaces bit-identical.
• Claim: apparatus-level source-lineage replacement; source-specificity not established.
A.8 Experiment 0007
• Freeze: 38-file executable closure; two separate auditor lineages; eleven roots committed in one entropy acquisition.
• Positive power control: six auditor/null cells, all failed the 320/512 threshold.
• Exchangeable negative: all passed.
• Metadata leakage: all exactly chance and passed.
• Legacy transport: passed.
• Replay: 4,096 persisted branch quartets exact.
• Incident: eight empty jail directories blocked terminal sealing; scientifically secondary to power failure.
• Final state: FORENSIC_AUDIT_PASS_CONTROLS_FAIL_RESEARCH_LOCKED.
• Claim: the terminal path was blind to planted higher-order structure; no research-source conclusion.
Appendix B. Concept glossary
Apparatus gravity — the tendency of a growing experimental apparatus to make its own maintenance and validation the dominant research activity.
Capacity expansion — a new operational primitive or process that makes previously infeasible discriminations, interventions, predictions, or transfers feasible within a resource budget.
Control inversion — the condition in which the production of admissibility evidence consumes more generative capacity than the phenomenon under study.
Endogenous co-genesis — mutual evolution of worlds, observers, actions, memories, questions, and provisional standards, without a single terminal human-authored target ontology.
Epistemic shell — the external, non-generative layer that freezes lineages, applies controls, performs independent translation, and limits human claims.
Lineage replacement — blind creation and pre-contact freezing of a replacement component under a restricted interface, followed by one integration with the existing apparatus.
Meta-ontological closure — variation among objects without corresponding variation in the language that defines which objects and variations are possible.
Observer debt — the accumulated obligation to determine which regularities were induced by measurement, preprocessing, representation, or selection.
Ontological overfitting — apparent diversity or discovery confined to a fixed designer-authored space of possible worlds, observers, and hypotheses.
Terminal transduction — conversion of frozen operational consequences into human language after generation and validation, with explicit accounting of loss.
Translation asymmetry — the possibility that a cross-world survivor is encoded by the translation bridge rather than generated by either world.
Unknown-by-process — an operational relation or capacity whose generation does not depend on its later human naming and whose consequences survive relevant representational replacement.
Verifier capture — narrowing of generation to objects legible to a pre-existing evaluator.
Appendix C. Proposed INGE 2.0 artifact schema
Each frozen bundle should contain:
| Field | Required content |
|---|---|
| bundle_id | Opaque commitment derived after freeze, not a semantic name. |
| ancestry | Parent commitments, borrowed components, mutations, extinctions, and migration events. |
| world | Constructor, update laws, resource model, and executable state commitment. |
| observer | Sampling/intervention procedure, costs, output types, and known blind spots. |
| action | Admissible interventions and their resource accounting. |
| memory | Compression or state-retention mechanism and transformation history. |
| question | Executable fork, horizon, branch contract, and local adoption history. |
| local_verifier | Provisional game, challenger lineage, budget, and failure cases. |
| capacity_claim | Machine-internal dependency record; no English semantics. |
| transfers | Attempts across world and observer families, including failures. |
| descriptor_history | Partitions used, retired, replaced, or defeated. |
| containment | Execution boundaries, resource caps, and external interfaces. |
| terminal_status | Unreviewed, quarantined, assay-frozen, passed, failed, or uninterpretable. |
Appendix D. Selected references
External primary literature and research reports
[1] Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292. https://arxiv.org/abs/2408.06292
[2] Yamada, Y., Lange, R. T., Lu, C., Hu, S., Lu, C., Foerster, J., Clune, J., & Ha, D. (2025). The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search. arXiv:2504.08066. https://arxiv.org/abs/2504.08066
[3] Romera-Paredes, B., et al. (2024). Mathematical discoveries from program search with large language models. Nature, 625, 468–475. https://doi.org/10.1038/s41586-023-06924-6
[4] AlphaEvolve Team. (2025). AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind. https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
[5] Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv:2505.22954. https://arxiv.org/abs/2505.22954
[6] Lehman, J., & Stanley, K. O. (2011). Abandoning objectives: Evolution through the search for novelty alone. Evolutionary Computation, 19(2), 189–223. https://doi.org/10.1162/EVCO_a_00025
[7] Mouret, J.-B., & Clune, J. (2015). Illuminating Search Spaces by Mapping Elites. arXiv:1504.04909. https://arxiv.org/abs/1504.04909
[8] Wang, R., Lehman, J., Clune, J., & Stanley, K. O. (2019). Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions. arXiv:1901.01753. https://arxiv.org/abs/1901.01753
[9] Colas, C., Karch, T., Sigaud, O., & Oudeyer, P.-Y. (2020/2022). Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: a short survey. Journal of Artificial Intelligence Research, 74, 1159–1199. https://arxiv.org/abs/2012.09830
[10] Hughes, E., Dennis, M., Parker-Holder, J., et al. (2024). Open-Endedness is Essential for Artificial Superhuman Intelligence. arXiv:2406.04268. https://arxiv.org/abs/2406.04268
[11] Huh, M., Cheung, B., Wang, T., & Isola, P. (2024). The Platonic Representation Hypothesis. arXiv:2405.07987. https://arxiv.org/abs/2405.07987
[12] Schmidhuber, J. (2010). Formal theory of creativity, fun, and intrinsic motivation (1990–2010). IEEE Transactions on Autonomous Mental Development, 2(3), 230–247. https://doi.org/10.1109/TAMD.2010.2056368
[13] Schapiro, S., Shashidhar, S., Gladstone, A., et al. (2026 version). Combinatorial Creativity: A New Frontier in Generalization Abilities. arXiv:2509.21043. https://arxiv.org/abs/2509.21043
[14] Cao, Y., & Yang, H. (2026). Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI. arXiv:2607.09560. https://arxiv.org/abs/2607.09560
[15] Tang, Y., & Yang, Y. (2026). AI Research Agents Narrow Scientific Exploration. arXiv:2605.27905. https://arxiv.org/abs/2605.27905
[16] Paolo, G., et al. (2026). TerraLingua: Emergence and Analysis of Open-endedness in LLM Ecologies. arXiv:2603.16910. https://arxiv.org/abs/2603.16910
[17] Momennejad, I., & Raileanu, R. (2026). A Compositional Framework for Open-ended Intelligence. arXiv:2606.15386. https://arxiv.org/abs/2606.15386
Project-internal executable and methodological sources
[P1] INHUMANT NATIVE GENERATIVE SUBSTRATE 0.1 — canonical README and Experiments 0002–0005 artifacts.
[P2] EXPERIMENT 0006 — Implementation, Executable Freeze, and Controls-Only Report 1.0.
[P3] EXPERIMENT 0006 — Separate Pre-Phase-2 Authorization Audit 1.0.
[P4] Minimum Nonhuman Demonstrator 0.1 — Methods and Result.
[P5] Lineage Replacement Test: Dynamic B 0.1 — Methods and Result.
[P6] EXPERIMENT 0007 — Source-Specific Causal Residue: Executable Implementation and Controls-Only Result 1.0.
[P7] ASI Mechanics: Canonical Recompilation and Evidence Architecture 1.0 and Reference Implementation and Simulation Program 0.1.
[P8] Project corpus: ASI Physics / Syntophysics & Ontomechanics, Computronium, Inhumant, and related ASI World texts, used as internal conceptual pressure rather than external empirical authority.
Terminal note
This report is itself an N2 object: a human-readable compression of a larger executable and conceptual ancestry. It should not be fed back into the generative interior as a target ontology. Its recommended architecture is a hypothesis to implement and falsify. If INGE 2.0 merely generates better versions of the concepts named here, it will have repeated the failure this report describes.
END OF MEMORANDUM



