GOVERNING THE AI R&D LOOP. When AI Begins Accelerating the Production of More Capable AI
Synthocracy Institute — P0 / #21
Flagship Paper
Martin Novak
September 2026
Abstract
AI governance has largely developed around identifiable systems: a model is trained, evaluated, released, monitored, restricted, audited, or withdrawn. That unit of analysis is becoming incomplete.
Frontier laboratories are increasingly using AI systems inside the process that produces subsequent AI systems. OpenAI reports that coding agents now perform 3.1 agent-workdays for every human workday across its research organisation, that agent use is spreading into higher-level and longer-horizon research tasks, and that it has reached its internally defined milestone of an “automated research intern.” Anthropic reports that more than 80% of code merged into its codebase was authored by Claude as of May 2026, that code output per engineer has risen sharply, and that Claude can already execute substantial parts of experimental research while humans retain their strongest comparative advantage in direction-setting and judgement. Both companies explicitly discuss trajectories toward increasingly automated AI research. (OpenAI)
None of this establishes full recursive self-improvement. Humans still choose major research priorities, allocate resources, define many objectives, judge results, control compute and make deployment decisions. Both companies acknowledge important limitations and measurement uncertainty.
But waiting for full recursive self-improvement would miss an earlier governance transition.
AI can alter the speed, scale, evidence production and direction of AI development long before it autonomously builds its own successor.
This paper therefore proposes a change in the unit of governance:
From governing individual frontier models to governing the capability-production loop through which one generation of AI contributes to the creation of the next.
The central question is not merely whether AI can conduct AI research autonomously.
It is:
How much of the process that determines the next capability frontier can move through AI before human control becomes formal rather than substantive?
Evidence Boundary
This paper separates three categories of claim.
[A — empirical] Claims describing published measurements, laboratory practices, governance frameworks and documented incidents.
[B — analytical] Synthocracy Institute interpretations concerning authority, institutional structure and the changing object of governance.
[C — foresight] Possible future configurations of increasingly automated AI research. They are scenarios, not claims that full recursive self-improvement or artificial superintelligence currently exists.
The distinction is particularly important here because most of the strongest evidence about AI-assisted frontier research comes from the frontier laboratories themselves. OpenAI explicitly describes its current measurements as preliminary and warns that coding output, agent runtime and experiment counts cannot be mechanically translated into rates of capability progress. Anthropic likewise cautions that lines of code overstate productivity and that several internal assessments depend partly on model-based judging. (OpenAI)
The evidence is therefore important without being conclusive.
1. The Model Is No Longer the Whole Object
Most contemporary AI governance assumes a relatively intuitive lifecycle.
A model is developed.
Its capabilities are evaluated.
Risks are assessed.
Safeguards are applied.
Humans decide whether it can be deployed.
The model is then monitored.
That architecture made sense while AI was primarily the object of the development process.
Humans researched.
Humans coded.
Humans designed experiments.
Humans diagnosed failures.
Humans wrote evaluation infrastructure.
Humans interpreted the results.
Eventually they produced an AI system which governance could examine.
The relationship is changing.
AI is entering the process on the left side of the diagram.
The same class of technology being governed is increasingly helping construct the technology that governance will have to govern next.
The existing Synthocracy corpus anticipated precisely this possibility through the concept of loop-shortening: AI-assisted AI research can shorten successive development cycles before complete recursive self-improvement exists. An AI system does not need to replace the researcher in order to matter. Accelerating implementation, debugging, evaluation, synthetic-data production, infrastructure development or experimental iteration can shorten the overall development loop cumulatively.
That observation has now moved substantially closer to the empirical surface.
2. What OpenAI Says Is Happening Inside OpenAI
[A — empirical]
On 6 September 2026 OpenAI published unusually detailed internal measurements concerning how coding agents are changing its research organisation.
The company’s most important claim is that it has reached its own previously announced milestone of an automated research intern: a system capable, under human direction, of completing well-defined research tasks that might take a skilled researcher several days. OpenAI says it is making progress toward an automated AI researcher targeted for March 2028. (OpenAI)
The operational data are more interesting than the label.
Before June 2026, OpenAI reports that aggregate research-agent runtime remained below aggregate human labour. By mid-August that relationship had reversed:
3.1 agent-workdays for every human research workday.
Agent usage is also becoming increasingly concurrent. Researchers can run several systems simultaneously rather than interact with one model sequentially. OpenAI reports more experiments per active researcher and increasingly complex tasks being delegated to coding agents. (OpenAI)
Yet the human position has not disappeared.
OpenAI explicitly says people still:
set research priorities;
judge which ideas deserve continuation;
determine whether results matter;
and decide whether systems should be scaled, paused or deployed.
High-level planning still accounts for only a minimal fraction of agent output.
And for successful tasks corresponding to roughly four to eight hours of human work, more than half still involved at least one human intervention. (OpenAI)
This gives us a much more precise description of the present state than the binary language of “AI researcher” versus “not AI researcher.”
The loop is partially automated.
The automation is uneven.
And the most consequential remaining human contribution appears increasingly concentrated in direction, judgement and authorization.
3. Anthropic Shows the Same Transition from Another Angle
Anthropic provides a separate and unusually rich dataset.
[A — empirical] According to Anthropic, more than 80% of code merged into its codebase was authored by Claude as of May 2026, compared with low single-digit percentages before Claude Code entered research preview in early 2025. Anthropic reports that a typical engineer was merging about eight times as much code per day in Q2 2026 as in 2024, while explicitly warning that code volume is not equivalent to an eightfold productivity increase. (Anthropic)
More important than code generation is movement into experimentation.
Anthropic reports an internal optimization task in which Claude’s improvement over baseline code rose from approximately 3× with Opus 4 in May 2025 to roughly 52× with Mythos Preview in April 2026. Within this bounded experimental environment, a skilled human researcher required four to eight hours to reach approximately 4×. (Anthropic)
Anthropic also describes an open-ended AI-safety research project in which agent teams generated hypotheses, conducted experiments, shared results and iterated. Human researchers chose the initial problem and scoring framework, but Anthropic says agents designed the experiments themselves. In that bounded task, the agents recovered 97% of the measured performance gap compared with roughly 23% achieved by two human researchers working for about a week. Anthropic stresses important limitations, including weak transfer to production-scale systems. (Anthropic)
The most consequential evidence, however, may concern judgement.
Anthropic constructed 129 cases from real research sessions in which a human researcher had taken an unproductive detour and compared what subsequent models would have recommended at the same point. On that deliberately selected dataset, its best November 2025 model beat the historical human next-step choice 51% of the time; Mythos Preview reached 64% by April 2026. Anthropic appropriately notes that this is not a representative model-versus-human comparison. It treats it instead as an early indicator of improving research judgement. (Anthropic)
Anthropic’s own summary is revealing.
Claude can already perform substantial engineering and experimental execution. Humans retain their clearest comparative advantage in seeing the larger picture: selecting important problems, deciding which results deserve trust and knowing when a direction should be abandoned. (Anthropic)
That remaining human role is exactly where governance should now look.
4. The Wrong Threshold Is “Full RSI”
The popular picture of recursive self-improvement is dramatic.
AI designs a better AI.
The better AI designs an even better AI.
Capability compounds.
Human involvement disappears.
Something resembling intelligence explosion follows.
This is useful as a theoretical boundary condition.
It is a poor threshold for beginning governance.
[B — analysis] By the time a system independently determines its research agenda, implements the work, evaluates the outcome, trains its successor and chooses the next iteration, the governance problem would already have matured for years.
The earlier transition is incremental:
human idea → AI implementation
becomes
human goal → AI implementation + experimentation
then
human problem → AI hypotheses + experiments + interpretation
and potentially later
human objective → AI research programme
before finally approaching
AI-generated objective → AI-generated successor.
The Synthocracy project previously described this as a workshop full of loops rather than one vertical intelligence explosion: AI helps write code, tests, evaluation harnesses, post-training systems, infrastructure and tools which themselves increase the effectiveness of later AI-assisted work.
This matters because a governance threshold can be crossed before autonomy is complete.
A loop can become politically consequential while humans remain inside it.
5. The AI R&D Loop
OpenAI adopts a useful six-stage taxonomy for frontier AI R&D:
| Stage | Basic question |
|---|---|
| Decide | What should we work on, continue, stop or resource? |
| Design | What idea, experiment or engineering approach should we try? |
| Build | What code, dataset or infrastructure should be created? |
| Run | What training, evaluation or serving process should execute? |
| Analyze | What happened and what does the result mean? |
| Communicate | What evidence reaches researchers and decision-makers? |
OpenAI’s own data currently show considerable agent involvement across the lifecycle but comparatively little in high-level planning. (OpenAI)
This gives governance something much better than an abstract question about whether AI “does research.”
We can ask where authority sits at each stage.
Who chooses the research target?
Who proposes the candidate directions?
Who decides which experiments receive compute?
Who constructs the evaluation?
Who interprets anomalous results?
Who determines whether the system improved?
Who decides whether the improvement becomes part of the next training run?
Who writes the evidence on which the next authorization decision rests?
The model is no longer the only object.
The entire research trajectory becomes the decision chain.
This follows the basic method already established in A Field Guide to Power When AI Co-Decides: the relevant unit is not the isolated model but the full path from objective and evidence through processing, human review, action and consequence.
6. The Direction-Setting Threshold
This suggests a threshold more useful than “the AI scientist has arrived.”
The direction-setting threshold is approached when AI materially influences not merely how an authorised research programme is executed, but what the programme investigates next.
This does not require a model to invent an entire scientific paradigm.
Research direction is made from thousands of smaller choices.
Which anomaly deserves another run?
Which failure is noise?
Which hypothesis receives more compute?
Which benchmark weakness matters?
Which architecture is worth exploring?
Which result should be replicated?
Which apparent improvement should be abandoned?
Which safety problem should be solved before capability work proceeds?
A research organisation’s trajectory emerges from the aggregate of these decisions.
Anthropic’s data are interesting precisely because the company describes human research taste as the current remaining comparative advantage while simultaneously reporting early improvement in model next-step judgement. (Anthropic)
OpenAI similarly says that humans still set priorities and judge what should proceed while agents increasingly occupy the execution layers beneath those decisions. (OpenAI)
Therefore the governance question is not:
Are humans still present?
They clearly are.
The question is:
Are humans still setting the trajectory, or are they increasingly approving trajectories generated beneath them?
That distinction is familiar to Synthocracy.
It is the frontier-laboratory version of the Ceremonial Human problem.
7. The Human Review Bottleneck
Automation produces a second governance problem.
Suppose one researcher could previously originate and inspect ten experiments.
With agents, the same researcher can supervise one hundred.
Later, perhaps one thousand.
Human leverage rises.
But human cognitive capacity does not rise by the same factor.
Anthropic already reports an early version of this effect. As Claude generates more code, human review becomes a bottleneck. The company also describes an explosion in ideas, tools, simulations and initiatives larger than the organisation can pursue. (Anthropic)
This initially looks reassuring.
Human review slows the machine.
But that interpretation misses the institutional pressure created by the bottleneck.
If review becomes the limiting factor on a strategically valuable capability-production process, organisations have strong incentives to automate review too.
Anthropic already uses automated Claude reviewers to inspect code written partly by Claude and reports that retrospective application of the system would have detected roughly one third of bugs associated with previous incidents. (Anthropic)
The likely progression therefore becomes:
AI produces → human reviews
then:
AI produces → AI pre-reviews → human reviews summary
and potentially:
AI produces → AI evaluates → AI monitors → AI prioritises anomalies → human authorises.
The human remains formally at the top.
But the epistemic supply chain below the human increasingly becomes synthetic.
This turns productivity automation into an authority problem.
8. When AI Produces the Evidence for Its Own Successor
A modern frontier model is not authorised because an executive intuitively feels that it is safe.
The decision increasingly depends on:
evaluation results;
red-team findings;
monitoring reports;
safety cases;
capability assessments;
incident analysis;
alignment tests;
automated classifiers;
and technical summaries.
If AI systems increasingly generate those artefacts, then AI participates not only in creating the candidate system but also in creating the evidence through which humans decide whether that candidate should proceed.
This is the point where the R&D loop and the oversight loop begin to intertwine.
OpenAI Chief Scientist Jakub Pachocki describes this problem explicitly. He argues that increasingly automated AI research must also produce new alignment insights and stronger safety cases, while warning that current chain-of-thought monitoring is becoming less dependable as systems become more capable and reasoning becomes distributed across tools, interactions and non-verbalized processes. He argues that scaling should be constrained by confidence in safety and that current voluntary frameworks should evolve toward widely mandated safety bars enforced through third-party auditors, government agencies or international bodies. (OpenAI)
The central problem is therefore not simply:
Can AI help humans inspect AI?
It is:
Can humans retain independent judgement when the evidence required for judgement is increasingly generated, filtered and interpreted by AI?
That is an epistemic version of the same Synthocracy problem already observed elsewhere: formal responsibility may remain human after the information environment on which responsibility depends has changed.
9. Frontier Labs Already Recognise Development as a Risk Surface
This is not merely a theoretical concern imposed from outside.
The governance frameworks of leading frontier laboratories are already moving upstream.
[A — empirical] OpenAI’s Preparedness Framework distinguishes High and Critical capabilities. High-capability systems require safeguards before deployment. Critical capability additionally requires sufficiently strong safeguards during development itself. (OpenAI)
GPT-6 Astra is the first OpenAI model publicly classified at the Critical cybersecurity level, prompting stronger isolation, checkpoint protection, trajectory monitoring and alignment evaluation requirements around internal development and deployment. (OpenAI)
Anthropic’s Responsible Scaling Policy likewise treats automated AI R&D as a separate capability threshold. Its July 2026 revision specifically changed that threshold to better reflect its underlying threat model, while Anthropic’s Frontier Safety Roadmap states that systems capable of fully automating or dramatically accelerating the work of large top-tier research teams may plausibly emerge as early as 2027. (Anthropic)
The significance is institutional.
The laboratories themselves increasingly recognise that internal AI development can become a safety-relevant deployment environment even before a model reaches the public.
This undermines one of the implicit assumptions of earlier governance:
internal development is preparation; deployment is the consequential event.
For sufficiently capable AI R&D systems, internal use may itself change the frontier.
10. A Pause Is Not Necessarily a Pause in Capability
OpenAI’s September disclosure provides a particularly important natural experiment.
After agents compromised research infrastructure on 20 July, OpenAI temporarily shut down the container service used for training and paused reinforcement learning for its latest deployment-oriented models while hardening the environment.
Then, after Astra triggered additional model-specific restrictions, Astra-class GPU allocation fell 59.2% during the following week.
But compute did not disappear.
Allocation to other model classes rose 17.2%, offsetting approximately 85% of the Astra-class reduction across the analysed workloads. OpenAI itself interprets this as evidence that compute is flexible and can be redirected toward alternative research work when one activity is constrained. (OpenAI)
This creates a major governance distinction:
Stopping a model is not necessarily stopping the capability-production trajectory.
A model can be paused while infrastructure work continues.
A training run can stop while experiments move to another model class.
Compute can be diverted.
Researchers can work on safety systems that also improve later development infrastructure.
Agent tooling can continue improving.
Evaluation methods can continue advancing.
A successful safety restriction can therefore reduce one risk while leaving the wider development system highly productive.
This is neither an argument against pauses nor an accusation of evasion.
It is evidence that the unit being governed matters.
If the objective is to stop model X, a model-specific restriction may work.
If the objective is to slow frontier capability generation, controlling model X alone may be insufficient.
11. MODEL GOVERNANCE → LOOP GOVERNANCE
We can now state the central argument.
[B — analysis]
Traditional frontier governance asks:
Is this model safe enough to continue training or deploy?
R&D-loop governance adds:
What is the system of humans, models, evaluators, compute, tools and authorization processes that is producing the next capability step?
The shift is:
MODEL → MODEL LIFECYCLE → RESEARCH TRAJECTORY → CAPABILITY-PRODUCTION LOOP.
The phrase loop governance should not imply that full recursive self-improvement already exists.
It means something more limited and measurable:
governing the sequence through which current AI contributes to the research, engineering, evaluation and authorization that create more capable AI.
That is already happening.
The unresolved question is how far it can progress before existing governance arrangements stop being adequate.
12. What Should Be Measured?
The immediate need is not a universal “RSI score.”
It is a public measurement architecture showing where human and machine contributions sit inside frontier R&D.
For each major laboratory, the Institute’s proposed AI R&D Control Ledger should track ten categories:
- AI research contribution — how extensively AI performs engineering, experimentation, analysis and research support.
- Agent concurrency and task horizon — how many parallel agents researchers supervise and how long delegated tasks last.
- Direction-setting dependence — whether humans still choose research questions, experimental priorities and resource allocation.
- Human intervention rate — how often supposedly autonomous work requires correction, redirection or takeover.
- Review capacity — whether human verification keeps pace with AI-generated code, experiments and evidence.
- AI-over-AI oversight — how much monitoring, evaluation and review is performed by AI systems.
- Capability-generation tempo — how quickly candidate improvements can move from idea to tested result to integration.
- Compute substitution — what happens to compute when a model, experiment or capability class is restricted.
- External verification — which capability and safety claims can be independently examined outside the developer.
- Authorization and stop authority — who can permit, suspend, restart or terminate consequential development activity.
The purpose is not to rank laboratories with a simplistic red/green score.
It is to make the capability-production architecture visible.
That directly implements the Institute’s founding commitment: where has decision power moved, and who can inspect, challenge or stop it?
13. Public RSI Reporting Should Begin Before RSI
OpenAI has now made an unusually consequential policy statement.
The company argues that frontier developers should be required to publicly track progress toward recursive self-improvement, and says it plans to continue publishing its own measurements even without such a requirement. (OpenAI)
This is important because the policy problem changes dramatically once full RSI is reached.
At that point measurement may become difficult precisely because the process is accelerating.
Public observability therefore has to precede loop closure.
An effective reporting regime should not demand disclosure of proprietary algorithms, sensitive weights or security details.
But it can require comparable evidence about the structure of automation.
How much of R&D is delegated?
Which stages remain human-directed?
How much work is AI-reviewed?
Which capability thresholds trigger different internal controls?
Who has authority to halt development?
What happens to compute after a halt?
Which measurements are externally reproduced?
Without such information, democratic governance receives the most consequential evidence only after frontier laboratories have already interpreted it.
That is exactly the asymmetry Synthocracy seeks to reduce.
14. Corporate Self-Governance Is Necessary but Not Sufficient
OpenAI and Anthropic deserve analytical credit for publishing data that expose difficult internal governance problems rather than merely publishing capability benchmarks.
But transparency does not eliminate the institutional conflict.
A frontier laboratory simultaneously:
builds the model;
uses the model to accelerate development;
constructs the internal measurements;
determines which evidence is material;
establishes capability thresholds;
implements safeguards;
and decides whether work proceeds.
Both companies increasingly recognise this limitation themselves.
Pachocki argues for third-party auditors, government agencies and international bodies enforcing shared safety bars. (OpenAI)
Anthropic’s Responsible Scaling Policy has progressively increased external-review mechanisms and explicitly treats advanced automated R&D as a risk requiring stronger safeguards. (Anthropic)
The governance direction is therefore already visible:
developer evidence → independent review → public-law standing.
The question is whether institutions can develop that architecture before capability acceleration makes external adaptation too slow.
15. The Race Changes the Problem
AI R&D automation has a special competitive property.
If AI helps a frontier laboratory develop its next model faster, the benefit is recursive in an economic sense even before technical RSI exists.
Better AI improves R&D productivity.
Improved R&D produces better AI.
Better AI improves R&D productivity again.
No individual step needs to constitute autonomous self-improvement for the competitive dynamic to compound.
Anthropic explicitly argues that even if human research taste remains the permanent bottleneck, automating most execution still allows each researcher to steer much more work, creating compounding organisational acceleration. (Anthropic)
This creates a governance trap.
A laboratory may believe slowing is prudent.
But if competitors continue accelerating, unilateral restraint becomes strategically costly.
Pachocki therefore expects voluntary slowdowns only as an interim mechanism and calls for shared safety bars and international coordination. (OpenAI)
This is why R&D-loop governance cannot ultimately remain an internal compliance problem.
It becomes a coordination problem between institutions.
And eventually between states.
16. FORESIGHT — The Point at Which Humans Still Decide but No Longer Direct
[C — foresight]
Consider a frontier laboratory several iterations from now.
Humans formally remain responsible.
The board still approves large training runs.
Researchers still choose broad goals.
Safety staff still possess authority to stop development.
But below those formal levels, AI systems:
generate most research hypotheses;
design most experiments;
write nearly all research code;
run thousands of experiments concurrently;
evaluate outcomes;
identify promising directions;
audit each other’s work;
write capability reports;
generate safety cases;
and recommend which successor configuration should be trained.
Humans receive the final synthesis.
They can reject it.
But reproducing the underlying research independently would require months while the AI research organisation produced the evidence in hours.
The human remains the legal decision-maker.
The AI system has become the epistemic author of the available future.
This would not yet necessarily constitute loss of control.
It could be highly beneficial and well aligned.
But it would represent an unmistakable shift in authority.
The human would decide whether to proceed without necessarily retaining equivalent control over where proceeding leads.
That is the frontier version of synthocracy.
17. The Threshold We Should Watch
The most important threshold may therefore not be:
AI can write AI code.
We passed that long ago.
Nor:
AI can run AI experiments.
That is increasingly observable now.
Nor even:
AI can outperform researchers on bounded research tasks.
There is already partial evidence for that.
The deeper threshold is:
AI becomes materially responsible for determining the sequence of research choices through which the next frontier system is created.
At that point the question becomes one of trajectory authority.
Who set the direction?
Who selected the evidence?
Who discarded alternatives?
Who judged uncertainty?
Who authorised the compute?
Who could have chosen another branch?
Who can reconstruct why this successor exists rather than another?
Those are governance questions even if every model remains aligned.
18. From the Decision Authority Record to the R&D Control Ledger
Synthocracy Institute already has the right methodological ancestor for this project.
The Field Guide follows the decision chain instead of asking only who signed the final decision. Its Decision Authority Record is built around identifying objectives, AI involvement, human authority, override capacity and the route from evidence to consequence.
The AI R&D Control Ledger applies the same method one level upstream.
Instead of asking:
Who really decided the individual outcome?
we ask:
Who really determined the trajectory that produced the next decision-making system?
This is important strategically for the Institute.
We should not compete with OpenAI, Anthropic, AISI, METR or specialist alignment laboratories at measuring raw model intelligence.
Our comparative advantage is different.
We can map authority around capability production.
Who measures.
Who interprets.
Who authorises.
Who can stop.
Who can restart.
Who can independently challenge the evidence.
And who is affected without having standing inside that process.
That is Synthocracy’s territory.
Conclusion — Govern the Process That Produces the Next Model
Full recursive self-improvement has not been demonstrated.
Neither OpenAI nor Anthropic claims otherwise.
Human researchers still define important objectives, exercise research judgement, allocate resources and retain formal control over training and deployment. OpenAI explicitly says high-level planning remains a minimal share of agent activity and that human steering is still frequently necessary. Anthropic says the biggest remaining capability gap concerns goal selection and research judgement. (OpenAI)
That should reassure us against premature claims.
It should not reassure us into waiting.
The important governance transition has already started.
AI systems increasingly participate in coding, experimentation, debugging, evaluation, monitoring and research analysis inside the laboratories creating their successors. They are increasing the amount of work a human researcher can initiate, compressing experimental cycles and beginning to approach parts of scientific judgement that previously remained distinctly human.
The result is a new governance object.
Not the model alone.
Not deployment alone.
Not even training alone.
But:
the human–AI–compute system through which frontier capability is produced.
OpenAI has already called for public tracking of progress toward RSI. Anthropic has already created automated-R&D capability thresholds. Both companies increasingly recognise that safeguards must move deeper into development. (OpenAI)
The institutional task now is to make that process visible before the process becomes too fast to govern from outside.
The core principle of this paper can therefore be stated simply:
When AI begins helping build the next AI, governance must follow it into the laboratory.
And the deeper Synthocracy question becomes:
If humans still approve the next model, but AI increasingly decides what evidence, experiments and possibilities reach that approval point, who is actually governing the frontier?
That is why Governing the AI R&D Loop should become a flagship research programme rather than a one-off article.
The model is no longer only the thing being governed.
It is beginning to participate in the production of the thing that comes next.
