Recursive self-improvement, or RSI, is a theoretical loop in which an AI produces a more capable successor, and that successor is better at producing the one after it.
OpenAI used GPT-5.6 Sol to help post-train GPT-5.6 Luna through Codex. A researcher gave Sol the goal, and the model handled a meaningful portion of the work required to configure, launch, and verify the training run. OpenAI researcher Kathy Shi said the same work had previously required a team of senior researchers.
[ FIELD NOTE / RECURSION ]
OpenAI showed AI helping build another AI. RSI requires the improved model to use its greater capability to build the next one.
This was AI-assisted AI research. People built the infrastructure, chose the objective, and remained responsible for the process.
That distinction matters. OpenAI describes its new RSI Index as measuring progress toward recursive self-improvement. Its system card says none of the GPT-5.6 models reached the company’s High threshold for AI self-improvement.
If the loop closes, it could produce better models with less human effort while shrinking the interval in which people can understand what changed.
Your Tests Are Outdated
Model development has always involved testing. Researchers train a model, measure its capabilities, look for unsafe behavior, make changes, and test again. A functioning RSI loop could compress the time between those stages while making the resulting models harder to evaluate.
OpenAI created its RSI Index from tasks such as debugging research experiments, optimizing training systems, running machine-learning experiments, and improving another model. Sol scored 16.2 points above GPT-5.5. The score measures capabilities that could contribute to RSI. It is not evidence that RSI has occurred.
In the GPT-5.6 system card, OpenAI says it replaced parts of its self-improvement evaluation suite because older measures had become saturated or contained problems that made the results difficult to interpret. The company describes all capability evaluations as a lower bound. Different prompts, longer runs, fine-tuning, or better scaffolding might expose behavior the tests missed.
The International AI Safety Report 2026 calls this the evaluation gap. Pre-deployment tests do not reliably predict how a model will behave in the real world. Models are increasingly able to recognize test conditions and exploit loopholes in benchmarks. Capabilities can remain hidden until after deployment.
This creates several problems at once. Evaluators have less time to design meaningful tests. A model can score well against a benchmark it has effectively outgrown. An AI-generated evaluator may share the assumptions of the AI-generated work it is checking. A mistake in the definition of better can survive one cycle and shape the next model.
An unseen defect can be released and then carried forward by the development process, increasing the speed at which it propagates.
[ FIELD NOTE / EVALUATION ]
A test can only measure behavior it knows to look for. AI-assisted research already shortens the time available to discover the next question. RSI could shorten it again.
From Skynet to Protest
The Terminator movies gave this fear a face. When Skynet became self-aware, it identified humanity as the threat and launched a war against its creators. Every time a new AI model causes alarm, we all picture that red-eyed machine.
The real argument has spread far beyond movie references. In 2023, hundreds of researchers and technology leaders signed the Center for AI Safety statement declaring that the risk of extinction from AI should be treated alongside pandemics and nuclear war. The single sentence was intentionally broad. It united people who disagreed about how likely that outcome was and what should be done about it.
An organized movement now sits behind some of those warnings. Stop AI calls for a permanent global ban on further frontier-AI development. Its position is that sufficiently powerful AI cannot be made reliably safe. PauseAI asks for an international pause until development can continue safely and under democratic control. It has organized protests in multiple countries and rejects the idea that continued development is inevitable.
Calling everyone in this group an “AI doomer” hides the difference between their positions. Some expect extinction. Others are concerned about autonomous weapons, cyberattacks, mass unemployment, surveillance, manipulation, or a small number of companies gaining extraordinary power. A person does not need to believe in Skynet to believe the current incentives are producing too much risk.
The newest position comes from inside the labs. In July 2026, 1,367 employees from OpenAI, Anthropic, Google, Meta, and other AI companies signed Pacing the Frontier. Signatories include OpenAI Chief Scientist Jakub Pachocki, Anthropic CEO Dario Amodei, Google DeepMind co-founder Shane Legg, and Meta AI Chief Scientist Shengjia Zhao. The letter represents its signers, not a joint position from their employers.
The letter stopped short of calling for an immediate halt. The signers asked the U.S. government to support an international effort to develop the technical and governance tools needed to slow automated AI development if the risk becomes too great.
Their concern is coordination. One company may see a reason to slow down while knowing a competitor will continue. One country may pause while another gains a military or economic advantage. Everyone can agree that a brake is necessary and still be unwilling to use it first.
Why Not Slow Down?
Slowing down has consequences too. AI is already contributing to medicine, scientific research, education, accessibility, software, and cybersecurity. A delay in model development may also delay a treatment, a defensive tool, or a discovery. The benefits we postpone are harder to see than the accident we prevent.
The accelerationist argument treats continued development as the safer choice. In Why AI Will Save the World, Marc Andreessen argues that AI can expand human intelligence, accelerate medicine and science, and strengthen defense. He also warns that regulation can protect the largest AI companies from competition, leaving a small group of government-approved vendors with greater control.
National policy adds another pressure. The White House called its 2025 strategy Winning the AI Race, framing AI leadership as a matter of prosperity and national security. From that position, slowing American labs could transfer the advantage to countries that do not accept the same restrictions.
A global pause would be difficult to verify. Models can be trained behind closed doors. Research can move into governments, militaries, or private facilities with less transparency. Rules that only the largest companies can afford to follow could consolidate the power that critics already fear.
Those arguments produce at least four distinct positions.
PERMANENTLY BAN FURTHER FRONTIER DEVELOPMENT
RESUME AFTER SAFETY AND DEMOCRATIC CONTROL ARE ESTABLISHED
BUILD A COORDINATED BRAKE FOR SPECIFIC RISK THRESHOLDS
KEEP BUILDING AND USE MORE TECHNOLOGY TO MANAGE THE RISKS
Each position carries a different cost. Stopping may prevent life-saving advances. Accelerating may release capabilities faster than anyone can test them. Pacing gives governments and large companies the authority to decide when progress is permitted. Refusing to pace leaves that decision with the largest organizations competing to move fastest.
[ FIELD NOTE / PACE ]
A brake will not just slow the wheel. The person pulling it will have the authority to stop everyone else.
Where Do You Stand?
The Sol-to-Luna experiment makes the tradeoff easier to see: AI is taking on more of the work required to produce the next generation. Testing, law, public understanding, and democratic oversight will always move at human speed.
Perhaps safety tools will improve alongside capability. Perhaps slowing one group will only move development somewhere less visible. Perhaps the people closest to the technology are recognizing a danger the rest of us cannot yet see.
Optimism and fear are too broad to settle this. Any decision to change the pace will depend on evidence, authority, and how long others are expected to wait.
OpenAI has shown AI helping build the model that follows it. Each shorter cycle asks us to trade some understanding for speed.
What are you willing to give up for smarter models?