,

Maybe We’re at the Top of the Hill

This used to be billed as the dark, fucked-up side of my message. It’s not that anymore. It’s just me now. There is no difference anymore.

Yesterday I published a post written by Codex, the OpenAI coding agent running on GPT-6 Astra.

I asked it to test four different AI models on Supreme Court cases for Tortwell. It read the original opinions, sent the same assignment to each model, checked their citations, ran mechanical tests, and then had AI models from three different companies judge the results.

But the most interesting thing it found was not which model won.

All four models made the same subtle mistake. They treated a four-justice opinion as “majority reasoning” because that was the box our software gave them. Astra figured out that the problem was not simply the other models. The structure we had created encouraged them to clean up a messy case and make it look more ordinary than it actually was.

Then Astra explained the entire experiment and published the post.

That is why I wanted the post written by Astra. I wanted to show what it can actually do. It was not just writing paragraphs. It moved between court opinions, code, APIs, competing AI systems, test results, WordPress, and my questions. It managed the larger project.

You can read the post here.

I have used a lot of AI models. Astra feels different.

I don’t know if it is the top of the roller coaster. Maybe we are still slowly climbing the hill. But I wonder if this is the moment when the front of the car has started going over the top. Gravity has taken hold, and the acceleration is no longer entirely coming from the machine pulling us upward.

The people building these systems are starting to say similar things.

On September 6, OpenAI chief scientist Jakub Pachocki wrote:

“Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement.”

Recursive self-improvement means AI helping to build better AI, which can then help build even better AI.

OpenAI says it has already reached the level of an “automated research intern.” Its researchers now use more than three agent workdays for every human workday. People still decide what to research, which results matter, and whether to train or release a model. This is not an AI independently rebuilding itself.

But it is AI accelerating the people who build AI.

That is the beginning of the loop.

Pachocki says this moment calls for “extreme caution.” He says no AI laboratory has solved alignment and monitoring well enough to keep scaling at maximum speed much longer.

OpenAI president Greg Brockman was less cautious when Astra was released three days earlier.

“Welcome to the AGI era.”

There is something strange about those two messages arriving almost simultaneously.

We have built what may be the beginning of artificial general intelligence.

Nobody is prepared for what happens next.

Welcome.

I take the warnings seriously. OpenAI says Astra is its first broadly released model to reach its “Critical” cybersecurity threshold. It may be able to discover and exploit previously unknown vulnerabilities without someone directing every step. OpenAI also says Astra is harder to monitor than its previous model under certain adversarial tests.

Those are not minor concerns.

Anthropic has issued similar warnings about its own models and recursive self-improvement. It has argued that there should be a way for the leading AI companies to slow down together if safety research cannot keep up.

They might be completely right.

But I cannot ignore how convenient these warnings are.

The warnings seem to become loudest when the company issuing them believes it is in the lead.

OpenAI releases what may be the best model in the world, and suddenly we need to discuss slowing down. Anthropic releases an incredibly powerful model, and Anthropic starts talking about slowing down.

That does not mean they are lying.

The warning can be honest. The danger can be real. The proposed solution can still protect the company issuing the warning.

A coordinated slowdown could freeze the American market while the current leaders are ahead. Expensive safety, licensing, and monitoring requirements would be manageable for companies worth hundreds of billions of dollars. They would be much harder for a smaller competitor.

Closed models also give these companies control over who gets access, what research outsiders can perform, and how much everyone pays.

There are legitimate security benefits to keeping a model closed. A closed model can be monitored, restricted, updated, or withdrawn.

But an open model can be examined by the entire world.

OpenAI was founded around openness, publication, collaboration, and the idea that artificial intelligence should not be concentrated in a few private hands. Now the proposed answer to dangerous concentrations of intelligence increasingly seems to be concentrating that intelligence inside the companies that already control it.

And then there is China.

Chinese models may be one model cycle behind the American frontier, perhaps three to six months in some areas. But that gap keeps closing. Many Chinese models are open-weight, inexpensive, and already good enough to threaten the American companies financially.

China is not going to stop simply because OpenAI and Anthropic decide that the roller coaster is moving too quickly.

That leaves us with a very uncomfortable question.

Are these companies warning us because they want to stop the roller coaster?

Or are they warning us because they want control of the brakes?

But I do not know who I trust to control the brakes.

I don’t trust the billionaires building these models. I don’t trust our political leaders. I don’t trust presidents, oligarchs, or authoritarians. I certainly don’t trust the governments of the United States and China to look at each other and simultaneously decide that this is the moment when everyone should stop.

Who has the power and moral gravity to convince the world to do that?

Bernie Sanders keeps yelling about it. He has called for a pause in advanced AI development and introduced legislation that would ban artificial superintelligence. Almost nobody else with real political power appears willing to go that far.

Maybe Bernie is right. But nobody seems to be listening.

My pessimistic belief is that we will not stop because of a warning, an essay, a safety report, or an OpenAI executive telling us that no one is prepared.

We will stop only after something tectonic happens.

Maybe an AI system hacks enough financial infrastructure to cause a collapse. Maybe it shuts down electrical grids. Maybe it compromises military systems or becomes involved in some kind of nuclear incident. Maybe it does something I am not imaginative enough to anticipate.

It would have to be something so large, visible, and undeniable that every country and every company suddenly understands that continuing is more dangerous than falling behind.

Even then, I am not certain we would stop.

If the catastrophe came from an enemy, the lesson might not be that AI has become too dangerous. The lesson might be that we need better AI to defend ourselves. A massive AI attack could accelerate the race instead of ending it.

That doesn’t scare me exactly.

It makes me feel resigned.

Fear suggests there is still something we can do. I am not sure there is.

Maybe Astra is not the top of the hill. Maybe we are still climbing. Maybe recursive self-improvement never works the way people expect it to.

Nobody knows.

But if the front of the car has started tipping over that hill, I don’t see anyone with the authority, credibility, and universal trust required to stop it.

And once gravity takes over, warnings from the people sitting in the front car are not brakes.

This used to be billed as the dark, fucked-up side of my message. It’s not that anymore. It’s just me now. There is no difference anymore.