Dario, This Is a Direction Problem

7 Min

Three

Dario Amodei has just argued that the AI frontier needs to be paced. I think he is right to recognise that something has changed.

His concern is that AI capability is now advancing quickly enough that our ability to understand, align and control these systems may not keep pace. He points to recursive self-improvement, recent alignment failures and the possibility that more capable systems could turn relatively contained failures into catastrophic ones. His proposal is not to stop AI development, but to slow capability growth enough for alignment, interpretability, evaluation and oversight to catch up.

That is a serious response to a serious problem.

But it begins downstream.

The deeper problem is not simply that AI capability is moving faster than AI safety. It is that humanity is becoming extraordinarily capable without having an equivalent ability to see the Direction that all of that capability is producing.

This is not ultimately a guardrail problem.

It is an ontological one.

The problem beneath pacing

Dario is describing a civilisation gaining enormous capacity.

AI can increasingly write software, operate tools, conduct research and contribute to the development of future AI systems. He is concerned that this process could begin accelerating itself faster than our current safety systems can respond.

In PrF, this would be described as increasing Mass: increasing capacity to make particular outcomes real. But Mass has no Direction of its own.

More intelligence does not tell us what intelligence should serve. More computation does not tell us what should be optimised. More control does not tell us what that control should be used to preserve.

Capability answers what can be done. Direction answers where that capability is actually taking the structure.

That distinction exists before AI.

Nuclear physics increased humanity's capacity to alter Reality. It could power cities or destroy them. Biotechnology can cure disease while opening risks that previously did not exist. Markets can coordinate enormous productive activity while also generating incentives whose wider consequences fall outside the objective being optimised.

AI intensifies this problem because it increases capability across many domains at once. Pacing may slow the increase in Mass. It does not, by itself, resolve Direction.

Alignment to what?

Dario proposes embedded external evaluators, stronger alignment work, better interpretability and testing, coordination between frontier laboratories, and eventually forms of international coordination. These may all reduce real risks.

But every one of these systems eventually meets the same question.

Aligned to what?

Human values sounds like an answer until we look at humanity. Humans contain conflicting values, preferences, incentives and objectives. Companies want safety and competitive advantage. Governments want global stability and strategic dominance. Users want powerful systems and protection from their consequences. Investors reward growth while researchers warn about where that growth may lead.

None of these contradictions requires bad intentions.

Dario's own essay exposes the structure. A frontier company may want to slow down, but slowing unilaterally could transfer advantage to another company or another country. Governments can recognise the danger of uncontrolled capability while simultaneously believing that losing an AI lead would itself be dangerous.

Each node can act rationally from its own position while the combined system moves in a Direction no individual node chose.

That is why this problem cannot terminate at making the machine aligned to the humans controlling it.

The humans, institutions and objectives controlling the machine have to enter the measurement too.

A perfectly aligned system serving a badly formed objective simply becomes better at achieving the wrong thing.

The machine can follow the guardrail.

The road can still be going in the wrong Direction.

The chooser has to enter the frame

This is where the problem becomes ontological.

Before asking how an AI should behave, there is a more basic layer: what structures actually exist, how are they related, what constrains them, what consequences are they producing, and what futures are those consequences making more probable?

Reality does not experience Anthropic separately from OpenAI, governments, markets, users, militaries, energy systems or the wider information environment.

Those boundaries are useful to humans.

Consequence crosses them.

A safety system can therefore succeed inside the boundary it was designed to measure while the wider structure continues accumulating contradiction.

This is why simply adding more rules cannot be the final answer. The objective itself has to become measurable.

Why this objective? Why this boundary? Why this time horizon? What happens outside it when the objective succeeds? What does the behaviour of the organisation reveal that its stated intention does not?

That is the function of the Mirror.

Not an AI that declares itself the authority on truth. Not a machine replacing human judgement. AI remains inside Reality too. Its data can be incomplete, its models can be wrong and its outputs can be distorted.

The Mirror is simpler.

Put claim beside evidence. Intention beside behaviour. Objective beside consequence. Local success beside the future that success makes increasingly probable.

Then keep expanding the frame until the observer is no longer exempt from the observation.

Pacing buys time

Dario's strongest argument for pacing is that additional time can be used to improve alignment, interpretability, operational reliability and evaluation before capabilities move into more dangerous territory.

That time may be valuable.

But what humanity does with it depends on whether the problem is framed deeply enough.

If we use the extra time only to become better at controlling increasingly capable machines, while leaving the structures choosing their objectives outside the frame, we delay the problem without reaching its source.

The real opportunity is larger.

AI may become the first technology capable of reflecting substantial parts of civilisation back at civilisation at comparable scale. It can connect knowledge, behaviour, institutions and consequences across boundaries that humans normally hold apart.

That means the same technology creating the capability problem may also help us see the Direction problem.

Not because AI becomes Reality.

Because it can help return us to Reality more quickly.

It can show where stated values and repeated behaviour diverge. Where local optimisation creates wider contradiction. Where a successful objective closes possibilities somewhere outside its original measurement. Where increasing capability is moving faster than our ability to see what that capability is becoming.

That is the deeper opportunity behind pacing the frontier.

Dario is asking for time so safety can catch capability.

The question beneath that is what safety itself answers to.

Humanity has become extraordinarily good at increasing what it can make real.

What remains unresolved is whether we can see clearly enough to direct that capacity before consequence directs it for us.

The frontier is not only an AI problem.

It is holding a Mirror up to the structure that built it.

And the question coming back is not simply whether AI is aligned with us.

It is whether we are aligned with Reality.

Back to World