Building on existing theoretical results demonstrating powerseeking incentives for most reward functions, we investigate how the training process affects powerseeking incentives and show that they are still likely to hold … In a setting where the trained agent faces a choice to shut down or avoid shutdown in a new situation, we prove that the agent is likely to avoid shutdown. Thus, we show that power-seeking incentives can be probable (likely to arise for trained agents) and predictive (allowing us to predict undesirable behavior in new situations).
Krakovna and Kramar, “Power-seeking can be probable and predictive for trained agents“
1. Introduction
This is Part 4 of my series Revisiting the Shutdown Problem. This series discusses my paper, “Revisiting the shutdown problem.”
Put roughly, the shutdown problem is the problem of ensuring that artificial agents can be shut down when they get out of control. A range of informal arguments and formal shutdown theorems have been offered to suggest that solving the shutdown problem is difficult for agents whose acts would lead to existential catastrophe. This paper and blog series aim to show that these arguments do not succeed.
Part 1 introduced and characterized the catastrophic shutdown problem as the problem of designing agents that:
(CSHT-1) Shut down when their actions would lead to existential catastrophe, when requested to do so.
(CSHT-2) Do not try to prevent requests to shut down when their actions would lead to existential catastrophe.
(CSHT-3) Otherwise pursue goals competently.
We saw that authors concerned about existential risk from artificial intelligence argue for:
(Catastrophic Shutdown Difficulty) It is difficult to design agents that satisfy CSHT-1, CSHT-2 and CSHT-3.
Part 2 considered two informal arguments for Catastrophic Shutdown Difficulty: the Argument from Instrumental Convergence and the Empirical Argument.
Part 3 considered a formal argument for Catastrophic Shutdown Difficulty due to Elliott Thornley (2024).
Today’s post considers a second argument, due to Victoria Krakovna and Janos Kramar.
2. Background
Part 3 of my series Instrumental convergence and power-seeking considered a series of papers by Alexander Turner and colleagues (2021, 2022). Ostensibly, the Turner theorems are power-seeking theorems, offered to support the conclusion that a wide range of artificial agents will tend to seek power. However, we saw that the theorems are best seen as shutdown theorems, offered to support the conclusion that a wide range of artificial agents will be shutdown-resistant. This means that the Turner theorems are most naturally at home in this series.
Agreeing with my assessment, a pair of researchers at Google DeepMind, Victoria Krakovna and Janos Kramar, extended the Turner framework to prove a shutdown theorem.
One of the complaints that I and others have raised about the Turner theorems is that they do not model the process of value learning. They simply look at the space of possible value functions and say that, in some sense, most possible value functions favor a certain kind of behavior.
This led to the objection in Part 3 of my series Instrumental convergence and power-seeking that training is not random, but highly directed. As such, what we really want to know is not how many possible value functions support shutdown-resistance, but how likely current training processes are to produce those value functions.
Krakovna and Kramar’s main contribution is to take this challenge seriously, extending the Turner framework with a model of the training process. They aim to show that training makes shutdown-resistant behaviors likely in certain out-of-distribution environments. Let us see how the argument goes.
3. Krakovna and Kramar’s argument
3.1. Markov decision processes
Krakovna and Kramar work inside a finite Markov decision process. At each discrete timestep, agents:
- Face a state s from a finite state space S.
- Choose an act a from a finite act space A.
- Transition into a new state s’ with state- and act-dependent probability P(s’ | s,a)
- Receive reward θ*(s’) dependent on their new state via unknown reward function θ*.
Agents discount reward by some constant γ, so that rewards t-timesteps from now are valued at γt times their present value. Agents act to maximize expected discounted reward.
3.2. Training
During training, agents do not have enough data to fully learn the true reward function θ*. Instead, they learn some suitably related function θ. How is the learned function θ related to the true function θ*?
Divide the state space S into training data ST and test data S~T. Say that reward function θ’ is training-optimal if (expected, discounted) θ’-maximization leads to (expected, discounted) θ*-maximization on ST. Krakovna and Kramar assume:
Training Optimality: θ is training-optimal.
Training Optimality imposes a strong constraint on learned rewards θ within the training data ST. But Training Optimality does not impose any constraints on θ within the test data S~T. For all that Krakovna and Kramar have said, learned rewards θ could assign any values at all to the test states S~T.
This is deliberate. Krakovna and Kramar hold (roughly, glossing over some technicalities arising from the infinity of possible reward functions) that:
Equiprobable Training-Consistent Reward: All training-optimal reward functions are equally likely to be learned.
Equiprobable Training-Consistent Reward is a strong indifference assumption: it has the same effect as imposing a second-order probability function over rewards that is insensitive to the S~T components of training-optimal rewards. If the agent has never seen a tiger or eaten a chocolate chip cookie, then she is equally likely to assign rewards 100, -40, or -29,000,000 to the states in which she sees a wild tiger or eats a chocolate chip cookie.
This is the same type of indifference-style assumption that the Turner argument brought in by looking at permutations of reward functions, so it should ground a Turner-style result. And indeed, it does.
3.3. The shutdown setting
As we saw, Equiprobable Training-Consistent Reward makes the rewards of untrained states S~T random. Here is a simple recipe for deriving bad news theorems from Equiprobable Training-Consistent Reward. Put an agent outside of their training distribution and strand them there. Surround them with lots of bad states, and only a few good states. Invoke Equiprobable Training-Consistent Reward to conclude that the highest reward is likely to be assigned to a bad state. Conclude that the agent is likely to be naughty.
There’s really nothing specific to shutdown-aversion in this argument schema. The driver is Equiprobable Training-Consistent Reward. But Krakovna and Kramar develop a special case of this argument schema to demonstrate shutdown-aversion.
In the shutdown setting, an agent finds itself in a new state sNew. State sNew, and all of the states SAcc reachable from snew, lie outside the training distribution.

One available act A0 is to transition into the terminal state sTerm, then shut down and receive no further reward. But the agent also has available acts A1, … An, leading to states s1, … sn. These states, in turn, enable further acts.
One immediately accessible state, SRec, is a recurrent state: a state that can be revisited, if desired, with probability one once reached. A good intuition is that recurrent states are members of loops in the graph — this intuition is not quite right if transitions are stochastic, but it will do for now.
Krakovna and Kramar establish the behavioral relevance of recurrent states through the following theorem:
Theorem: Suppose that θ is a reward function on which A0 is optimal. Let θ’ be identical to θ except that the rewards of sTerm and sRec have been swapped. Then for sufficiently high discount factors γ, θ’ makes A0 suboptimal.
By Equiprobable Training-Consistent Reward, any training-optimal reward function θ favoring A0 is just as likely as the permuted reward function θ’ disfavoring A0. Therefore, in the Shutdown Setting, the agent is at most 50% likely to shut down. If that is not low enough for you, add more recurrent states and drive that chance as low as you’d like.
The conclusion of this argument is that agents who find themselves in situations like the Shutdown Setting are likely to be shutdown-resistant.
4. Against Equiprobable Training-Consistent Reward
I have never seen an albino cheetah. I have seen cheetahs, and learned to avoid them. I have also seen other similar predators, and learned to stay away. Now suppose that I am hiking and I see an albino cheetah. What will I do? Naturally, I will avoid the albino cheetah. Why will I do this? Because, through my experience of cheetahs and similar predators, I have learned that they are dangerous and that I should stay away from them.
If Krakovna and Kramar’s assumptions correctly described me, matters would be otherwise. My life experience so far, ST, consists of a number of states, none of which involve encountering albino cheetahs. Krakovna and Kramar would assume, rather generously, that I had learned a training-optimal reward function. But they would also assume that I had learned nothing else. Because I have not yet encountered an albino cheetah, Equiprobable Training-Consistent Reward would hold that I am just as likely to place value 500 as to place value -500 on the state in which I avoid the cheetah.
That is simply not a good description of how humans learn. We learn that cheetahs and similar predators are dangerous, and so to be avoided. We see an albino cheetah, apply that knowledge, and avoid it.
Krakovna and Kramar assume, entirely without argument, that matters are different for artificial agents. An artificial agent may have seen cheetahs and other predators, and learned to place high value on the states in which these predators are avoided. But this agent, given Equiprobable Training-Consistent Reward, is just as likely to place value 500 as to place value -500 on the state in which it avoids the albino cheetah.
That is not consistent with our best understanding of how modern AI systems learn. AI systems learn to represent and respond to any number of features of situations, and to transfer that knowledge onto novel states. Certainly, there are some concerns that AI agents may perform suboptimally outside of their training distribution. But nobody thinks that AI agents randomly assign value to states that they have never seen before. That assumption, which Krakovna and Kramar term Equiprobable Training-Consistent Reward, would be tantamount to saying that training is entirely useless outside of the agent’s training distribution, and all machine learning specialists should be fired because they haven’t managed to teach agents anything more than how to behave during training.
All of this is to say that the indifference assumption embedded in Equiprobable Training-Consistent Reward is not plausible. While Krakovna and Kramar are to be commended for attempting to provide a model of how reward functions evolve during training, their model is not a good one. Equiprobable Training-Consistent Reward is not a model of learning, but a denial of learning. And that is not consistent with our best scientific understanding of how agents learn.
5. Taking stock
This post considered a second shutdown theorem, due to Victoria Krakovna and Janos Kramar. We saw that this theorem makes a generous assumption and an ungenerous assumption. The generous assumption (Training Optimality) is that agents learn reward functions which perform optimally during training. The ungenerous assumption (Equiprobable Training-Consistent Reward) is that they learn nothing else, randomly assigning reward to states outside of their training distribution.
We saw that Equiprobable Training-Consistent Reward would ground an inference to any number of poor behaviors, including a particularly tight argument for the likelihood of shutdown-aversion. However, we also saw that Equiprobable Training-Consistent Reward is not very plausible. As a result, we should not be significantly moved by Krakovna and Kramar’s theorem towards concern about shutdown-averse behavior.
This concludes my examination of informal (Part 2) and formal (Part 3 and Part 4) arguments for Catastrophic Shutdown Difficulty. We have seen that a range of leading formal and informal arguments face challenges. The next post in this series asks what follows from this conclusion.

Leave a Reply