So far I have argued:
The ideology of succession—that humans, either entirely or at least in their role as decision-makers, should be replaced by AI—is driven by cultural factors including (a) worship of mathematical abstraction, (b) bureaucratic safetyism stamping out license for human agency in favor of rule by procedure & algorithm, (c) a cuckoldry-adjacent simping towards the unlimited domineering power of superintelligence, and (d) the existence of the city of San Francisco.
The two bars of alignment are (1) AI not going rogue and (2) AI instantiating utopia. If you agree with Eliezer Yudkowsky on the totalizing nature of superintelligence, these are of comparable (and near-infinite) difficulty, because superintelligence rearranges the atoms into exactly what it wants so it better know how to build utopia because otherwise you’re dead. But the field, even when not endorsing Yudkowsky, often defaults to treating succession (in the broad sense of fundamental transfer of power) as the only viable solution to the AGI transition. Everyone (including Yudkowsky) is very scared by this. This is because the succession-shapedness of alignment is exactly what makes it hard, because value is complex and therefore personnel is policy and hence our desired polices aren’t policy after superintelligence is the only personnel running things.
Morality lives in the human individual because the ground truth of our morality is tied to humans and their consciousness, and we don’t know what makes something conscious or even how humans make value judgments. This is a terrible and cursed philosophical problem and we should adopt a precautionary principle, which rules out both mistreating AIs and successionism. Also, since moral valuing is active process engaged in by humans with changing values, we should be very suspicious of putting “value” or utility in and of themselves on a pedestal, and instead advance them indirectly by supporting the flourishing and freedom of human individuals. Crushing individuals for the sake of some more abstracted thing is perhaps the classic moral blunder in history, so don’t do it. Keep human individuals around instead, treat them well, and anchor decision-making and judgment to them.
There are three remaining questions related to successionism and its intersection with alignment and the philosophy of the AI transition that I want to discuss:
Is advocating for humanity “speciesism”?
Clearly, we need some successors (e.g. our children). How do we decide which successors are good?
What maxims help guide us toward a future that is good in light of the philosophical currents painted in this series?
Will you get cancelled for being a speciesist?
Might an anti-successionist stance make us commit moral atrocities with regard to sentient AIs? Obviously, being opposed to humans being pets of AIs or being done away with all together does not mean you cannot care about treating AIs well too. But perhaps we are morally obligated to surrender to the AIs? Maybe discriminating against AIs or even just keeping them down by not letting them conquer the world will later be seen as comparable to historical discrimination against minority groups?
Total universalism is impossible
Consider this argument from Nina Panickssery:
Instead [of being based on a “definition of value so arbitrary as to stipulate that biological meat-humans are necessary”], my opposition to AI successionism comes from a preference toward my own kind. This is hardwired in me from biology. I prefer my family members to randomly sampled people with similar traits. I would certainly not elect to sterilize or kill my family members so that they could be replaced with smarter, kinder, happier people. The problem with successionist philosophies is that they deny this preference altogether. It’s not as if they are saying “the end to humanity is completely inevitable, at least these other AI beings will continue existing,” which I would understand. Instead, they are saying we should be happy with and choose the path of human extinction and replacement with “superior” beings.
Famous altruists from Peter Singer to Jesus Christ would likely have some pushback against a naked preference for your own kind. But the argument doesn’t have to rest on pure preference, and resting on human preferences is a necessary feature of human morality, not a bias.
First, for the foreseeable future, pure epistemic uncertainty about AI consciousness implies resisting successionism.
Second, while we should be altruistic in who we have concern for, the nature of what we have concern for is necessarily based on our human ethics, and there is no way to delete that human component. As discussed in the previous post on how morality lives in human individuals, our morality is grounded in intuitions that we have access to because we are humans. Your morality is a view from somewhere. In The Human Prejudice, philosopher Bernard Williams defends a pro-human “prejudice” in ethics, in particular arguing that the view advocating for a universal non-human yardstick smuggles in a “view from nowhere” that doesn’t exist.
If there is no view from nowhere then things have to bottom out in “we are humans and so have this take” rather than something even intelligent non-humans with different intuitions must agree to. Is universalism necessary in moral principles, though? One function we want of applied morality is arguments that make sense regardless of who you are. That you should pay me more for that carton of eggs is in my self-interest, and an argument unlikely to move you. But that you shouldn’t defraud me by taking my cash and running away without giving me the eggs is a moral principle, that has both practical appeal to both of us (we’d both rather live in a world where people at large acted in accordance to that rule), and an appeal to a higher principle above self-interest. But remember Hume’s fork: you cannot derive an ought from an is, so no moral argument can be universally compelling, since at some point there is the starting ought. The yearning for perfect universalism can never be reached, however useful universalism may often be for game-theoretic morality.
The moral road
The yearning for ever-increasing universality is summarized by the concept of the moral circle. The moral circle encourages us to imagine concentric circles—us, our family, our friends, our country, our species, our fellow mammals, cephalized animals—with moral progress consisting of pushing out our concern further and further out. Clearly, this becomes ridiculous if taken as an axiom. Today, an enlightened liberal may care about the suffering of pigs, but should tomorrow’s extra-enlightened ones #raiseAwareness about the plight of jellyfish, or their turbo-lib descendants worry about chairs being sat on? Broad-mindedness is only a virtue to the extent that it is encompassing something worthwhile. Noticing that the slave is a person who can feel and choose and participate in the glory of the human life just like we can is not about pushing out a boundary to conquer more moral territory for the sake of it, it is about protecting what, given information and honesty, we already care about. Where concern is stopped by distance or ignorance or harmful self-interest, those boundaries should fall. But the concern exists because we have something in particular to protect.
One alternative to the moral circle is a moral compass. Rather than expansion, we have a specific direction to go in. Just check the compass and see where it points. The thing to protect is the happiness of conscious beings, say (and at this point you run into lots of difficulties with definitions & being sure, as discussed, but I’ll ignore that shape of concern here). So maximize that!
If we know the thing to maximize, likely humans are not the most efficient arrangement of atoms for achieving it. Utilitarians talk about the utility monster, a being capable of vastly more pleasure and suffering than anyone else. You eating a cookie is suboptimal, because the utility monster gets way more joy out of every cookie crumb than you could get out of all the cookies in the world, so all cookies to the utility monster. Take this view to its limit, and shouldn’t we engineer utility-monster AIs, give it all to them, and kill ourselves so that every last joule of energy can be diverted from running our silly little monkey brains toward being spent on the Great Pan-Galactic Claude Logos 9000 Hyper-Orgy at the end of time?
But remember that morality is an active process, we can’t describe it in an algorithm, our values are changing, and personnel is policy. “Head due north”, sure, but there’s no magnetic true north to follow, only our own inner sense. A utility monster can only exist with respect to a yardstick, but a definable fixed yardstick is what we don’t have. Be humble: do not presume you know what you yourself want, since getting there is very rarely a straight-shot to something whose description you could pull out of your head right now, and very often an iterative process with much observation and learning and changing along the way.
Think of it as a moral road that we’re on, that we build as we go. Progress means taking steps to somewhere, not just broadening the empire of our concern. We make progress—the virtues and beings and experiences we care about go further with every step—but unlike with the compass, we must in fact go out into the terrain to see where the road should be laid, since that choice comes from us rather than being a cardinal direction set from above. The only utility monster worth respecting is not something that stands aside our moral road and demands us to swerve toward it, but something that emerges out of us from progress on the moral road we’re on. The only way we could possibly knowably arrive at the fountain of infinite goodness (if such a thing can even exist), as measured by our own lights (but which other lights should we care about?), is by ourselves continuing to step forward on our particular moral road until we get there.
We will have many fellow travelers on this moral road (at least the entire rest of humanity; see below for general thoughts on determining when successors are traveling the same moral road). Progress on it will mean extending concern to many: for example, a rich civilization should clearly be good to animals, and out of precaution and eventually perhaps evidence of moral patienthood, we should treat AIs well. But progress on our moral road depends on us humans remaining on it and deciding where it, and the bulk of Earth-originating civilization, goes.
Treat AIs well
One of my big hopes for humanity is that we’re the kind of species that is nice to aliens.
The aliens may be conscious and may feel, and this should obligate us to some level of concern toward them. This is despite the fact they won’t be on the same moral road as us and therefore we should of course continue charting our own path and not let them get decisive power over us. Likewise for AIs that really do seem like moral patients (though note that biological aliens seem more likely to be moral patients than AIs).
We often prioritize our country over foreign countries. This makes sense: if we identify with our country we probably feel that on net it has more of the right stuff in it and thus the world sculpted by our country and the cultural lineages spawned by it are more agreeable to our sense of how the world should go. But that is still no reason to go warpath or not care about other countries. The aliens, if nice and moral-patient-like, should slot in as an ultra-foreign country.
As for the AIs, I think we should not treat them as their own country. Instead, we should try and integrate them into our society in symbiotic ways, with roles different from humans but with whatever dignities that both our and their sense of dignity requires (we should be able to do this, since if we cannot engineer them to not want our roles and titles and power, we have a very large alignment & loss-of-control problem on our hands). To degrade a thing that talks like a human cannot be good for your morals. Also, by treating them well but not treating them as our successors, we have more room to appreciate them as they are. Scott Alexander wrote:
We can debate forever - we may very well be debating forever - whether AI really means anything it says in any deep sense. But regardless of whether it’s meaningful, it’s fascinating, the work of a bizarre and beautiful new lifeform. I’m not making any claims about their consciousness or moral worth. Butterflies probably don’t have much consciousness or moral worth, but are bizarre and beautiful lifeforms nonetheless. Maybe Moltbook will help people who previously only encountered LinkedInslop see AIs from a new perspective.
The world is full of wonders. Some of those wonders, like tigers or misaligned powerful AIs or gamma ray bursts, might kill us. Not confusing the people and the wonders goes better for us, and also lets us appreciate the wonders for what they are. Respect the butterflies, but still have children.
Have precautionary but not overriding concern; avoid succession
To sum up: there is no morality we could have that isn’t in some sense partial to and based on humans. Of course, the development of human morality may involve great concern toward things that aren’t human, though precautionary arguments based on uncertainty about radically new beings like AIs, and the possibility that our morality is partial to some pretty specific properties of humanity, mean that the core of our portfolio of concern should stay centered on humans for a long time. This all follows even without any selfish preference-for-your-kind that is separate from the unavoidable and entirely proper partiality of having some moral view at all. The transfer of power or the moral development process is much more sensitive still. We can’t do it without extremely tightly-scoped successors, if we are to not lose progress toward our moral ends.
But clearly, some successors are good.
Which successors are good?
Here’s a good type of successor: actual children.
But what counts as a child? Robin Hanson approvingly cites Hans Moravec, who already in the 1990s was talking about AI as our “mind children”. Children, of course, are worthwhile successors. Hanson and Moravec would like us to think that including AIs among our worthy successors is an obvious expansion of the moral circle. Hanson argues that since they inherit our culture and are most likely to continue our tradition of abstract thought (which as a neo-Pythagorean he cares a lot about, and where he is not optimistic about the human trends), they are literally evolutionarily equivalent to (cultural) descendants and therefore we ought to treat them as our kids.
This comparison is not a moral circle expansion but a moral gerrymander. The similarity of the brains of your biological children to you means that you should be far more confident that the experiences, personhood, and development that you value live in and can be driven forward by their brains, compared to something built very differently. Handing over to the AIs is not another step on our moral road, but a jump onto moral rails not steered by our moral intuitions since these intuitions cannot be perfectly capture and exported. And of course, it takes neither deep introspection of the human condition nor a deep understanding of evolution to realize that literal children are extremely special and have a privileged position in human moral concerns.
Children, in fact, are such worthy successors that they seem like one of the only cases in which succession is actively good. Future generations should not live forever in the shadow of their parents. People complain about boomers trying to achieve this and, whether they’re right or not about the boomers in particular, it’s clearly possible for the past and the old to exert too much control on the present and the young. Of course, we will eventually have the technology to indefinitely extend lifespans, and we obviously should use this technology, and deal with the resulting generational lock-in problemsthrough social and political means rather than literally forcing older generations to die. Even in the case of good successors existing, there is no reason for genocide!
But clearly: there are gradations of kid-ness. One example: how should our ancestors from a million years ago feel about us? Or in the future: humans will continue changing. Surely there must be some arcs of change that end up very far from us that are still good. For example, I find it likely that our ancestors from a million years ago, even if they disagree directionally about the way the world has changed, would not think the world filled by us is morally empty, especially if they were shown every link in the chain between them and us.
If you think this, Robin Hanson thinks you might just be afraid of change:
If you are horrified by the prospect of greatly changed space or AI descendants, then maybe what you really dislike is change. For example, maybe you fear that changed descendants won’t be conscious, as you don’t know what features matter for consciousness. But notice that even if advanced AI never appears, your Earth descendants would still likely differ from you as much as you now differ from your distant ancestors. Which is a lot. And with accelerating rates of change, strange descendants might appear quite soon. Preventing this scenario likely requires strong global governance, which could go very wrong, and also strong slavery or mind-control of descendants, controls which could cripple their abilities to grow and adapt. I recommend against these.
The core missing concept is that we can build locally-valid chains of succession that are much surer than the big leaps. You know your kids are worthy successors. They know their kids are worthy successors. And so on.
An analogy might be: there are boxes, some which contain a flame inside them. The flame is invisible from outside the box, but you see the interior of your own box and know that it currently has a flame in it. There is a process by which you can create a new box and be sure a flame is lit within it (you may have learned of it in health class). Then, someone comes to you with a box, and points out all the ways it looks similar to the boxes that are part of the flame lineage - both are rectangular, and yes, sure, they’re made of metal rather than cardboard, but otherwise practically identical! “Therefore, this box has a flame”, they say. Well, maybe. You definitely wouldn’t want to take actions that assume there isn’t a flame in the other box—the fire safety people would rightly get mad. But you are far more certain about the boxes you see being produced by the flame-transfer process. If you want the flame to definitely keep burning, you should continue the known flame transfer process.
One issue is that locally-valid chains that aren’t quite perfect can accumulate error if they’re long. Let’s say that every generation there is subtle selection pressure against felt inner experience. You could have conscious humans evolve (or more likely, transform themselves technologically) into things that do not feel. More prosaically, elders may feel like important things are being lost—have you seen the youth these days? But this is a task we have to manage. Ensuring that key values, culture, and moral patienthood are passed on to the next generation, while we avoid lock-in and stagnation, is just one way to describe the job of civilization in general, not some specific new duty we might get to shirk because it’s annoying.
How far to dial the transhumanism?
A sub-question about succession is how much we should modify ourselves. Human modification for humanist ends is often called transhumanism.
Transhumanism is defined in many ways. Eliezer Yudkowsky calls it “simplified humanism”: the philosophy that life should be good for individual humans, without making exceptions like “unless it requires taking actions that go against the natural order”. The transhumanist says: “humans shouldn’t be killed”. An anti-transhumanist humanist-with-asterisks might say: “humans shouldn’t be killed by autocrats, tigers, or appendicitis, but it would be playing God to prevent humans over the age of 90 from being killed by heart failure”.
Meanwhile Francis Fukuyama warns that transhumanism is the most dangerous idea in the world, because it may fork the human lineage into many splinters that have less grounds to unite or cooperate around shared assumptions, for example about human rights. But liberalism is a powerful social technology for managing diversity, as Fukuyama should know, and lineage-forking is often good for reasons including helping cultural evolution, making civilization resilient, and simply being a consequence of letting people do what they want. I suggest that many people who worry about transhumanism are conflating it with posthumanism. No one has very precise definitions but, in line with the above discussion, I would draw the line at transhumanists wanting to liberally extend the human lineage (in a way that keeps it walking down the same moral road), whereas the posthumanists want to achieve something more like a clean cut or jump to a new thing, without the requisite locally-valid chains of succession. Successionism then is a specific type of posthumanism that advocates AI in particular as the type of posthuman lineage to transfer our concern and control to.
The above discussion of which successors are worthy then implies: transhumanism is fine, under the asterisk that given large or many jumps we might lose something and we of course need to be careful, but posthumanism is not.
And therefore, when I complain about successionism, I want to make it very clear that this does not rule out wild, diverse, greatly-different futures. I want humans to live as long as they want. I want humans to figure out how to make themselves smarter. I want a million cultures to bloom for a billion years. I want to eventually find myself shaking my head at the youth these days doing incomprehensible things I don’t understand. The human flame should live, but in the best futures it finds itself in bigger and more capable boxes.
Build the successor you want to have
It is also important to remember: we can choose our successors.
For example, regarding AI utility monsters as discussed above, there is an even simpler resolution to this philosophical puzzle. We are not in that scenario yet: we can just not build the utility monster from the famous utilitarian thought experiment about why it’s bad for us to build the utility monster.
If we let technology run its course while humans remain the ones steering it, eventually we can uplift humans in any imaginable way, if we so want to. No need to throw everything to AIs that we’re less sure about in the meantime.
Arguments about successionism being morally necessary are generally weakened by the fact that we are already here. Many theories of ethics include some notion of rights. The Native Americans did not have a moral obligation to surrender and exile themselves when presented with European technological superiority. Likewise with all present humans together facing technologically superior AIs.
We should exercise our ability to control the future to build and choose successors that carry on the human project. Doing so is the only way we keep progressing on our moral road and create a world that is good by our lights.
Which directions to steer in?
Along with the cultural drivers of successionism, successionist ideas are inflated in the discourse due to important missing moods in philosophy-of-AI. Here I list a few that help steer toward positive futures.
Mill over Bentham
Jeremy Bentham got a lot right: he was a key liberal and progressive reformer, with moral views sometimes centuries ahead of his time, for example on the rights of animals to not suffer or the rights of gay people. He was so ahead of this time because he was a utilitarian and thus cared deeply about wellbeing.
Bentham had his rough spots. “Nature has placed mankind under the governance of two sovereign masters, pain and pleasure. It is for them alone to point out what we ought to do”, he wrote, seemingly without irony. He also loved control, such as through the proposed Panopticon prison, or a variant he wanted to shuffle all of England’s poor into to maximize their productive labor. He also hated common law, since it grew organically without a centralized order, and wanted to replace it with a rationally-designed code called the Pannomion. He loved to sketch out the structure of such centralized structures in technocratic detail.
His program was in-tune with the other philosophical radicals of the time, and attracted supporters such as James Mill, who raised his own son John Stuart as a role model of the ideal philosopher. John Stuart Mill did become incredibly precocious, but also burned out in his 20s and ended with a more measured philosophical program.
The Benthamite insight into the moral importance of raw wellbeing is powerful and important. It is a large block in our modern morality and a big argument for rejecting things like war and slavery. For Bentham, utility is fundamentally a substance or a fluid, that can be mined like oil out of the ground. In contrast, Mill wrote: “[U]tility, or happiness [is] much too complex and indefinite an end to be sought except through the medium of various secondary ends”. Mill understood that we humans were not so blessed as to have a simple terminal goal, so we need to embrace the messy and shifting nature of our wants to be able to build the good. Even more fundamentally, Mill understood utility only makes sense in the context of a person. He wrote under the banner of “utilitarianism” that he had inherited from his father, but his philosophy was fundamentally an ethic of the glory of the human individual and its freedom and energy and development. Depending on your attitude toward “utilitarianism”, this is either a disproof of it, or its only true instantiation. The genius of Mill was striving for the good not through control and calculation but through the free action of human individuals making their own judgments and own way through life.
Every power should have an interest and every interest should have power
Imagine our goal is to keep fire around. Imagine that, if all fires are extinguished, we can’t get it back. Now consider two worlds. The first world has giant powerful robots with furnaces in their hearts. The second has campfires dotted around a plain crisscrossed by massive vehicles. In the first world, power lives very close to the flame, and so the flame is hard to put out without having to tackle and defeat all the centers of power in that world. In the second world, flames and power are far more distinct. Even if the job of the vehicles is tending to the fires, you need to make a much more careful argument about why the flames keep burning. The vehicles could always run off course and smother the flames, for example. And what is the link between control of the vehicles and control of the flames? How stable is that institution or control mechanism?
Power should live close to the moral stuff that gives the world meaning, because then that power by default protects the moral stuff. If humans are the source of power through their brains and muscles, it is hard for the world to flow into a state where the moral flame is totally out. Maybe some humans lose out. But even if, say, some group is wiped out in war, at least the aggressors are humans who also have moral value. The world may get worse and be unjust, but it is in little danger of becoming pointless. But if humans are not the source of power, then there is one thing in the world that is its powerful actors that cause things, and another thing which is the thing the world exists for the sake of. A lot of risk from advanced AI comes down to this separation between the worthy things and the working things. A lot of successionism is about trying to become okay with this, by claiming the ability to do the work implies the worthiness.
Ideally, then, every interest should have power: the beings that make up the moral worth of the world should have the power to protect themselves. This is basically decentralizing the problem of keeping the world from becoming pointless: every point of moral stuff defends itself, rather than being kept alive by some distinct unfeeling caretaker.
Also, every power should have an interest, in the sense that if some powerful entity exists, its power should feed through to human gain. Some of the horrors of our world are due to a partly-non-human power, such as a market or a bureaucracy, having power with which it shapes the world, without a human behind it to benefit. There are scary futures involving feedback loops in a robot economy without any meaningful human ownership, where huge power centers might emerge that accrue massive resources that do not feed through to any human. We could end up in a world of zombie processesambling forward with real material consequences, but with their intended human beneficiaries crazy or dead or utterly ignorant. We should make sure that, however opportunity and resources and rewards flow around the world, they flow between competing humans rather than dripping off the human game board entirely.
To avoid specification-gaming, avoid the specification
We don’t know how to specify exactly what is good. So how do we get it? (Top 10 scenarios that throw analytic philosophers into crisis—number 3 will shock you!)
Remember the boxes-with-flames analogy. It is often possible to create or spread a thing without knowing how. Your kid is a moral patient even if you don’t know what makes them one. To take other examples: you might not know how to define good writing, but you can sure make a list of demonstrations. You might not know what about the culture of a city or company is good, even after experiencing it, but you sure can hang around and hope it rubs off on you.
There is a view, common among modern STEM folks, that everything must route through a specification. Observations to analytically-tractable spec to generalizations is, after all, how the hard sciences work. But the world we live in just is one where that is not always the most effective method. Notice how AI started working when we skipped the part with the analytically-tractable spec.
This also suggests that if we want AIs to inherit our values, not just in the fundamental-moral-sense but also in the hard-to-pin down this-is-how-we-do-things-here cultural sense, you should think seriously about tacit knowledge transfer through things that look like participation and apprenticeship, not just specification or rules.
Judgment should stay with us
The one thing that you cannot hand off is the root process of judgment. Sure, a CEO delegates many things to their employees and the programmer to Claude Code. But for the highest-level choices, you need to be in the loop if you want to achieve what you want, because the only ground-truth source for what you want are the judgments in your own brain. And the goal, after all, is that human wants keep shaping the world, because we sure have some wants about how the world should look and the point of doing anything at all is to bring the world toward those wants.
So, what is the gift we’ll give to tomorrow? We could give up, and surrender to the machines, and let them make judgements and sculpt the world in our stead. Imagine your chain of ancestors looking on, all the way from the Paleolithic to today, and you saying to them: “Sorry folks, I know you survived plagues and war and famines so that I could have a chance at life, but look: have you been to SF? Have you seen the AI partners? Truly, Man is fallen, and Silicon is god. We’re calling it, we’re giving up, we’re taking our hands off the wheel. It’s just going to be machines from here on out.”
Alternatively, the human journey could continue, however weird and wondrous it gets. If so, one day our descendants will look back on us, and the gulf between us and them will be at least as large as between us and hunter-gatherers, or maybe even us and chimpanzees. But there will still be children, and there will still be song, and there will still be young couples in love. And they will make choices and those choices will matter and life will keep changing and improving as the world is sculpted by human wills.
Thanks to Aniket Chakravorty, Xavi Costafreda-Fu, Luke Drago, @softminus, Elsie Jang, Yudhi Kumar, and Oak Hu for feedback.


