In this post I present a sketch of a grounding for morality that is human and active. It can be read standalone, or as part of a four-part series discussing the proper relation of succession—the handing away of power—to how humanity should deal with the coming of superintelligent AI. It is necessarily a sketch rather than a rigorous proof of every last point, intended to orient toward some important and often-neglected moral dimensions.
Successionism, as defined in the first part, is the ideology that says that a fundamental transfer of power (and maybe even experience itself) from humans to AIs is the right way to deal with superintelligent AI. The strongest of their arguments is that it may be very hard to avoid.
But as anyone not posturing for a bit knows, “is” doesn’t make an “ought” and might doesn’t make right. So are the successionists right, morally speaking?
I regret to inform you that answering this requires metaethics.
For successionism to be right, it must be possible to confidently divorce moral value from humans. However, human felt experience is the only sure window we have into the territory of which moral theories are the map. The meaningful transfer of that felt experience seems hard to do and be sure of, which is a reason for a strong precautionary principle. Not only are humans in the abstract indispensable to human morality but so is actively weighing options on your own and making choices and undergoing value change. This, and ghosts of history we should heed, argue that the active participation of the human individual is critical to our morality too.
Morality comes from humans
The fragility and slipperiness of moral value discussed previously often leads to one of two responses.
Sometimes the response is a mental panic that forces a retreat to the familiar solid terrain of equations and objectivity and abstract principles. I think this is behind answers like “the ultimate good is complexity” or “the fundamental value is intelligence”, or neo-Pythagoreanism.
The other response is that it’s too subjective: sorry, we just can’t say anything about values, anything goes and all is arbitrary.
I claim: much like your eyes give you access to evidence about the physical world we live in, your felt moral intuitions give you access to a moral world that your sense of “ought” is unavoidably based on.
The inescapability of moral intuition
Humans clearly have moral intuitions: you think that being in a nice forest, or being in love, or making someone smile are good, and you think being sick or being treated unfairly are bad.
However, it might feel reductive to argue that this is what morality is grounded in. There are two ways to disagree:
There is something “higher” that instead defines morality.
Most prominently this is the religious view. As I am already trying to derive the correct metaethical grounding of moral theories in one blog post, to keep scope reasonable I have elected to not also settle the question of God’s existence right here. I think the arguments below work without a God, but if you think that God is a faster path to similar conclusions, or explains why we have access to these moral intuitions, go ahead.
In these dark twilight years after Nietzsche delivered God’s eulogy or whatever, there are fundamentally two remaining types of things you can appeal to: properties of the universe (”facts”) and properties of your experience (”feelings”, “felt experience”, “intuition”, etc.).
The most pernicious form of the first type is some form of might-makes-right, such as the idea that evolution or thermodynamics necessarily point toward what is right. The best argument for this is that natural selection, and if not natural selection then at least thermodynamics, does in fact eventually win. This of course is ridiculous as an argument about an “ought”: if tomorrow we learned that entropy were reversible, or Moloch were slain forever and natural selection replaced with some artificial selection, there would be no change to the definition of right or wrong. The impulse to identify your morals with the winning physical principle comes from a combination of two factors: first, the laziness of wanting to win definitionally rather than on merit, and second (in the West), from a lingering neo-paganized Christian identification of the set-in-stone trajectory of the universe being one and the same with the source of its morality.
Once you’ve ruled out the “facts”—the idea that somewhere in the equations of electromagnetism, or written into the sands of Mars, or otherwise out in the physical world, there is the answer to what you should consider right—what you’re left with as a basis for morality is that it must relate to something about your felt moral intuitions, or at least something that lives in your experience of the world rather than the external world itself. You can argue about what aspect of the human mental world it is, or what values it implies, but that is where you end up.
Alternatively, even these things don’t matter and it’s all arbitrary. But you can’t actually remove the fact that you like some things and dislike other things, and feel compelled to pursue some ends and not others. You can disbelieve and be disappointed by the physical world as much as you want, or try to argue it away as your senses deluding you, but you are still stuck inside it. Similarly, your wants and preferences and instincts toward right and wrong are still there and still part of you however many edgy moral statements you try to endorse. You, as a human being in this world, cannot escape having a normative stance. And by virtue of it being your normative stance, i.e. your fundamental source of “ought”, i.e. the reason why you do or prefer anything at all, it cannot be insignificant to you. It must be the most essential thing there is, regardless of how vague or unsatisfying the nature of its source seems, much like a star cannot help but be bright and beautiful in a sky that is otherwise dark.
Note that this does not tell us what morality says, just like the fact that our senses are our only source of evidence about the physical world does not deliver us general relativity. Making moral decisions requires figuring out what the instincts imply, not just that they come from us. The first-order theory you could have here is that the obviously-instinctively-good things are the only good things and similarly for the bad things, and this is in essence Bentham’s utilitarianism. But the grounding of morality in human intuition does not deny more sophisticated moral theories any more than seeing a flat horizon every day denies a round Earth. (Later, I will mention several other properties required of a human moral theory that go against Bentham, and especially against the Benthamite momentum toward a view of value that ignores the valuer.)
We are the territory
In science, we try to make the “map” (our theories) fit the “territory” (the world), but at least the territory is always there. But when it comes to values and morals, because we are the territory, we could throw ourselves out.
We are the territory in two different ways:
One of the clear things that human morality says we should care about is human experience existing and being good (including that of others, though I will not make all the obvious arguments for a strong level of altruism here). Humans are the territory in the sense that they are “real“; their felt experience matters and is part of the set of terminally-salient moral objects of the world. Thus deleting humans who feel things does to the project of morality what deleting the world would do to science. (This is the type of moral grounding in humanity that experiential successionism argues against.)
The ability of humans to go around and experience things and have moral intuitions about the things they have experienced is our fundamental source of evidence for what the right moral theories are. We are the ones who can see the territory and walk our moral world, and therefore without us there is no more accumulation of new evidence about its shape. Thus deleting humans from observing things and having opinions and making choices would do to the project of morality what deleting observation of the world would do to science. (This is the type of moral grounding in humanity that control successionism argues against.)
But could these properties of humans not be transferred?
Consciousness is confusing
Could some other being have access to this same felt conscious experience that we consider valuable? To do that, they would have to at least be conscious. The importance of consciousness to moral value is hard to get around (though some successionists try to deny consciousness is real). A universe of rocks and dust, without any conscious observers, doesn’t seem like a place where there’s any reason for one configuration of rocks to be preferred to another. The thing that most moves us about animals, for example, is evidence of conscious experience in them. (Is consciousness, of any type or quality or degree, sufficient? This depends on what you mean by consciousness, and what exactly our felt moral intuitions say, though I suspect we all lean toward yes for some sensible definition of the terms. Here I argue simply for the weaker but sufficient claim that even verification of consciousness is hard.)
This dependence of morality on consciousness is a profoundly inconvenient fact.
In any given field of human understanding, progress is usually driven by some engine of verification. In math, we have formal proof. In science we have the prediction of observation. In ethics, as discussed, the engine of verification eventually grounds in moral intuition. But consciousness by its nature precludes direct access to the ground truth of anything but our own consciousness.
It seems like the best we can then do is to start from ourselves, and then reason about how similar various other things are. I am extremely certain other humans are conscious, because I know they’re built like me, act like me, and so on. Now go to chimpanzees: they’re quite similar, but they’re missing some things, like language and some brain size. Probably they have some consciousness. But what’s the experiment that confirms it? And consciousness is almost certainly not binary—but is it a scalar, or some sort of other structure? Are there different types?
Now, consider: dolphins, dogs, octopuses, shrimp, Claude Sonnet 3.6, a cactus, the economy, FedEx, the concept of the number 5. I think it’s much more likely it feels like something to be a dolphin than that it feels like something to be a shrimp, for example. But I can’t say much else.
That’s not very satisfying. This points, strongly, to a precautionary principle. Our actions should remain non-catastrophic across a wide range of assumptions about what makes something conscious. So yes, don’t torture the chickens. Also don’t torture the AIs. And, more than anything else, do not remove humans from the picture, because what if it is just humans, or things that are human-like in a very specific way?
If you strongly think some other conclusion is correct, I think you have insufficient epistemic humility. Consider how many times people have been wrong about questions with lower stakes and that are way less philosophically messed up. I cannot emphasize how cursed this entire area is to reason about. Every time I have a conversation about consciousness with people who (unlike myself) have properly thought about it, I hear about some new thought experiment that makes me feel like I’m in a Lovecraft story. The last one involved homomorphic encryption of an uploaded dog. The dog’s name was Fido. I don’t think he was having a good time.
Morality is an active process
So far we have talked about the difficulty of extracting the human from the moral. But there is another important axis, that even lots of well-meaning humanists miss, the recognition of which cuts out many successionism-flavored ideas.
The fundamental unit in morality is not “value” or “utility” as a homogenous substance that lives in minds and which the world should be structured to extract like an oil rig pumping oil. Instead, the fundamental unit is the active process of the mind of an individual experiencing, valuing, flourishing, and growing.
J. S. Mill wrote in On Liberty:
Human nature is not a machine to be built after a model, and set to do exactly the work prescribed for it, but a tree, which requires to grow and develop itself on all sides, according to the tendency of the inward forces which make it a living thing.
A tree wants for water, but the ideal tree is not a puddle. Water should flow in the xylem of the tree, but that it should do so is a “should” for the sake of the tree. Likewise we care about happiness and value and utility, but only so that they flow through and carve and build up a person.
The moral black box
One reason to think the sense of value only makes sense given an individual is that we seem to need the individual around to make value decisions.
Let’s set aside, for now, the whole consciousness thing. Just assume for example that verbal reports are a good proxy of moral judgment and felt experience. This is a lot of simplification, but there’s still the problem that we don’t actually know how our brains make value judgements. We cannot write down an algorithm that takes a description and knows very accurately what value judgment a given human would make in that circumstance.
For some intuition pumps on this, consider:
We want lots of sensory information to make judgments. People don’t like renting apartments they haven’t toured in person even if photos or a 3D tour are available. Why? Smell, sound, “vibes”. Also: people don’t like hiring people or making deals with people they haven’t met in person; trust is harder to establish.
Consider how ineffective it is to describe a piece of music or a work of literature, versus experience it. It is very hard to compress the quality of things into words.
Many people over-focus on abstracted moral dilemmas when talking about values, presumably because they’re more flashy and it’s easier to write philosophy papers about them. But even here: consider how hard it is to pin down when exactly autonomy violations are fine, or it is justifiable to wage a war (even if you’re a blanket libertarian or blanket pacifist, try giving a mathematically-rigorous definition of “autonomy violation” or “defensive war”).
Whoever goes around saying “behold, for I have figured out the full shape of what is right and wrong” has not actually done it. As the saying goes: never ask a man his salary, never ask a utilitarian to write down their utility function, and never ask a Kantian deontologist what to do when the Gestapo knocks on the door and asks if you’re hiding anyone.
If you can’t write down the algorithm, in a way that doesn’t involve asking the individual to consult the black box (from the perspective of the algorithm) of their moral instincts, then you need the individual to keep choosing.
Moral theories cannot fit perfectly
Alas, making moral choices is uncomfortable. Uncertainty sucks, and when it’s a moral question, not only is it draining but if you get it wrong there’s the guilt of maybe being a bad person too.
One way to avoid having to constantly judge and make choices would be to just make a few judgments, and then fit a theory to those that predicts the other judgments you might make. Why can’t you just do that? Because that theory is not the territory. The good scientist never stops making measurements, because the world is big and contains more than you think. A great scientist might like theories but they must love reality more. Empirically, the human moral world is also big and contains more than you think.
I don’t deny that you could train a pretty good predictor of human, or a specific human’s, moral judgments. Obviously LLMs agree quite well with humans about abstractly-described scenarios, and this is useful and should give us hope that we won’t necessarily miserably fail on alignment in every possible way. But as the whole point of science is the frontier where we’re confused, the whole point of the human moral project—in the sense that includes you expending judgment to figure out what you think is right day-to-day in your particular life—is the cases that require judgment and thought. Purely based on trends so far, you should not be hopeful of a final answer; the number of moral questions or amount of judgment required to get them right does not seem to be decreasing over time!
There is also a reason to think the frontier will always remain. Ultimately, moral evidence bottoms out in human felt experience. If the ground truth of morality is tied to intuitions in your brain, then to pull on those intuitions and let them fight it out in your mind is the bedrock of moral deliberation. A predictor of the results of the battle in your conscious mind that is not itself both conscious and accessing the same conscious experiences with which you weight things cannot generate new evidence, only fit existing evidence. The infinite creation of new circumstances, and the changing of the individual doing the judging through their experiences, means that the realization of the will of the individual can always be more perfectly achieved with the active participation of that individual. We cannot free ourselves from the agony of choice.
If you want things to go right, by whatever your particular lights of rightness in your moral world are, don’t throw yourself out. Stay involved: keep judging, valuing, deciding. Therefore, even if AI did all the work for us, humans should still be making value judgements themselves. However weird the future is otherwise, I want many someones to be walking around and looking at the world with their own eyes and having takes about whether it is good or bad.
Value change is fundamental
So: you must keep going through the agony of choice because that process is how you access the moral intuitions that everything else is built on, and there is enough complexity to your moral world that it can’t be fully captured without these continuous checks against ground truth.
But not only are your values complex, they’re also changing. In her book Aspiration, philosopher Agnes Callard argues that value change is a core part of how human values work.
Imagine you want a kid. You could phrase this as a fixed preference: you want to have a kid, and this preference becomes fulfilled when the kid is born. But that’s obviously wrong. After you’ve had a kid, you value that kid in particular. If someone offered you to switch that one for a different kid of higher utility, you would obviously refuse. You don’t know who this kid is before they’re born, and the kid keeps changing over time, so there’s clearly not some platonic ideal of “I love this exact child” that exists in your head before the kid is born, that is then fulfilled by it. Rather than you having fixed values that are fulfilled by having a kid, the far more natural description is that you start out with an inkling that there’s something very valuable in the direction of having a kid, and the value of loving that kid in particular is something that grows within you over time as you learn who that kid is.
Callard has other examples as well. Contra Faggella, when people fall in love with their partner, they love that particularperson, not just the bundle of “fulfill[ed] drives” and “good feelings” they cause. But before you met your partner, of course, you didn’t know what they were like—the fact that you value them is a change in your values compared to before. What you love is not a bundle of your static unchanging needs being met, but a specific person.
“Ah”, the successionist might say, “but these are just special edge cases that humans are weird about for obvious evolutionary reasons.” I don’t know about that, they seem like pretty core parts of the human experience!
“Okay, but if you really think about it, what you value isn’t the kid or your partner, but the things they make you feel, and those are constant regardless of the kid or partner”. Interesting relationship to your children & partner that you have there, but okay: let’s take Callard’s default example of learning to appreciate classical music. The person who walks into a music appreciation class literally cannot feel what they later learn to feel when they get really into classical music. The inner experience of listening to and liking, say, Chopin, is different to that of listening to and liking Bach. As I’m sure anyone would tell you, the texture of the feeling is different; the embedding vector for the two is different—pick your metaphor of choice. And what about the inner experience of being a believing Christian, or an enlightened monk, or a successful entrepreneur, or a hunter-gatherer? Are these all really just pulling on the same few basic emotional levers in interchangeable ways to each other, forming different linear mixtures of pleasure / satisfaction / joy? Or does each of these involve a different inner world, made of its own fabrics and own bricks? If you go from one to another—and remember that everyone experiences something comparable, whether growing up, loving, having kids, changing careers, changing worldviews—do you not clearly change your values?
To distill the argument:
The most straightforward model is that your preferences are over functions of your sensory perception. But clearly, the value to you of seeing your partner before you fell in love with them is entirely different from seeing the exact same sight after you’re in love. So if preferences are defined like this, they obviously change with time and experience.
Next, you could say that your preferences are functions over your inner state. But as I hope I’ve persuaded you, you can learn to experience and like new inner states. So even preferences defined on inner state can—and, for humans, do—change over time.
Finally, to try to claim a fixed and unchanging basis for preferences, you could retreat to some more abstract notion of preference—that whatever inner states you value, there is some quality they have (”utility”, “being-preferred”, “potentia”, “arglebargleness”) that is constant and unchanging. There are some theoretical reasons around coherence properties such that it makes sense to talk about an implicit utility ranking implied by actions. It sure would be nice if one existed in our heads. But we don’t have evidence that this axis is something real that actually exists and is feltin human heads. It is a theoretical construct; at best a useful mathematical framework, at worst an epicycle.
Therefore: your values change, your values contain referents to things outside your brain and inner state, both of those are important parts of you being a human, and there is very likely no concrete single yardstick pointing towards the good in your head. So whatever the process within you is that can value things morally is a changing and subtle thing that cannot be “exported” as a yardstick of utility that some other being could mechanically go forth and optimize.
Have you considered that we live in a society?
In this post, I focus on valuing as something done by a single individual, and the necessity of the individual to that process. This is not to deny the role of society. Empirically, individuals greatly benefit from others in figuring out how to pursue the moral good, and much progress in valuing we make comes to us through culture and contact with others. I have touched on how to think about the social process of moral progress in The Technology of Liberalism and Paul Christiano touches on the value of humanity as largely coming from humans collectively thinking for a long time (as quoted in The Two Bars of Alignment), but there is much more to figure out and say here beyond the scope of this post.
Is this really morality?
Hold on, why all this stuff about aesthetic experiences and inner feelings? Is this all a bit woo, and shallower than “real morality”, which is about things like “don’t break promises” and “play ‘cooperate’ in prisoner’s dilemma”?
I think it’s useful to separate morality into two parts:
“Game theoretic morality”. It is well-known that tit-for-tat is a very robust algorithm to run in repeated games. Concepts akin to honesty, trustworthiness, and altruism tend to emerge naturally in environments that reward collaboration. People spend thousands of words edging ever-closer to the uncrossable is/ought line starting from these arguments. Beren Millidge gave an excellent talk on how ecosystems of competing AI systems might recover these principles (at least with respect to each other, if not humans).
“Experiential morality”. The goodness or badness of your mental state that I’ve been talking about here.
Outside philosophy, most human talk about morality is about the former, because game-theoretic-morality cooperation principles are very important for dealing with day-to-day life. You probably consider others’ trustworthiness and your commitments to others many times a day. In contrast, you rarely need to question whether the entity you’re interacting with has inner experience. However, with AI we are heading into a future where that question will get a lot more confusing and a lot more uncertain (and that uncertainty is unlikely to abate soon, as argued above).
The obvious necessity of the experiential part of thinking about morality is shown by the thought experiment of imagining a universe containing nothing but trillions of commitment-honoring, altruistic, perfectly-cooperating non-conscious automata. There is nothing in such a universe more valuable than a single minute of a couple’s felt experience on a good date.
The individual as the bulwark of morality
I’ve argued moral value is:
Grounded in human felt moral intuition.
Tied to consciousness, perhaps the most philosophically confusing thing in the world.
Not specifiable in a closed-form algorithm that can be carried out in the absence of the individual.
Driven and tied to subtle and constantly-changing processes inside individual human brains.
Therefore, it’s a deeply fragile, subtle, and confusing thing. None of this diminishes its realness. But it means that values are like fish: slippery enough that if you try to grasp one directly you probably lose it. But conveniently, values come in buckets: the individual human. Values swim inside their brain. We have millennia of experience dealing with individuals. Even when we don’t know how to nourish an abstract value, we mostly know how to nourish an individual. The best way we know of making moral progress is to have a bunch of humans around, experiencing and thinking and living. Like fishermen dealing in buckets of fish rather than individual fish, the thing that works is not trying to grasp values directly and lift them up, but to deal with the familiar, dear-to-us actual human individuals whose heads the slippery values live in.
Many moral atrocities are downstream of placing value outside the individual. Historically, value was often placed outside the individual. This resulted, for example, in fascist dictatorships sacrificing real individuals for the glory of the fatherland, or communist dictatorships doing the same for the cause of the proletariat. Today we increasingly reach inside the individual and try to deal with subcomponents directly. This can be instrumental, as when Big Tech invests billions into hijacking your limbic system over the protestations of your higher self. It can also be well-intentioned philosophical momentum, as when some utilitarians speculate about tiling the universe with pleasure circuits, because they see the pleasure itself, rather than the individual surrounding it, as the thing that counts.
The surest way to avoid all of this is to retain the focus of moral concern at the individual. It is the proper functioning of the entire individual’s mind that results in access to moral intuitions. It is the entire mind, not just the pleasure circuits in it, that are able to participate in Callardian aspiration and value-change.
When they come in and try to break the individual—whoever “they” are this time and whatever justification they come with—resist! That is perhaps the one simple recipe that, if followed, would have cut out the most historical horror.
So: don’t invent god-concepts and then sacrifice everyone you love to them. Also don’t try too hard to reach underneath the shell of the individual and strip-mine whatever value-fluid you think exists beneath. Instead, help human individuals. Let them flourish and live. Let them explore. Let them gestate new things in their mind as they change and grow. Let them make choices and judge things, and let those choices and judgments have power and let them feel the effects, and let them learn from this interaction with each other and the world and themselves.
Thanks to Xavi Costafreda-Fu, Aniket Chakravorty, Luke Drago, @softminus, Elsie Jang, Yudhi Kumar, and Oak Hu for feedback.



Interesting post, I found the "we are the territory" point and the idea that morality is an active process compelling.