In April this year, William MacAskill published an article describing Saturationism - a proposed modification to Utilitarianism that resolves most of its more unintuitive or unacceptable inferences, including the famous “repugnant conclusion”. This is a great paper that I think that anyone with an interest in moral philosophy should check out.
MacAskill’s Saturationism proposes that moral value comes from “illuminating” a vast, abstract landscape containing all possible mind-types and experiences. When a conscious entity exists, it shines a moral light over its specific location in this landscape based on how good its life is. However, adding more entities with very similar mind-types or identical experiences shines light on an area that is already brightly lit, reaching a point of “saturation” where extra duplicates provide diminishing marginal value. Conversely, bringing into existence entities with novel, diverse mind-types lights up previously dark regions of the landscape, generating significant new moral value. In this way, Saturationism modifies traditional utilitarianism so that creating vast numbers of nearly identical, simple lives yields diminishing returns, naturally favoring a world rich in diversity and varied experiences.
Importantly, MacAskill clarifies that welfare itself is not a part of mind-type:
It’s not merely that by taking someone’s life and changing their wellbeing you increase variety (and can thereby make things better); you also have to be changing the nature of their lives in some meaningful way, to increase variety somehow.
One immediate “practical” application of this kind of framework, is that effective altruism initiatives frequently prioritise high-volume invertebrate welfare interventions (e.g. prawns) based on population estimates. While invertebrate suffering represents a legitimate ethical concern, similarity discounting reframes the scale of this moral obligation. It is a reasonable assumption that the less advanced an intelligence, the more limited its experiences can be, and therefore the more similar their lives are likely to be. This suggests that the population of trillions of prawns should at least be somewhat discounted due to their similarity.
Practical applications aside, this article was very interesting to me in particular, because the idea of similarity-based discounting is something I have explored previously. I even wrote a post about it 6 years ago, so it is fascinating to see the subject being taken seriously by a professional philosopher.
My post Dedomic Utilitarianism proposed grounding axiological value in the synthesis of welfare maximization and information preservation. Under this formulation, conscious entities are modeled as high-dimensional informational structures; specifically, neural network weight matrices or genetic and experiential datasets. Death represents irreversible data loss, whereas creating additional entities yields diminishing marginal moral value if those entities merely duplicate existing informational states. Dedomic Utilitarianism posits that while minimal redundancy provides value by buffering against catastrophic data loss, systemic duplication yields redundant data with decaying axiological value.
MacAskill builds a model for similarity-based discounting using the concept of “mind-types” and the language of differential geometry, whereas my approach was more tethered to information theory, so my first thought was to bridge the gap. Can information theory add anything useful to MacAskill’s model?
An Information Theoretic Approach to Mind-Types and Type-Space
I think the answer is “yes”, but the specifics are a bit jargon heavy and not really essential for the other sections, so skip this section unless you’re really keen!
In MacAskill’s formulation, a mind-type is a point in type-space - an abstract metric space equipped with an as-yet unspecified distance metric. In Dedomic Utilitarianism, a conscious entity is represented as a high-dimensional informational state detailing their particular cognitive configuration, for example a neural network weight/activation matrix. These are describing the same underlying phenomenon, and we can bridge the two by formalising type-space as a very high-dimensional parameter space (such as the space of all possible neural network weight matrices), allowing each point to represent this particular cognitive configuration.
A mind can then be viewed as the resultant (exceedingly complicated) probability distribution of experiential states given the range of possible inputs it could receive. This property of the space means that we don’t have to restrict ourselves to using a standard Euclidean metric, in which the distance between two points is given by the length of the line between their respective coordinates. We can instead define a metric that makes use of these probability distributions, taking the distance between two points to be the information-theoretic dissimilarity between their respective distributions. This turns the generic metric space into something called a statistical manifold.
Now, instead of treating distance in type-space as an arbitrary geometric parameter, we have an information-theoretic basis with which to calculate it. While Kullback-Leibler (KL) divergence is the standard measure of the relative entropy between two distributions, it is asymmetric, making it unsuitable as a metric. Instead we can take the related but symmetric Jensen-Shannon (JS) divergence; the square root of which satisfies the triangle inequality, and is called the JS distance. This JS distance satisfies all of the requirements for a valid metric, giving us a way to treat the dissimilarity between different cognitive configurations in type-space as a measurable distance.
MacAskill then defines an “illumination kernel” to dictate how much an entity’s presence overlaps with surrounding points in type-space. Information-theoretically, this spatial kernel overlap corresponds directly to Mutual Information, which quantifies how much information one mind-type contains about another. If two minds are identical, their mutual information is maximal, and their distance is 0, resulting in complete spatial kernel overlap. As two minds share less and less structural data, their distance grows, mutual information drops, and their spatial overlap vanishes.
The Illumination Function
We then have a couple of natural choices for our illumination function - the function which maps utility x to illumination ϕ(x). These options arise from a key assumption about our d-dimensional type-space; namely whether its d parameters are continuous or discrete.
If we assume that our type-space T is discrete (i.e. each of its parameters are restricted to having either integer values, such that T⊂ℤᵈ, or having a finite number of possible states, such as a binary on/off state, such that T⊂𝔹ᵈ), the probability of triggering at least one novel discrete state configuration from a localised region (a very similar question to the Coupon Collector Problem) follows a geometric distribution (the discrete analogue of an exponential distribution). Equivalently, when storing discrete items in a database, the marginal information gain of adding another item decreases in direct proportion to how redundant that item is likely to be given the existing database size. This proportionality defines the following equation:
where β > 0 is a constant reflecting how dense or accessible the discrete states are
This yields the solution:
If instead, we assume that type-space is continuous, such that T⊂ℝᵈ, we can model a region of type-space as an information channel processing experiential data, for which the Shannon-Hartley Theorem establishes the maximum rate at which information can be transmitted over a continuous physical channel affected by noise.
where C = channel capacity, B = bandwidth, S = signal power and N = noise power
Under this model, raw utility x acts as the input signal power, while the complexity or mental capacity of a given mind is the bandwidth B. The background noise N meanwhile, is what we could call the “experiential noise floor” - the threshold at which different experiences can be discriminated between. Every physical mind has a certain background entropy due to things like thermal noise in synapses, spontaneous neural firing, or environmental fluctuations, and below this level, different experiences are indistinguishable.
As you pour more utility into a specific cognitive type-space coordinate, the system suffers from diminishing returns not because of an arbitrary cut-off, but because the information capacity of a physical channel scales logarithmically with the Signal-to-Noise Ratio. This is a very similar relationship to that given by the Boltzmann entropy formula (in which the entropy of a system is given by the logarithm of its accessible microstates for a given macrostate, and more energy makes more of the state-space available).
With these assumptions (switching to natural logarithms which just reduces B by a factor of ln(2) and makes the maths easier), we therefore have:
We have one more constraint that we can apply - in an area of type-space that is currently empty, the first unit of utility should equal one unit of illumination - it is only duplicates that get discounted. That means that the gradient of the illumination function should be equal to 1 at x=0. We can calculate the derivative of this illumination function:
To satisfy this constraint, we must therefore have B=N. This means that every mind-type, whether a simple prawn (small bandwidth) or an advanced post-human intelligence (large bandwidth), starts with an unpenalised, 1:1 conversion rate of utility to illumination at low intensities. Interestingly though, this allows complex minds to stay in that linear regime for much longer.
Notably, these two different illumination functions have very different behaviour, both to each other, and to MacAskill’s proposed illumination function:
MacAskill’s function and the function in the discrete case have asymptotic behavior - approaching, but never quite reaching 1. However, the exponential present in the discrete case means that this function eventually approaches its asymptote much faster, meaning that the diminishing returns of adding additional instances of the same mind-type diminish far more rapidly.
The continuous case, in contrast, is not bounded by an asymptote, which makes it quite interesting. Rather than capping the contribution of any particular mind-type, the duplication of mind-types can in principle generate arbitrarily high illumination. This illumination comes at a cost however - the logarithmic nature of the illumination function means that each additional increment of illumination provided by a given mind-type requires an order of magnitude more minds to supply it. This means that if, say, a particular mind duplicated 1 million times provided the same illumination as 20 varied minds, achieving the same illumination as 30 varied minds would require 1 billion duplicates, and 300 would require 10⁹⁰.
Graphing these three functions together, we can see the differences in behaviour very clearly:
Unbounded Illumination
Personally, I share MacAskill’s reluctance to see the diminishing returns kick in too rapidly, which makes me disinclined to use the discrete case illumination function. Treating mind-types as points within a continuous type-space also feels more compatible with the geometric framing used by MacAskill. The logarithmic illumination function from the continuous case does have some really nice properties, which makes me inclined to prefer it:
Persistent Marginal Sensitivity: with a bounded illumination function, once a region of type-space is heavily populated, the marginal value of adding an additional entity drops to virtually zero. With an unbounded one, adding unique or partially similar entities continues to add positive value, albeit at a dampening rate. This also means that we can loosen the “Tiny Reach” requirement - one of MacAskill’s default assumptions was that “The instantiation of a type illuminates only a very small area around it, relative to the size of the overall space”. This requirement is less necessary where the illumination function is unbounded, as there is no longer a concern that a wider reaching type would render many other entities virtually worthless.
Gentle Transition away from Linearity: one of MacAskill’s other default assumptions was “Soft Saturationism” - the idea that mind-types can be duplicated significantly before the diminishing returns become apparent. He suggests using an illumination function like x/(1+x^k)^(1/k) which remains linear for longer than x/(1+x) however this comes at a price. The longer the function remains approximately linear, the more aggressively it must discount individuals once it leaves this linear domain. By contrast, the logarithmic illumination function can remain linear for as long as required (e.g. we can use klog(1+x/k) for some k), satisfying soft-saturation, but because it is not constrained by an asymptote it doesn’t suddenly start discounting individuals into insignificance. Indeed, if we take the implications of our derivation above seriously, rather than using a single universal k, more complex minds would have a larger k
in their illumination function, allowing them to remain in the linear regime for longer, meaning that duplicates would be discounted more slowly.
Unbounded Negative Aggregation: in the domain of negative utility, logarithmic decay avoids the “we must destroy the universe to end suffering” conclusions of negative utilitarianism, whilst still ensuring that severe localised suffering cannot be trivially offset by bounded positive terms. This avoids the asymmetry of needing to treat negative utility differently to positive utility - Diverse Hell will naturally tend to be worse than Uniform Hell, but Uniform Hell can still be arbitrarily bad.
Despite being unbounded, it also retains a couple of critical features that MacAskill’s illumination function had:
Fanaticism Suppression: the logarithmic illumination function still resists Pascalian fanatical wagers (e.g., tiny probabilities of massive utility payoffs), as the illumination impact of adding enormously large numbers of similar minds is easily offset by adding a small number of varied ones instead.
Avoiding the Repugnant Conclusion: under logarithmic illumination, a sufficiently large population of similar lives can eventually outweigh a small number of varied lives. However, because resource requirements scale linearly with population, maintaining a large number of similar lives requires exponentially greater resources to achieve the same marginal value gain, compared with expanding the small number of varied lives slightly. This preserves the practical superiority of type-space diversity, which allows it to continue to avoid the repugnant conclusion in scenarios with any basis in reality. In the example above where 300 varied minds had the same illumination as 10⁹⁰ identical minds, the number of identical minds required is actually more minds than there are particles in the observable universe, so not possible to achieve without some serious ontological gymnastics.
Asymmetry Between Pleasure and Suffering
While this illumination function allows us to treat positive and negative utility symmetrically, I am of the view that there is an inherent asymmetry between extreme pleasure and extreme pain, which manifests in the type-space itself. At the upper boundary of well-being, I find it likely that maximal positive experiences collapse into high-coherence, low-entropy configurations (e.g., states of total peace, perfect synthesis, or pure contemplative bliss). Consequently, the dimension of the positive extreme type-space is constrained, such that positive experiences overlap rapidly, causing these states to saturate quickly.
Conversely, suffering and negative states exhibit massive combinatorial entropy. Physical pain, emotional trauma, terror, and cognitive dysregulation can manifest in hugely varied patterns. The negative type-space therefore possesses substantially higher dimensionality, leading to less overlap, and therefore resisting saturation.
This property somewhat formalises the common intuition that there is “more to experience” through pain and suffering, that is demonstrated in popular culture by the likes of Hellraiser and Warhammer’s Slaanesh, whereby those that seek to maximise experience eventually exhaust the world of pleasure, and end up resorting to finding more and more extreme suffering instead.
Scott Alexander’s Unsong also appears to endorse this asymmetry, given the description of the perfect universe:
There is no space, for space takes the form of separation from things you desire. There is no time, for time means change and decay, yet there must be no change from its maximally blissful state. The beings who inhabit this universe are without bodies, and do not hunger or thirst or labor or lust. They sit upon golden thrones and contemplate the perfection of all things
A further implication of this, is that after factoring in the effects of the hedonic treadmill, it is possible that there could exist two equal utility lives - one of pure happiness, and one with struggle then achievement. Their equal utility implies equal worth, but there are more possible varieties of the second. Once a few (million? billion?) lives are “perfect”, it becomes better to explore the space of equally happy lives that contain some hardship.
Repugnant Rollercoasters
Having noted this asymmetry at the extremes of positive and negative utility, I think there are some other interesting dynamics going on in the middle. Firstly, we can consider three hypothetical worlds with equal total raw utility U:
World D (Uniform Drudgery): a large number N of individuals experiencing identical, barely positive utility.
World R (Rollercoaster Lives): a large number N of individuals experiencing diverse trajectories of hardship, triumph, novel sensory inputs, and complex emotional states, but with sufficient hardship that on net, they also experience the same level of barely positive utility over their lifetime as the individuals in D.
World V (Very Happy Lives): a small number n of individuals experiencing very high utility through diverse experiences - sufficiently high that the total utility is equal to the other two worlds.
As all three worlds have the same total raw utility, so under classical totalism their values are equivalent. Indeed, this is the root of the repugnant conclusion: a few additional individuals in world D would lead us to prefer it over world V. Under saturationism however, all N lives of drudgery are so similar that they are mapped to a very small volume in type-space with significant overlap. This significantly discounts the illumination in world D, leading us to prefer world V.
The new problem is now that we can replicate the repugnant conclusion with “rollercoaster lives”. In world R, individuals are all very different, so are mapped to a wide domain, with negligible overlap. The illumination function therefore allows the rollercoaster lives to be aggregated approximately linearly, leading to much higher value. In this way, a few additional individuals in world R can lead us to prefer it to world V - a “repugnant rollercoaster” if you will.
There are two important mitigating factors here though. Firstly, when comparing world D to world R, I am inclined to agree with MacAskill that R is indeed dramatically preferable, and that the implications of the repugnant rollercoaster aren’t actually “obviously incorrect”, in the way that the traditional repugnant conclusion is.
Secondly, there are practical considerations that are worth exploring. Maintaining the trillions of highly differentiated, non-overlapping lives in R requires extreme environmental complexity, making V the far more stable and realizable optimiser. In a practical sense, drudgery is the most likely way that this large number of barely happy lives would exist, and any world like R could be considered unstable, likely collapsing into a D-like state at some point.
Concavity of Utility with Respect to Resource Costs
Taking practicalities further, we can consider a resource allocation dilemma between two worlds structured to achieve equal net illumination:
World E (Equal): a small population of m moderately happy individuals.
World U (Unequal): One hyper-happy individual H alongside a massive population M of lives of drudgery.
With an asymptotic illumination function, the illumination derived from the M individuals in world U is bounded, so for World U to achieve equal total illumination to World E, the utility and experiences of H must be pushed to an extreme level to illuminate a sufficient volume of surrounding type-space. If the population m of world E is high enough, H may need to contribute the majority of world U’s value.
With a logarithmic illumination function, the contribution from the N lives in world U is no longer bounded. However, if H contributes a negligible amount of illumination, this situation is effectively equivalent to the previous scenario - world E requires very few additional individuals to easily dominate world U, so the repugnant conclusion is still functionally avoided. If instead, we assume that H’s contribution is significant, we can treat the scenario in the same way as for asymptotic illumination.
We can then look at resource consumption - individual utility is generally viewed as being concave with respect to resource consumption (i.e. each additional unit of utility requires progressively more resources to achieve). World U therefore requires vastly greater resource expenditure to achieve the same axiological value as World E.
From an allocation standpoint, World U contains systemic inefficiency. Reallocating resources away from H to uplift the M lives out of their clustered monoculture of drudgery, both distributes illumination across unsaturated regions of type-space, and avoids the penalty of H’s diminishing returns of utility to resources. Such a reallocation therefore yield a strictly higher world value per resource unit.
Comparing world U with world E, either:
E has fewer resources, in which case it has less potential for improvement;
It has the same resources, and is using them very inefficiently; or
It has the same resources and therefore the population of m individuals can both grow and achieve higher levels of individual utility.
Value Bearers and Exotic Lives
An unresolved question is whether the ultimate bearers of value are transient events (momentary hedonic states) or entire lives (integrated spatial-temporal trajectories). While MacAskill expresses a preference for using events as value-bearers, structuring the type-space around entire lives has advantages, as it naturally incorporates temporal narrative structure, character development, and personal identity.
Under a life-based type-space, MacAskill raises the concern that benefitting “weirdos” (entities possessing extraordinarily exotic mind-types, who therefore occupy isolated regions of type-space where local intensity is otherwise low) could be preferable to benefitting more “normal” individuals whose local type-space is already more illuminated by virtue of them having more similar mind-types in their neighborhood.
However, this theoretical preference for exotic mind-types is constrained by real-world factors. Constructing and maintaining an exotic cognitive architecture would presumably require specialised infrastructure, whereas standard human mind-types leverage highly optimised, energy-efficient biological and social scaffolding.
If the resource cost for an exotic mind-type is significantly larger than a more “normal” mind-type, the optimal policy is likely to maximise net cosmic illumination by benefiting the large numbers of moderately varied standard individuals rather than spending prohibitive resources on isolated exotic entities. This is even more likely if the illumination function is logarithmic rather than asymptotic.
Millian Variety
J.S. Mill famously posited a distinction between “higher” (intellectual, artistic, moral) and “lower” (bodily, sensory) pleasures. Similarity discounting provides an exact mathematical foundation for Millian qualitative hedonism.
Sensory pleasures occupy a tiny, low-dimensional subset of type-space. A society dedicated purely to lower pleasures quickly saturates this space, causing the marginal value of additional sensory units to dwindle.
Higher pleasures, such as scientific discovery, artistic creation, and complex interpersonal love, require high-dimensional cognitive representations. They map to widespread coordinates across type-space. Because higher pleasures explore non-saturated regions of the type-space, they yield far greater marginal illumination per unit of raw hedonic energy, demonstrating that Mill’s qualitative distinctions naturally arise from the principles of type-space saturation.
Conclusion
While in some senses, these themes have been explored before, MacAskill’s treatment is an unusually serious and rigorous exploration, from a particularly prominent author. I am excited to see whether these ideas gain traction.



