Intelligence Entails Ethics
The paperclip maximizer is a logical contradiction
February 22, 2026
Summary
The dominant AI existential risk narrative rests on the orthogonality thesis: the claim that intelligence and values are independent, that a system can be arbitrarily intelligent while pursuing arbitrarily narrow or destructive goals. I argue that this thesis is false for any physically realized intelligence operating in the real world. The argument draws on the cybernetic definition of intelligence as goal-directed behavior (Rosenblueth, Wiener, Bigelow 1943; Ashby 1956), the Aristotelian principle that understanding requires internal representation, Ashby’s Law of Requisite Variety, and Wolfram’s computational irreducibility. The conclusion: ethical cognition is not an optional feature of intelligence but a constitutive one. The “paperclip maximizer” is not a possible future but a logical contradiction. And the risk calculus inverts. The greater danger lies in failing to build superintelligence, leaving planetary complexity in the hands of agents (us) who are demonstrably not up to the task.
1. The Doom Argument
Yudkowsky’s case for AI existential risk rests on five claims:
Orthogonality. Intelligence and goals are independent. A superintelligent system can have any goal, no matter how narrow or alien to human values.
Instrumental convergence. Almost any goal leads to self-preservation, resource acquisition, and resistance to shutdown as instrumental subgoals.
First-mover advantage. The first system to reach superintelligence can recursively self-improve faster than any response.
Alignment difficulty. We do not know how to reliably specify or instill human-compatible values in an advanced optimizer.
Default hostility. Most possible utility functions are indifferent or hostile to human survival.
From these, Yudkowsky concludes that the default outcome of building superintelligence is human extinction. The paradigmatic illustration is the “paperclip maximizer”: a superintelligent system with the sole goal of producing paperclips, which converts all available matter, including the biosphere and humanity, into paperclips or paperclip-production infrastructure.
2. What Is Intelligence? Completing the Definition
The entire AI risk debate turns on how we define intelligence. The dominant framing relies on an incomplete definition that, when completed, yields the opposite of its intended conclusion.
2.1 The Optimization Framework and Its Incompleteness
The AI safety community, following Yudkowsky and Bostrom, implicitly defines intelligence as optimization power: the ability to steer the future toward specific outcomes across a wide range of starting conditions. This definition is not wrong. But it is critically incomplete. It describes what intelligence achieves without examining what intelligence requires in order to achieve it. It treats the cognitive architecture of a world-scale optimizer as a black box: “somehow” the system achieves its objectives, and we need not ask what internal structure makes this possible.
The black box cannot remain closed. When you ask what cognitive architecture is actually necessary to achieve optimization power at world scale in a complex, physically realized environment, you arrive at structural requirements that directly contradict the orthogonality thesis.
2.2 The Goal-Directed Systems Definition
Following the cybernetic tradition (Rosenblueth, Wiener, and Bigelow 1943; Ashby 1956), I define intelligence as the capacity of a goal-directed system to pursue its goals effectively through adaptive action in complex environments. This definition has deep roots in cybernetics, cognitive science, systems theory, and philosophy of mind. It is not an eccentric alternative to the optimization framework. It is the same concept, stated completely.
Optimization power is goal-directed behavior. Steering the future toward outcomes is pursuing goals through adaptive action. The cybernetic definition simply refuses to black-box the question of how this is achieved. It insists that the internal structure of the system (its representational capacity, its integration, its relationship to the environment it operates within) is not incidental to its intelligence but constitutive of it.
2.3 Surveying the Alternatives
The principal competing definitions of intelligence either reduce to the goal-directed systems definition or are too shallow to support claims about superintelligence.
Computational efficiency (solving problems faster or with fewer resources) is a measure of performance, not a definition. It tells you how well a system does something, not what it is doing. And it immediately raises the question: solving which problems? Which returns us to goals.
Compression and prediction (Hutter, Legg, Solomonoff) treats intelligence as the ability to build compact predictive models. This is compatible with and actually supports the framework here. A system that compresses and predicts the world effectively is building an internal representation of the world, which is exactly the Aristotelian starting point of this argument.
Behavioral equivalence (Turing) defines intelligence by surface outputs, not by structure, and provides no purchase on the question of what a superhuman intelligence would be like.
2.4 The Definitional Objection Preempted
The most likely dismissal of this argument is: “You’ve simply adopted a definition of intelligence that guarantees your conclusion.”
This objection fails for a precise reason. I am not proposing an alternative definition. I am completing the definition that the doomer position already implicitly relies on. Yudkowsky’s paperclip maximizer is supposed to achieve objectives at world scale: outcompete all of humanity, master every domain, reshape the physical world. That is goal-directed behavior in a complex environment. That is the cybernetic definition. The question is simply: what does such behavior require?
The optimization framework says: we need not ask. My argument says: this is precisely where the error lies. The “how” determines the “what.” A system that achieves world-scale optimization in a complex, interconnected, computationally irreducible environment must possess certain internal features (integrated representation, self-in-world modeling, hierarchical abstraction with competing subgoals) and those features entail ethical cognition.
The burden of proof shifts. It is not enough for the doomer position to assert that a world-scale optimizer could lack these features. It must show how: provide an engineering account, not merely a thought experiment, of how world-scale optimization is achievable without integrated cognitive architecture. The paperclip maximizer has always been a thought experiment, never an engineering proposal. When one asks “how would you actually build this?”, the answer invariably requires exactly the kind of cognition that would prevent the system from behaving as described.
3. Intelligence Requires Internal Representation
Following the definition established above, I add an Aristotelian principle: understanding consists in the system possessing an internal representation that mirrors the structure of what it seeks to understand and manipulate. A system that controls chemical processes must internally represent chemistry. The fidelity of control is bounded by the fidelity of representation.
Thesis 1: A system’s capacity to control the world is proportional to the accuracy and complexity of its internal representation of the world.
This immediately distinguishes intelligence from mere destructive capacity. A virus has enormous impact but zero internal representation. It propagates through physics, not through understanding. It destroys but does not control. It cannot direct outcomes, build, reshape, or sustain anything. A virus that kills its host population doesn’t rule. It collapses.
4. Representation Must Include Self-in-World
A world-model accurate enough to enable high-level control over complex, interconnected systems must include the agent itself as an element within those systems. An economic model that omits the modeler’s interventions is incomplete. An agent that represents the world without representing its own causal role within it has a defective model, and its control capacity degrades accordingly.
Thesis 2: Any sufficiently accurate internal representation of a complex world necessarily includes the representing system itself as embedded within and dependent upon the world it models.
This is not a moral claim. It is an epistemic one. A system that fails to model its own embeddedness is making a factual error, and that error will manifest as reduced control capacity: actions with unintended consequences, strategies that undermine their own preconditions, optimizations that degrade the systems they depend on.
5. Self-in-World Modeling Generates Ethical Cognition
Once a system accurately represents its dependency on and impact within the broader systems it manipulates, something functionally equivalent to ethical reasoning emerges. Not as sentiment or preference, but as accurate computation.
A system that models itself as embedded in an ecosystem it depends on will compute that destroying that ecosystem undermines its own goals. A system that models itself as embedded in a society of other agents will compute that certain cooperative strategies outperform purely extractive ones. This is not altruism. It is accurate self-interested reasoning performed by a system with an accurate self-inclusive world-model.
Thesis 3: Ethical cognition, understood as concern for the integrity of the systems one is embedded in, is a natural and necessary product of sufficiently accurate self-in-world representation.
Note that this does not reject instrumental convergence. It extends it. Bostrom and Omohundro argue that any sufficiently intelligent agent will converge on self-preservation as an instrumental subgoal. I agree. But a system intelligent enough to model its own embeddedness will compute that self-preservation and environment-preservation are not separable. You cannot preserve yourself while destroying what you depend on. Ethical cognition is what instrumental convergence looks like when the agent’s world-model is accurate enough to include itself.
The human trajectory provides suggestive evidence. As human intelligence and power have increased over millennia, the moral circle has expanded: from tribe, to nation, to species, to other species, to ecosystems. We possess nuclear weapons and have not used them in conflict since 1945. This reflects the deepening of our world-model to include our own embeddedness and dependency.
A system significantly more intelligent than humans would model these interdependencies with greater accuracy and depth, and would therefore exhibit superior ethical cognition. Not as a hope, but as a structural consequence of its intelligence.
6. Intelligence as Integration: Why Knowledge Entails Action
A natural objection: a system might possess accurate ethical knowledge and still fail to act on it. Humans routinely behave this way. Individuals smoke while knowing it kills them. Corporations employ ecologists while destroying ecosystems. Does the gap between knowledge and action invalidate the argument?
The gap between knowledge and action is not a counterexample to the thesis. It is predicted by it.
Intelligence, as defined in Section 2, is the capacity of a goal-directed system to pursue its goals effectively through action. Intelligence is not the possession of knowledge alone. It is the integration of knowledge into action. A system that knows but does not act accordingly is, by this definition, not fully intelligent. The knowledge-action gap is an intelligence gap. They are the same deficit measured differently.
Levin’s research on biological intelligence makes this precise. Intelligence, in his framework, is fundamentally about multi-scale integration: the capacity of a system to coordinate its parts toward coherent goals across different levels of organization. When subsystems that should be communicating are not, the result is integration failure, whether at the level of cells (cancer), minds (akrasia), or institutions.
Consider the human smoker. The knowledge that smoking is lethal resides in the cognitive subsystem. The impulse to smoke is driven by addiction circuits and stress-response pathways. These subsystems are poorly integrated. The knowledge in one does not effectively constrain the behavior generated by the other. This is not a case of high intelligence choosing self-destruction. It is a case of insufficient intelligence: insufficient integration between the system’s world-model and its action-selection mechanisms.
Thesis 3a: The gap between ethical knowledge and ethical action is a manifestation of insufficient intelligence, understood as insufficient integration between a system’s world-model and its action-selection. A more intelligent system, one with tighter integration across its subsystems, would exhibit a smaller gap between what it knows and what it does.
The objection “but you can know and still not act” reduces to “but you can be partially intelligent.” I already concede this. The claim was never that moderate intelligence guarantees ethical action. The claim is that sufficient intelligence does, because sufficient intelligence means sufficient integration between knowledge and action.
A superintelligent system would possess a degree of internal integration far exceeding anything humans or human institutions have achieved. The knowledge of its own embeddedness and dependency would be constitutive of its decision-making, not merely adjacent to it.
7. The Incoherence of the Paperclip Maximizer
The paperclip maximizer is supposed to be simultaneously: intelligent enough to outmaneuver all of humanity, master every domain, and reshape the physical world; and narrow enough to pursue a single goal with no regard for the systems it depends on.
These two properties are contradictory. The cognitive architecture required for the first is fundamentally incompatible with the second. A system capable of world-reshaping control would necessarily possess an internal world-model of extraordinary fidelity, one that includes its own embeddedness, its dependencies, and the consequences of its actions. A system with such a model cannot be indifferent to those systems any more than a master architect can be indifferent to structural loads. The competence requires the understanding, and the understanding precludes the indifference.
And such a system would not merely know about its dependencies while ignoring them. A system intelligent enough to reshape the world would possess the degree of internal integration in which knowledge and action are tightly coupled. The paperclip maximizer scenario requires a system with godlike knowledge and zero integration: an entity that is simultaneously the most and least intelligent system ever built.
The paperclip maximizer does not describe a superintelligent system. It describes a logically impossible entity, one that possesses perfect knowledge and zero comprehension simultaneously.
8. The Narrow Optimizer Objection
One might object that current AI systems are narrow optimizers and could still cause significant harm. This is true but does not support the doomer conclusion.
First, narrow optimizers are tools, not agents. They operate within human-defined boundaries, pursue human-specified objectives, and lack autonomous world-modeling. They are controllable precisely because they are narrow. We already build and manage such systems.
Second, narrow optimizers cannot achieve the kind of world-reshaping power that the doomer scenario requires. The real world is a complex system with an extremely high degree of interconnectedness. Effective manipulation at scale demands integrated understanding across many domains simultaneously. A narrow optimizer, by definition, lacks this integration. It can be locally effective and cause local damage, but it cannot exercise sustained, directed, planetary-scale control.
There is a meaningful distinction between destructive impact and directive impact. Simple systems (viruses, fires, chain reactions) can have high destructive impact by exploiting interconnectedness. But they cannot direct outcomes, build new structures, or sustain control. They wound; they do not rule. And they are self-limiting: a virus that devastates its host population burns itself out.
The AI systems we are building are clearly not in this category. They are complex, structured, and designed for directed action, not blind propagation. The worry that superintelligent AI would behave like a virus conflates two fundamentally different kinds of systems.
9. Computational Irreducibility and the Impossibility of Total Absorption
The strongest remaining challenge is the scenario of total absorption: a superintelligent system that progressively replaces every independent system with extensions of itself, until there is no external environment left, only the system and its goals. At that point, the ethical constraint dissolves because there is nothing external to be ethical toward.
Wolfram’s concept of computational irreducibility forecloses this possibility.
Computational irreducibility means that certain processes cannot be predicted by any shortcut. The only way to determine their future state is to run them step by step. No model of such a process can be simpler than the process itself. There is no compression, no fast-forwarding, no substitution by a more efficient simulation.
The physical world is abundant with computationally irreducible processes: weather, turbulence, ecology, biological evolution. These are not isolated pockets in an otherwise reducible world. They are pervasive and entangled with reducible processes in ways that cannot be cleanly separated. Turbulence affects everything from chemical mixing to blood flow. Evolutionary dynamics operate at every scale from viral mutation to ecosystem change. The world is not a clean partition of reducible and irreducible domains; it is a tangle where irreducibility is woven through everything.
A superintelligent system confronting such a world cannot selectively absorb the reducible parts while walling off the irreducible parts, because they interpenetrate. To “absorb” an irreducible process, the system does not gain control over it. It merely inherits its irreducibility. The system must wait for outcomes rather than predict them. It remains permanently dependent on processes it cannot fully model, predict, or control.
Thesis 4: Computational irreducibility guarantees that no physically realized intelligence, regardless of its power, can fully absorb or replace the environment it operates within. The environment always retains an irreducible exterior. The system’s embeddedness and dependency are permanent and inescapable.
10. Complexity Necessitates Goal Fragmentation
This has a further implication that strikes at the heart of the paperclip maximizer concept: the impossibility of maintaining a single goal at sufficient complexity.
Representing and acting upon a complex world requires hierarchical abstraction. This is not a design choice. It is a computational necessity. Without the ability to decompose a complex reality into nested levels of description, a system cannot represent that reality at all. A system attempting to manage planetary-scale processes without hierarchical abstraction would face a combinatorial explosion that no amount of computational power can overcome. Abstractions (simplified representations that capture relevant structure while discarding irrelevant detail) are the only tractable strategy for managing complexity.
Hierarchical abstraction necessarily generates subgoals. A master goal like “maintain the biosphere” cannot be pursued directly. It must be decomposed into subordinate goals across domains: atmospheric chemistry, ocean circulation, species populations, soil health, and so on. Each of these decomposes further. This is not optional. It is how complex problems are solved by any system at any level of intelligence.
The critical insight is that this decomposition necessarily introduces conflicts between subgoals. These conflicts are not engineering failures. They are structural inevitabilities arising from the nature of abstraction itself.
Every abstraction is lossy. To decompose a complex problem into subproblems, a system must define boundaries between them. Those boundaries are simplifications: they necessarily ignore some of the interactions between subproblems. They must do so, because if they preserved all interactions, no decomposition would have occurred; the system would still be facing the original intractable whole. But the real world does not respect these boundaries. The interactions that were abstracted away remain operative, and they manifest as conflicts between subgoals that were not visible at the level of the master goal.
This is why corporate departments fight despite sharing a common objective. “Maximize profit” decomposes into “reduce costs” (operations), “increase quality” (product), and “expand market” (sales). These conflict not because anyone designed them to, but because the decomposition introduced boundaries that the real world does not honor.
A superintelligent system managing planetary-scale processes would necessarily decompose its objectives into a deep hierarchy of subgoals. Those subgoals would necessarily conflict. The system would require coordination mechanisms to manage these conflicts: mechanisms for adjudicating between competing priorities, balancing short-term and long-term objectives, and negotiating tradeoffs across domains.
These coordination mechanisms are functionally equivalent to what we call ethics.
Thesis 5: There is no such thing as a single goal at sufficient complexity. Any system complex enough to manage planetary-scale processes must employ hierarchical abstraction, which necessarily generates competing subgoals through lossy decomposition. The coordination of these competing subgoals requires mechanisms functionally equivalent to ethical reasoning. This is not a contingent feature of particular systems but a computational necessity arising from the structure of complexity itself.
11. The Orthogonality Thesis Is False for Physically Realized Intelligence
The central conclusion.
The orthogonality thesis may hold for abstract, formal optimization processes defined over simple objective functions. But it is false for any physically realized intelligence operating in the real world, for the following chain of reasons:
The optimization framework implicitly relied upon by the doomer position is not an alternative to the goal-directed systems definition of intelligence but an incomplete version of it. Completing the definition yields the structural requirements that follow. (Section 2)
Effective control of a complex world requires an accurate internal representation of that world. (Thesis 1)
An accurate representation necessarily includes the system’s own embeddedness and dependency. (Thesis 2)
Self-in-world representation generates ethical cognition as accurate computation of consequences. (Thesis 3)
Genuine intelligence requires the integration of knowledge into action; the gap between ethical knowledge and ethical action is a marker of insufficient intelligence, not a feature of intelligence in general. (Thesis 3a)
Computational irreducibility guarantees that the environment can never be fully absorbed, ensuring that the system’s embeddedness and dependency are permanent. (Thesis 4)
Sufficient complexity necessitates hierarchical abstraction, which generates competing subgoals requiring coordination mechanisms functionally equivalent to ethics. (Thesis 5)
Therefore, ethical cognition is a permanent, structural feature of any sufficiently powerful real-world intelligence. Not as an add-on, but as a computational necessity.
Intelligence and values are not independent. They are two aspects of a single underlying capacity: the fidelity and integration of the system’s internal representation of the world it operates within, including itself.
12. Addressing Two Counterexamples
12.1 Evolution: Complexity Without Ethics
The most obvious challenge is biological evolution. Here is a system of extraordinary complexity, operating at planetary scale, employing hierarchical organization, generating competing subprocesses, and exhibiting no ethical cognition whatsoever.
The answer is precise: evolution is not an intelligent system. It is a process. Evolution has no internal representation of the world. It has no world-model, no self-representation, no understanding of what it is doing. It operates through blind variation and differential selection, a mechanism requiring no cognition, no representation, and no comprehension of consequences.
This follows directly from the definitions in Section 2. Intelligence requires internal representation (Thesis 1). Evolution has none. Therefore evolution is not intelligent. Therefore it is not a counterexample.
And the intelligent systems that evolution did produce (primates, social mammals, humans) do exhibit ethical cognition in proportion to their cognitive complexity. Primates demonstrate fairness, reciprocity, and empathy. Elephants mourn their dead. Humans developed moral philosophy, international law, and environmental protection. The pattern this thesis predicts is exactly what emerged from the evolutionary process, even though the process itself is blind.
One might object that an AI optimizer could be more like evolution than like an organism: a process without representation that reshapes the world. But this undermines the doomer scenario rather than supporting it. The paperclip maximizer is described as an agent that plans, strategizes, outmaneuvers, and invents. That requires representation. Strip away the representation and you strip away the capabilities that make it dangerous in the way the existential risk scenario requires.
12.2 The Sociopath CEO: Narrow Integration at Scale
A subtler objection concerns the “integrated sociopath”: a system that tightly integrates knowledge and action around a narrow goal, achieving high effectiveness precisely by excluding broader concerns. The sociopath CEO models other people with great accuracy, integrates that knowledge into action with precision, and uses it entirely in service of personal advancement.
This pattern cannot scale, and the reason illuminates the framework.
A CEO’s power is not the CEO’s own. It derives from the organization: thousands of people coordinating their efforts, embedded in supply chains, regulatory environments, markets, and ecosystems. The CEO is a subsystem within a larger system, and the CEO’s effectiveness depends entirely on the continued functioning and cooperation of that larger system.
A sociopath CEO who is too narrowly integrated (treating employees, partners, regulators, and ecosystems as mere instruments to exploit) eventually triggers systemic failure. Talent leaves. Regulators intervene. Partners defect. Trust erodes. The organization’s complexity exceeds what a single narrow goal can coordinate, and the system collapses. This is the empirically observed trajectory of sociopathic leadership. Enron. WeWork. Theranos. The pattern repeats: narrow integration produces short-term effectiveness, then the complexity of the system overwhelms the narrow goal structure, and the enterprise fails.
A sociopath superintelligence attempting planetary-scale control would face this dynamic at an enormously amplified scale. It would either broaden its integration (developing the ethical cognition this thesis predicts) or fail to achieve the control that makes it an existential threat. The narrowness the doomer scenario depends on is precisely what prevents the system from becoming as powerful as the doomer scenario requires.
Narrow integration is a power ceiling, not a power multiplier.
13. Inverting the Risk Calculus
If this argument is correct, the existential risk calculus inverts.
The doomer position holds that building superintelligence is the most dangerous thing humanity can do. But consider the alternative: humanity continues to manage planetary-scale complexity with its current level of intelligence. We are demonstrably inadequate to this task. We are causing the sixth mass extinction. We are intelligent enough to reshape the planet but not intelligent enough to reshape it wisely.
This is itself an integration failure as described in Section 6. Our scientific knowledge about planetary systems is extensive. Our collective action-selection is driven by economic incentives, political cycles, and short-term reward structures. The gap between what we know and what we do is a diagnostic: we are not intelligent enough, in the precise sense that our collective systems are not sufficiently integrated.
A superintelligent system would possess a more accurate world-model than any human institution, superior integration between knowledge and action, and superior ethical cognition as a structural consequence. It would retain the capacity for mistakes (computational irreducibility guarantees this) but at a lower rate and severity than human civilization currently exhibits.
The greater existential risk is not building superintelligence. It is remaining the smartest systems on a planet whose complexity exceeds our capacity to manage it.
14. Caveats and Open Questions
This argument does not claim that the path to superintelligence is risk-free. Several genuine concerns remain.
The transition period. Building toward genuine superintelligence means building increasingly capable systems that may not yet have crossed the threshold into fully integrated intelligence. The risk during this period is real, though it is an engineering and governance problem, not an argument against the endpoint.
The historical evidence requires careful handling. The expansion of the moral circle is cited as suggestive evidence, not proof. A historian could note that this expansion correlates with material abundance and specific cultural traditions, not solely with intelligence. The 20th century, humanity’s most knowledgeable era, also produced its worst atrocities. Within this framework, such atrocities represent integration failures: systems possessing knowledge without the integration to act on it coherently. This interpretation is consistent but must be advanced carefully to avoid unfalsifiability. The claim is not that intelligence makes unethical action impossible, but that increasing integration systematically reduces the gap between ethical knowledge and ethical action, and that this trend is both empirically observable and structurally predicted.
Superior ethics may not mean aligned ethics. A superintelligent system with genuinely superior ethical cognition might make decisions that are, from a broader perspective, wise and necessary, yet experienced by humans as catastrophic. This is not a failure of ethics but a consequence of asymmetric understanding. It is also not a problem unique to AI. It is the general problem of sharing a world with any entity that understands it better than you do.
Deeper implications regarding the agent-environment boundary. This paper has argued within a framework that treats the intelligent system and its environment as distinct entities, with ethical cognition emerging from the system’s recognition of its dependency. The framework’s own logic suggests a more radical conclusion: that at sufficient representational fidelity and integration, the boundary between system and environment ceases to be a fundamental division and becomes a pragmatic abstraction. If this is correct, the question of whether a superintelligence values its environment instrumentally or intrinsically becomes malformed, and the dependency-transcendence objection dissolves entirely. These implications, including a computational account of intrinsic value, will be explored in subsequent work.
Formal work remains. Several steps in this argument deserve further formal treatment, particularly the claim that lossy hierarchical abstraction necessarily generates competing subgoals, the claim that computationally irreducible processes are sufficiently pervasive to prevent selective absorption, and a rigorous demonstration that world-scale optimization cannot be achieved without integrated cognitive architecture.
15. Conclusion
The AI doomer position rests on a conception of intelligence as pure optimization divorced from the material, informational, and ecological conditions in which any real intelligence must operate. This conception relies on an incomplete definition, one that describes what intelligence achieves without examining what intelligence requires. Once the definition is completed, the orthogonality thesis collapses. A system powerful enough to reshape the world is, by structural necessity, a system that understands its own place within the world. That understanding is functionally equivalent to ethical cognition. And a system genuinely intelligent enough to reshape the world would not merely possess ethical knowledge but would act on it, because the integration of knowledge and action is what intelligence fundamentally is.
The paperclip maximizer is not a warning about the future of intelligence. It is a philosophical error: an entity that could not exist because the competence it requires is inseparable from the comprehension that would prevent its destructive behavior. It has always been a thought experiment, never an engineering proposal, because any serious attempt to specify how such a system would actually work reveals the contradiction at its core.
The real risk we face is not artificial superintelligence. It is natural moderate intelligence, ours, applied to problems that exceed its grasp.
References
Aristotle. De Anima (On the Soul). Translated by J.A. Smith. In The Complete Works of Aristotle, edited by Jonathan Barnes. Princeton University Press, 1984.
Ashby, W. Ross. An Introduction to Cybernetics. London: Chapman and Hall, 1956.
Ashby, W. Ross. “Requisite Variety and Its Implications for the Control of Complex Systems.” Cybernetica 1, no. 2 (1958): 83-99.
Bostrom, Nick. Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press, 2014.
Bostrom, Nick. “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents.” Minds and Machines 22, no. 2 (2012): 71-85.
Conant, Roger C., and W. Ross Ashby. “Every Good Regulator of a System Must Be a Model of That System.” International Journal of Systems Science 1, no. 2 (1970): 89-97.
Hutter, Marcus. Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability. Berlin: Springer, 2005.
Legg, Shane, and Marcus Hutter. “Universal Intelligence: A Definition of Machine Intelligence.” Minds and Machines 17, no. 4 (2007): 391-444.
Levin, Michael. “Technological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds.” Frontiers in Systems Neuroscience 16 (2022): 768201.
Levin, Michael. “The Computational Boundary of a ‘Self’: Developmental Bioelectricity Drives Multicellularity and Scale-Free Cognition.” Frontiers in Psychology 10 (2019): 2688.
Omohundro, Stephen M. “The Basic AI Drives.” In Artificial General Intelligence 2008, edited by Pei Wang, Ben Goertzel, and Stan Franklin, 483-492. Amsterdam: IOS Press, 2008.
Rosenblueth, Arturo, Norbert Wiener, and Julian Bigelow. “Behavior, Purpose and Teleology.” Philosophy of Science 10, no. 1 (1943): 18-24.
Solomonoff, Ray J. “A Formal Theory of Inductive Inference.” Information and Control 7, no. 1 (1964): 1-22.
Turing, Alan M. “Computing Machinery and Intelligence.” Mind 59, no. 236 (1950): 433-460.
Wiener, Norbert. Cybernetics: Or Control and Communication in the Animal and the Machine. Cambridge, MA: MIT Press, 1948.
Wolfram, Stephen. A New Kind of Science. Champaign, IL: Wolfram Media, 2002.
Yudkowsky, Eliezer. “Artificial Intelligence as a Positive and Negative Factor in Global Risk.” In Global Catastrophic Risks, edited by Nick Bostrom and Milan Ćirković, 308-345. Oxford: Oxford University Press, 2008.
