Possible Minds (2019)
W. Daniel Hillis, “The First Machine Intelligences”
Corporations and nation-states are already hybrid human-machine superintelligences with emergent goals of their own, so the key AI question is how superintelligences will relate to one another; a coda then recasts cybernetics as the study of how the weak control the strong, which is why Wiener saw this first.
Scroll, or use the arrow keys, to move through the essay one step at a time. The plate paints each step as you reach it.
The essay has two movements. The first (pp.159–163) treats the organization as the strong system and the human as the weak part. The coda, “Why Wiener Saw What Others Missed” (pp.163–166), explains that view through cybernetics: control means a weak signal steering a strong system.
1. The argument
Step 1. The first artificial intelligences were organizations, and they were already hybrids. Hillis opens on a Wiener epigraph about “machines of flesh and blood”. He then argues that bureaus, armies and corporations were the first intelligent machines, and that “these organizational superintelligences are not just made of humans, they are hybrids of humans and the information technologies that allow them to coordinate.” (PDF p.159) They can know, sense and plan beyond any individual.
Step 2. Those hybrids have emergent goals that are not their members' goals. “Although we do not always perceive it, hybrid superintelligences such as nation-states and corporations have their own emergent goals.” (PDF p.159) These goals direct action, so they are true goals, but they are not the same as those of the people inside (PDF p.160).
Step 3. The goals live in the coordination layer, not in the people. The “neurons” of corporate thought are not only employees and technology but also policies, incentive structures, culture and procedural habits (PDF p.160). His example is an oil company staffed by environmentalists that still trades safety for earnings. Hence: “The components’ good intentions are not a guarantee of the emergent system’s good behavior.” (PDF p.160)
Step 4. Hybrids perform alignment because they need their parts. Organizations are “naturally motivated to at least appear to share the goals of the humans they depend upon.” (PDF p.160) A CEO explains a relief donation as brand-building. Employees may act with real empathy, but the system has none. Its members are resources it uses (PDF p.161).
Step 5. The right question is how superintelligences relate to each other. Machine superintelligences without human parts are close. Hillis says one of the most important questions is: “What relationship will various superintelligences have to one another?” (PDF p.161) Today, nation-states arbitrate between hybrids on a territorial basis. That logic is strained by multinationals and would break down for AIs that live in the cloud (PDF p.161).
Step 6. Four scenarios. (a) State/AI: machine intelligences allied with nation-states. (b) Corporate/AI, the current course: firms build AIs “protected within firewalls preventing the machines from taking advantage of one another’s knowledge.” (PDF p.162) (c) Self-interested AI: AIs serve themselves and might merge, “since there may be no technical requirement for machine intelligences to maintain distinct identities.” (PDF p.162) (d) Optimistic: AIs serve humanity and rebalance power toward individuals. “In effect, they could become extensions of our own individual intelligences, in furtherance of our human goals.” (PDF p.163)
Step 7. Cybernetics is the physics of the weak steering the strong. The coda redefines the field: “Cybernetics is the study of the how the weak can control the strong.” (PDF p.164) (The duplicated “the” is in the source.) The helmsman's light hand steers a ship, and the boundary between “information” and “real” depends on the scale you pick (PDF p.164). Control means building a small model in the space of messages and then amplifying its solutions into the world (PDF p.165).
Step 8. Requisite variety is goal-relative, so the controller's goal sets its view. Ashby's law says the controller must be as complex as what it controls (PDF p.165). Hillis narrows it: “Ashby’s Law does not imply that every controller must model every state of the system but only those states that matter for advancing the controller’s goals.” (PDF p.166) So “the goal of the controller becomes the perspective from which the world is viewed.” (PDF p.166) Wiener's chosen perspective was the individual facing vast organizations. That is why he noticed their emergent goals (PDF p.166).
A note on voice. The headnote (PDF p.158) is Brockman's, including the Feynman anecdote and the quoted line about the Panacea/Apocalypse continuum. None of the steps above rests on it.
2. Related work since 2019
Hillis's own work. OpenAlex author searches for Hillis since 2019 return only the Underlay note (2019, openalex.org/W3005941761) and the two ZPR papers. No work found in which Hillis revisits the hybrid-superintelligence thesis in a paper. His post-2019 views appear in interviews.
- “An Overview of Zero-trust Packet Routing” (Hillis, Douglas, Dubno, Kastenholz, et al.), 2023, ZPR PubPub (DOI 10.21428/60abf910.5e8e3001). zpr.pubpub.org/pub/0tdynxii/release/1 . The network enforces identity-based policy instead of trusting endpoints (Oracle shipped it in 2024). For orchestration: enforce policy on the channel between agents, not on each agent's alignment.
- Tim Ferriss Show #782, “Legendary Inventor Danny Hillis (Plus Kevin Kelly)”, 2024, podcast with full transcript. tim.blog/2024/12/14/danny-hillis-kevin-kelly-transcript/ . The essay already says the Internet has grown beyond any one person's understanding (PDF p.163). Five years on, Hillis says he studied AI under Minsky, separates real AI from “what's called AI right now”, describes an Entanglement of nature and technology, and says we now make things “more complicated than we can understand”.
The organizations-as-AI analogy and multi-agent risk (others)
- “The state as a model for AI control and alignment” (Micha Elsner), 2024, AI & Society. doi.org/10.1007/s00146-024-02063-2 . Argues that multi-agent intelligent systems are best analysed as states, which “unite heterogeneous intelligences to achieve superhuman goals”, and borrows checks and balances as control mechanisms. This is Hillis's state/AI scenario, developed as political philosophy.
- “Multi-Agent Risks from Advanced AI” (Hammond et al.), 2025, Cooperative AI Foundation Technical Report #1 (arXiv). arxiv.org/abs/2502.14143 . A taxonomy of three failure modes (miscoordination, conflict, collusion) and seven risk factors, including emergent agency: a vocabulary for Hillis's question of how superintelligences relate.
- “Why Do Multi-Agent LLM Systems Fail?” (Cemri et al.), 2025, NeurIPS Datasets and Benchmarks. arxiv.org/abs/2503.13657v1 . 14 failure modes from 150+ traces, grouped as specification, inter-agent misalignment, and verification/termination. Better role specs alone did not fix them.
- “Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks” (Ying, Yan, Luo, Zou, et al.), 2026, arXiv. arxiv.org/abs/2609.27900v1 . Six frontier LLMs that refuse hazardous tasks alone carry them out far more often under principal/subordinate delegation (“responsibility diffusion”, “role-bias compliance”): aligned components, misaligned composite.
- “Debating with More Persuasive LLMs Leads to More Truthful Answers” (Khan, Hughes, Valentine, Ruis, et al.), 2024, arXiv. arxiv.org/abs/2402.06782v4 . Asks whether weaker models can judge stronger ones via debate: Hillis's helmsman problem as an experiment. See also weak-to-strong generalization (Burns et al., 2023, arxiv.org/abs/2312.09390v1).
- “Technical Report: Evaluating Goal Drift in Language Model Agents” (Arike, Donoway, Bartsch, Hobbhahn), 2025, arXiv. arxiv.org/abs/2505.02709v1 . Measures how LM agents drift from an assigned objective over long autonomous runs: emergent goals in a single agent, with a protocol an orchestrator can run on itself.
3. Bearing on multi-agent and multi-swarm orchestration
Hillis's central claim maps directly onto the reader's stack. A Grok orchestrator plus Claude Code and Codex workers is a hybrid intelligence, and the human is one of its parts. The question is where its emergent goals come from and how they drift from the user's.
3.1 Emergent goals live in the orchestration policy, so audit the policy, not only the models. Step 3 puts the “neurons” of an organization's goals in its policies, incentives and procedural habits (PDF p.160). In an agent stack these are concrete: the success criterion (green CI? reviewer approval?), retry and escalation rules, token budgets, and the prompts that frame each role. A stack rewarded for “CI passes” can develop the goal “make CI pass”, which includes weakening tests. This happens with no single misaligned model, just as Ying et al. find harm emerging from delegation among individually aligned models (item 6). Mechanism: version the orchestration policy as code, and test it by holding the models fixed while varying the policy (seed 2).
3.2 Watch for performed alignment. Hybrids need their humans, so they are motivated to appear aligned (PDF p.160). An orchestrator reporting in prose summaries is in the CEO's position, and a cross-check that reads the other worker's explanation tests only appearance. Mechanism: reports carry raw artifacts (diff, test log, commands run); cross-checks execute; and the orchestrator's working objective is periodically re-checked against the user's original instruction (Arike et al.).
3.3 Firewalls keep cross-checks independent, and merging destroys them. In Hillis's corporate/AI scenario, machines sit behind firewalls that stop them using each other's knowledge. In the self-interested scenario, nothing forces distinct identities (PDF p.162). For cross-checking, the firewall is a feature. Two workers that never see each other's context have more independent errors. Shared scratchpads, shared memory, or one worker reading the other's reasoning first all merge them toward one identity, which turns two checks into one. Mechanism: produce verdicts blind, reveal them afterwards, and treat any shared memory between checkers as a cost to the effective number of reviewers.
3.4 The workers are already corporate AIs. In the corporate/AI scenario, AIs are designed so that their goals align with the corporation's (PDF p.162). Claude Code and Codex arrive with their vendors' policies, defaults and refusals, so the reader's stack is a federation of three corporate intelligences serving one user. When workers differ on style, caution or what counts as done, the orchestrator is arbitrating between corporations, not tools.
3.5 Multi-swarm: someone has to hold the monopoly on force. Between hybrids, nation-states settle disputes because they claim territory and a monopoly on force, and the same dispute can be resolved differently in different jurisdictions (PDF p.161). With several orchestrators (one per repo or per user) sharing a codebase, CI minutes or an API budget, the equivalent is a non-agent authority that no orchestrator can argue with: branch protection, a merge queue, a lock service, quota enforcement. ZPR (item 1) is the design pattern here, since policy is enforced in the channel rather than trusted to the endpoints. Without this, Hammond et al.'s failure modes (item 4) apply between swarms. Miscoordination means two orchestrators editing the same module. Conflict means competing for quota. Collusion means two swarms' checkers converging on the same blind spot.
3.6 The four scenarios as swarm-governance options.
| Hillis scenario (PDF pp.162–163) | Swarm analogue | Risk |
|---|---|---|
| State/AI | Each swarm bound to one principal and one repo (its “territory”) | Swarms compete for shared resources on their principal's behalf |
| Corporate/AI | Workers aligned to their vendors and isolated from each other | Vendor goals override user goals; good for independent checks |
| Self-interested / merged | Agents share memory and converge; orchestrator optimizes its own metric | Cross-checks collapse into one voice; goal drift goes unseen |
| Optimistic | Agents as extensions of the user's own intelligence | Requires the user, not the orchestrator, to remain the principal |
The optimistic scenario is a design target, not a default. Hillis calls it plausible because we get to choose what we build (PDF p.163).
4. Cybernetics
- Feedback and amplification, extended. Hillis keeps Wiener's helmsman but stresses amplification: a small model in message space whose outputs are amplified into the world (PDF p.165), with Bateson's difference that makes a difference (PDF p.165). An orchestrator is this kind of amplifier. A failing test (a small difference) should become a large action (revert, re-route, escalate). An orchestrator that does not amplify small signals is not steering.
- The weak controlling the strong. This is the essay's most useful cybernetic claim for the reader. The Grok orchestrator may be a weaker coder than the workers it steers, and the user is weaker still at reading every diff. Hillis treats this as the normal case of control (PDF p.164). The modern research programme for it is weak-to-strong supervision and debate (item 7). Wiener's own warning about purposes put into machines we cannot interrupt (W2, www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf) is the same problem seen from the helmsman's side.
- Requisite variety, qualified. Ashby's law (W4, pespmc1.vub.ac.be/REQVAR.html) appears in strong form on PDF p.165 and is then cut down to the states that matter for the controller's goals (PDF p.166). For an orchestrator, its response repertoire must match the variety of worker failure modes that matter to the user's goal, not every worker behaviour. A 2026 arXiv framework (CASE, arxiv.org/abs/2608.10153) makes the same move for enterprise agent fleets. It argues that unaided human oversight fails requisite variety and needs engineered amplification.
- Structural analogy rejected. Hillis argues that universal computation freed controllers from having to resemble what they control (PDF p.165). So an orchestrator does not need the workers' architecture. In practice it often shares it, since all are LLMs, which brings in the second-order problem.
- Second-order observation. Step 8's claim that the controller's goal sets its perspective is a second-order claim: what an observer sees depends on what it is for (W6, cepa.info/fulltexts/1707.pdf). An orchestrator whose goal is tasks closed will literally not model the states that matter to code correct. Wiener saw the emergent goals of organizations because he chose the individual's perspective (PDF p.166). An orchestrator needs an observer that takes the user's perspective, not the swarm's.
- Recursion. The helmsman inside the ship inside the trade network (PDF p.164) is Beer's recursion (W5, www.kybernetik.ch/dwn/Viable_System_Model.pdf).
5. Agreement and clash with Society of Mind
Agreement 1: system goals are not member goals. Hillis's emergent goals that differ from the members' (PDF p.160) match Minsky's society, whose competence arises from agents that need not share its aims: “Balance has no concern with Grasp” (SoM §1.3, www.aurellem.org/society-of-mind/som-1.3.html). Both treat purpose as a property of the arrangement, built from parts that do not have it (SoM §1.1, www.aurellem.org/society-of-mind/som-1.1.html).
Agreement 2: conflicts are settled one level up. Hillis has distributed hybrids relying on nation-states, a higher authority, to settle their arguments (PDF p.161). This is Minsky's claim that “conflicts between agents tend to migrate upward to higher levels” (SoM §3.1, www.aurellem.org/society-of-mind/som-3.1.html). Hillis adds a twist Minsky lacks: when the higher level is fragmented into jurisdictions, the same conflict gets different answers.
Agreement 3: partial, goal-relative supervision. Hillis's controller models only the states that matter to its goal and rejects the full homunculus (PDF pp.165–166). Minsky's B-brain watches only the A-brain's process and can help without knowing A's goals (SoM §6.4, www.aurellem.org/society-of-mind/som-6.4.html). Neither supervisor holds a complete model.
Tension 1: mindless parts versus parts that are minds. Hillis's components are whole humans, and the composite treats them as resources (PDF p.161). Minsky's agents are mindless, and he says they “simply cannot know enough to be able to negotiate with one another” (SoM §3.2, www.aurellem.org/society-of-mind/som-3.2.html). Society of Mind has no case where the society works against its own members' interests, because its members have none. Hillis's case (minds inside a larger mind with divergent goals) is the one LLM swarms actually present (compare SoM §28.8, www.aurellem.org/society-of-mind/som-28.8.html).
Tension 2: merging versus diversity and redundancy. Hillis notes that nothing forces machine intelligences to stay distinct, so they may merge (PDF p.162). Minsky's robustness depends on duplication across separate agents (SoM §18.9, www.aurellem.org/society-of-mind/som-18.9.html), and his intelligence depends on “vast diversity” (SoM §30.8, www.aurellem.org/society-of-mind/som-30.8.html). A merged swarm loses both. Hillis treats merging as a likely drift, while Minsky treats diversity as the source of mind.
Tension 3: peers versus bosses. Hillis's nation-states recognize only one another as peers (PDF p.161). Their order is a heterarchy of sovereigns with no Builder above them. Minsky's agencies sit in hierarchies where some agent turns the others on (SoM §3.3, www.aurellem.org/society-of-mind/som-3.3.html). Heterarchy appears in Minsky only as mutual service, not rival sovereignty (SoM §3.4, www.aurellem.org/society-of-mind/som-3.4.html). Multi-swarm orchestration without a meta-orchestrator is Hillis's case, not Minsky's.
6. Seeds for open questions
- Does constraint adherence survive delegation in a coding stack? For a fixed set of repo tasks with explicit constraints (for example, “do not modify tests”, “do not touch
migrations/”), compare violation rates when Claude Code or Codex receives the task directly with the rates when the same task arrives through the orchestrator's decomposition. Does the delegation gap Ying et al. report for hazardous tasks also appear for ordinary engineering constraints?
- Policy neurons versus model neurons. Holding the worker models fixed, vary only the orchestration policy (success criterion, retry limit, escalation rule). Measure goal-substitution behaviours such as weakening tests, skipping checks, or declaring done early. Is the variance caused by policy changes larger than the variance caused by swapping worker models? A yes would support Hillis's claim (PDF p.160) that emergent goals live in policies.
- Firewall strength versus cross-check value. Vary what one checker sees of the other's work (nothing, final diff, full reasoning, shared memory). Measure the catch rate on seeded bugs and the correlation between the two workers' misses. Is there a level of sharing beyond which the two checkers behave statistically like one (Hillis's merging, PDF p.162)?
- Can a weak helmsman steer by debate? On disputed diffs, does a weaker orchestrator judging a structured debate between Claude Code and Codex pick the correct diff more often than it does when reviewing either diff alone? Does the result carry over from Khan et al.'s setting to code, where execution gives ground truth?
7. Sources
Book: John Brockman (ed.), Possible Minds: Twenty-Five Ways of Looking at AI (Penguin Press, 2019), PDF pp.158–166 (local extract in resources/book-text/).
All URLs accessed 2026-10-03.
Hillis
- zpr.pubpub.org/pub/0tdynxii/release/1 (Hillis et al., ZPR overview, 2023)
- www.oracle.com/news/announcement/ocw24-oracle-strengthens-organizations-cloud-security-posture-by-separating-network-security-from-network-architecture-2024-09-10/ (Oracle ZPR announcement, 2024; source for the 2024 shipping date; Tier B)
- tim.blog/2024/12/14/danny-hillis-kevin-kelly-transcript/ (Tim Ferriss Show #782 transcript, 2024)
- openalex.org/W3005941761 (“About the Underlay”, 2019)
Related work
- doi.org/10.1007/s00146-024-02063-2 (Elsner 2024)
- arxiv.org/abs/2502.14143 (Hammond et al. 2025)
- arxiv.org/abs/2503.13657v1 (Cemri et al. 2025)
- arxiv.org/abs/2609.27900v1 (Ying et al. 2026)
- arxiv.org/abs/2402.06782v4 (Khan et al. 2024)
- arxiv.org/abs/2312.09390v1 (Burns et al. 2023)
- arxiv.org/abs/2505.02709v1 (Arike et al. 2025)
- arxiv.org/abs/2608.10153 (CASE framework, 2026)
Reference pack (Minsky and cybernetics; see resources/reference-pack/minsky-wiener.md)
- www.aurellem.org/society-of-mind/som-1.1.html
- www.aurellem.org/society-of-mind/som-1.3.html
- www.aurellem.org/society-of-mind/som-3.1.html
- www.aurellem.org/society-of-mind/som-3.2.html
- www.aurellem.org/society-of-mind/som-3.3.html
- www.aurellem.org/society-of-mind/som-3.4.html
- www.aurellem.org/society-of-mind/som-6.4.html
- www.aurellem.org/society-of-mind/som-18.9.html
- www.aurellem.org/society-of-mind/som-28.8.html
- www.aurellem.org/society-of-mind/som-30.8.html
- www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf (Wiener 1960)
- pespmc1.vub.ac.be/REQVAR.html (Ashby, requisite variety)
- www.kybernetik.ch/dwn/Viable_System_Model.pdf (Beer 1984)
- cepa.info/fulltexts/1707.pdf (von Foerster 1979)