Possible Minds (2019)
Neil Gershenfeld, “Scaling”
AI's history is not a sequence of fashions but a sequence of threshold-type scaling laws (Shannon for communication, von Neumann for computation, depth for representation), and the next domain those laws will conquer is fabrication itself.
Scroll, or use the arrow keys, to move through the essay one step at a time. The plate paints each step as you reach it.
1. The argument
Step 1. The boom-bust cycles hide one continuous activity. Gershenfeld counts five AI booms (mainframes, expert systems, perceptrons, multilayer perceptrons, deep learning) and says they differ less than their marketing suggested: “Each of these stages was heralded as a revolutionary advance over the limitations of its predecessors, yet all effectively do the same thing: They make inferences from observations.” (PDF p.151)
Step 2. The right lens is how a method scales. “How these approaches relate can be understood by how they scale—that is, how their performance depends on the difficulty of the problem they're addressing.” (PDF p.151) Booms began with demos in narrow domains and busts came when the demos met less-structured problems. The steady progress underneath rests on the difference between linear and exponential functions (PDF p.151).
Step 3. Shannon's threshold theorem made communication scale. Analog telephone calls degraded with distance. Shannon showed that symbols behave differently because errors can be detected and corrected. If noise is above a threshold, errors are certain. But “if the noise is below a threshold, then a linear increase in the physical resources representing the symbol results in an exponential decrease in the likelihood of making an error” (PDF p.152). The mechanism is multiplicative: “Each symbol sent multiplies rather than adds to the certainty” (PDF p.152). Reliable networks then turned everyone into a data generator, which solved AI's knowledge-acquisition problem (PDF p.153).
Step 4. Von Neumann did the same for computation. Analog computers degraded with time, accumulating errors as they ran. Von Neumann's 1952 result showed “it was possible to compute reliably with an unreliable computing device by using symbols rather than continuous quantities ... That's what makes it possible to have a billion transistors in a computer chip, with the last one as useful as the first one.” (PDF p.153) That exponential growth in computing solved AI's second problem, processing the data.
Step 5. Depth did it for representation, at the price of exactness. The third problem was deriving rules without a programmer per task. “Wiener recognized the role of feedback in machine learning, but he missed the key role of representation.” (PDF p.153) A linear increase in network depth gave an exponential increase in expressive power, and learned representations constrain search against the curse of dimensionality. The cost is that the exact best answer is given up for one that is merely good enough (PDF p.154).
Step 6. Trust comes from testing interfaces, not from inspecting internals. Gershenfeld describes a project where AI pioneers kept moving the goalposts on data scientists who solved their problems without explaining the solutions. His reply: brains and chips alike are opaque inside, and “We come to trust (or not) brains and computer chips alike based on experience that tests them rather than on explanations for how they work.” (PDF p.154) Engineering is shifting from imperative to declarative design: describe what you want, then search automatically for designs that satisfy the goals and constraints (PDF p.154).
Step 7. Biology stores procedures, not blueprints. The Hox genes are genes that regulate other genes in developmental programs, and “Nothing in your genome stores the design of your body; your genome stores, rather, a series of steps to follow that results in your body.” (PDF p.155) The Hox genes mark out a productive region for evolutionary search, which he calls a parallel to search in AI. AI, by contrast, has a mind-body problem because it has no body (PDF p.155), and evolution changed form and program together.
Step 8. Digitize fabrication, and the same threshold logic applies to matter. After communication and computation, the next thing to digitize is fabrication. Digital materials are “constructed from a discrete set of parts reversibly joined with a discrete set of relative positions and orientations. These attributes allow the global geometry to be determined from local constraints, assembly errors to be detected and corrected” (PDF p.155). The parts can be generic: “What's interesting about amino acids is that they're not interesting.” (PDF p.155) Twenty or so part types are enough to build robots and computers. His lab's goal is an assembler that can assemble itself from the parts it assembles (PDF p.156), the experimental form of von Neumann's self-reproducing automata and Turing's morphogenesis. He calls this scarier than runaway AI because it moves intelligence into the physical world, and also more hopeful, because designs could be shared globally and produced locally (PDF p.156). The essay closes on a composition ladder from atoms to civilizations: “This grand evolutionary loop can now be closed, with atoms arranging bits arranging atoms.” (PDF p.157)
A note on voice. The headnote on p.149 is Brockman's. It reports Gershenfeld telling the group that he hated The Human Use of Human Beings and that Wiener, in worrying about automation, missed how access to the means of automation can empower people (PDF p.149). That is Brockman's account, so none of the steps above rests on it.
2. Related work since 2019
Gershenfeld and the Center for Bits and Atoms (CBA)
- “Self-replicating hierarchical modular robotic swarms” (Abdel-Rahman, Cameron, Jenett, Smith, Gershenfeld), 2022, Communications Engineering. www.nature.com/articles/s44172-022-00034-3 . A voxel-based material-robot system capable of “serial, recursive (making more robots), and hierarchical (making larger robots) assembly”. A planner chooses whether each robot builds, reproduces, or “evolves” into a bigger robot by searching the tree of alternatives for minimum timesteps. This is the essay's self-assembling assembler made concrete, and it is directly a swarm-composition result.
- “Discretely assembled mechanical metamaterials” (Jenett, Cameron, Tourlomousis, Parra Rubio, et al.), 2020, Science Advances. doi.org/10.1126/sciadv.abc9943 . A peer-reviewed instance of the “digital materials” programme of PDF p.155: engineered properties from a small set of discrete, assembled part types.
- “Algorithmic Approaches to Reconfigurable Assembly Systems” (Costa, Jenett, Kostitsyna, Abdel-Rahman, Gershenfeld, et al.), 2020, arXiv. arxiv.org/abs/2008.11925v1 . Addresses “the algorithmic problems in scaling” reconfigurable structures built by small robots that traverse the structure they modify: planning for many simple agents working on a shared lattice.
- “Hierarchical Discrete Lattice Assembly: An Approach for the Digital Fabrication of Scalable Macroscale Structures” (Smith, Richard, Kyaw, Gershenfeld), 2025, arXiv. arxiv.org/abs/2510.13686v1 . Argues that large-scale fabrication systems are “still typically complex, expensive, and unreliable” and proposes “simple robots and interlocking lattice building blocks” instead, which is the essay's claim that reliability should come from the discreteness of the parts rather than from the precision of the machine.
- “Speech to Reality: On-Demand Production using Natural Language, 3D Generative AI, and Discrete Robotic Assembly” (Kyaw, Smith, Jeon, Gershenfeld), 2024 (arXiv; ACM 2025). arxiv.org/abs/2409.18390v7 . A pipeline from speech through 3D generative AI to discrete robotic assembly. It is the essay's final line in miniature: a generative model (bits) whose output must be re-digitized into discrete, assemblable parts (atoms) before a robot can build it.
The threshold argument applied to AI systems (by others)
- “More Agents Is All You Need” (first authors: Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu), 2024, arXiv. arxiv.org/abs/2402.05120v2 . Finds that “simply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiated”, with gains that depend on task difficulty. This is the redundancy half of the threshold argument in its simplest form: more voters buy reliability, but only to a degree set by the task.
- “Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems” (first authors: Lingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis), 2024, NeurIPS 2024 (arXiv 2403.02419). proceedings.neurips.cc/paper_files/paper/2024/hash/51173cf34c5faac9796a47dc2fdd3a71-Abstract-Conference.html . Shows that majority-vote accuracy “can first increase but then decrease as a function of the number of LM calls”, because more calls help on “easy” queries and hurt on “hard” ones. This is the threshold of §3.2 measured per query: below it, voting helps; above it, voting drives the system toward the wrong answer, so a task mixing both kinds has an optimal number of voters.
- “Correlated Errors in Large Language Models” (Elliot Myunghoon Kim, Avi Garg, Kenny Peng, Nikhil Garg), 2025, ICML 2025 (PMLR 267). proceedings.mlr.press/v267/kim25e.html . Across more than 350 LLMs, “models agree 60% of the time when both models err”, and “larger and more accurate models have highly correlated errors, even with distinct architectures and providers.” This limits the argument directly: Shannon's multiplicative gain assumes independent errors (§3.3), and switching providers does not supply that independence.
3. Bearing on multi-agent and multi-swarm orchestration
The essay's strongest contribution to orchestration design is not an analogy. It is a set of conditions under which adding unreliable components makes a system more reliable. Shannon and von Neumann state them, and Gershenfeld repeats them carefully (PDF pp.152–153). Each one maps to a specific design decision in an orchestrator that dispatches to Claude Code and Codex and has them cross-check.
3.1 A cross-check is a repetition code, and two copies can only detect. Running the same task through two workers and comparing outputs is the simplest code. With two copies you can detect a disagreement but cannot correct it, because there is no majority. A pair therefore gives the orchestrator an error flag, not an answer. Correction needs either a third independent vote or a deterministic oracle (a test suite, a type checker, a reproducible benchmark) that plays the role of the decoder. In practice the oracle is the better decoder: it is cheap, and its errors are not correlated with the models' errors. One caution from the logged searches: a 2026 preprint, "Blind to the Pivotal Vote" (arxiv.org/abs/2608.06940v1), reports that adding a test-suite signal to an LLM judge panel "produced no distinguishable change in the panel's effective-vote count". Its title argues that such aggregate metrics miss the specific votes where verification helps. The claim that an oracle decorrelates errors should therefore be measured at the pivotal cases, not assumed from aggregate agreement.
3.2 The threshold is per task class, and above it more agents make things worse. The theorem says that redundancy helps only if each component's error rate is below threshold. Above it, Shannon found that errors are certain (PDF p.152). For majority voting, the threshold is a per-worker error rate below one half on that task. On a task class where both models are usually wrong (an unfamiliar API, a subtle concurrency bug), adding voters drives the majority toward the wrong answer. An orchestrator should estimate per-class error rates from its own logs and route below-threshold classes to voting, above-threshold classes to oracle-checked iteration or to a human.
3.3 Errors must multiply, which requires independence. The claim that each extra symbol multiplies rather than adds to the certainty (PDF p.152) holds only when failures are independent. Two LLMs trained on overlapping corpora are not independent coins (the same caveat appears in the reference pack, M15). Correlation shrinks the effective number of voters and raises the threshold. Heterogeneity is therefore a resource the orchestrator should spend deliberately: different model families, different prompts, different tool access, and above all different kinds of check (a model reviewer plus an executable test), so that their failure modes overlap as little as possible.
3.4 Re-digitize at every handoff. Shannon and von Neumann's other condition is that the state is carried by symbols that can be restored, not by continuous quantities that drift. Analog computing degraded with time and accumulated errors as it ran (PDF p.153). A pipeline where agents pass free-text summaries to each other is analog in exactly this sense: each hop adds paraphrase noise and nothing restores the signal. The digital alternative is to make every handoff a discrete artifact with a pass/fail check: a diff that applies, a build that succeeds, tests that pass, a JSON object that validates against a schema. The orchestrator's job at each boundary is restoration, snapping the state back to a valid symbol before the next stage, so that errors do not accumulate with chain length. This is the billion-transistor property, where the last component is as useful as the first (PDF p.153), applied to the hundredth agent in a pipeline.
3.5 Digital materials as a specification for work units. The four properties on PDF p.155 work as a design checklist for task decomposition:
| Digital-material property (PDF p.155) | Work-unit analogue |
|---|---|
| Discrete set of parts | Small set of task types with typed inputs and outputs |
| Reversibly joined | Each unit lands as a revertible commit or patch |
| Discrete relative positions | Fixed interfaces and file ownership, so units fit only one way |
| Global geometry from local constraints; errors detected and corrected | Integration correctness follows from per-unit contracts, so a bad unit is caught and replaced locally rather than discovered end to end |
The claim that twenty or so part types suffice (PDF p.155) suggests the orchestrator needs a small, stable vocabulary of worker roles more than a large zoo of specialists.
3.6 Multi-swarm: build, reproduce, or evolve. The 2022 swarms paper gives each robot three choices: build the structure, build another robot, or build a bigger robot, and picks the combination that finishes fastest. For a swarm of swarms the analogue is concrete: a worker can do the task, the orchestrator can spawn more peer workers, or it can spawn a sub-orchestrator that runs its own swarm. The threshold logic adds a constraint the robotics paper does not need to state. In a concatenated code each level must itself be below threshold, or the hierarchy amplifies error. A sub-orchestrator's output is a symbol to the meta-orchestrator and must be restored (verified) at that boundary before it is composed with others. Gershenfeld's closing ladder of composition (PDF p.157) only works because each level has its own way of holding its parts together.
3.7 Trust by interface. Step 6 (PDF p.154) argues for evaluating agents by testing their outputs rather than reading their reasoning. For cross-checks this means a reviewer agent should run the code and the tests, not only read the other agent's explanation, which shares that agent's blind spots. Declarative design (PDF p.154) is also the right shape for the orchestrator's prompt: state goals and constraints, and let the workers search.
4. Cybernetics
This is the most openly anti-cybernetic essay in the set. Gershenfeld says of Wiener's field, “I've never understood what that is” (PDF p.152). His charge is that Wiener saw feedback and missed the two things that made AI work: digital error correction and representation (PDF pp.152–153).
- Rejected: feedback as the central idea. Feedback is acknowledged and demoted. On his account representation, not regulation, solved the problem of rules (PDF p.153).
- Used without the name: restoration as regulation. Digital error correction is a regulator in the cybernetic sense. Each restoring stage compares the received state to a set of valid states and pushes it back, which is negative feedback applied per symbol. Minsky's difference-engine (SoM §7.8, M9, www.aurellem.org/society-of-mind/som-7.8.html) is the same structure. Gershenfeld's contribution is the scaling law, the claim that this regulation gets exponentially better for linear cost below threshold. That quantitative claim is absent from the cybernetic sources in the reference pack.
- Requisite variety, reframed. Ashby's law says that only variety can destroy variety (W4, pespmc1.vub.ac.be/REQVAR.html). Gershenfeld's twenty part types (PDF p.155) say the regulator's variety can come from combination of a few generic parts rather than from many specialised ones. For an orchestrator, the variety of its responses to worker failure (retry, re-route, re-decompose, escalate, revert) matters more than the number of distinct worker types.
- Two time scales. Wiener warned about two control operators on very different time scales acting together (W2, www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf). Gershenfeld's analogue is analog error accumulating over time (PDF p.153). A slow orchestrator over fast workers needs restoration at the workers' time scale, through automated checks, because it cannot regulate every step itself.
- Observation from outside. Trust built on experience that tests a system (PDF p.154) is a first-order, black-box stance: the observer stays outside the observed system. Second-order cybernetics (W6, cepa.info/fulltexts/1707.pdf) notes the catch for LLM swarms. When the tester is itself an LLM, the observer is made of the same material as what it observes, and its tests inherit correlated errors (§3.3).
- Wiener's automation worry, inverted. Per Brockman's headnote, Gershenfeld thinks Wiener missed that access to automation empowers people (PDF p.149). In the essay itself he says Wiener did not question assumptions about work that break down once consumption can be replaced by creation (PDF p.156).
5. Agreement and clash with Society of Mind
Agreement 1: generic, mindless parts. Gershenfeld's amino acids, deliberately uninteresting parts of which twenty types suffice (PDF p.155), match Minsky's programme: “minds are built from mindless stuff, from parts that are much smaller and simpler than anything we'd consider smart” (SoM §1.1, www.aurellem.org/society-of-mind/som-1.1.html). Both locate capability in composition, not in the components.
Agreement 2: robustness through redundancy. Minsky's argument that duplicating each function in ten agents makes total loss as unlikely as “ten tossed coins would all come up tails” (SoM §18.9, www.aurellem.org/society-of-mind/som-18.9.html) is a repetition-code argument in informal form. It is the same multiplicative logic as Shannon and von Neumann's threshold results (PDF pp.152–153). Both silently assume independent failures, which is the assumption that LLM swarms violate (§3.3).
Agreement 3: control by regulators that do no work. The Hox genes are genes that regulate other genes (PDF p.155) and store procedures rather than a design. Minsky's Builder “does no physical work but merely turns on Begin, Add, and End” (SoM §3.3, www.aurellem.org/society-of-mind/som-3.3.html), and Papert's Principle says growth comes from “new administrative ways to use what one already knows” (SoM §10.4, www.aurellem.org/society-of-mind/som-10.4.html). An orchestrator is a Hox gene in this sense: a short program over generic workers.
Tension 1: one principle versus many. Gershenfeld reads the whole of AI history as one principle, scaling, applied three times and about to be applied a fourth (PDF pp.151, 155, 157). Minsky says the opposite: “The power of intelligence stems from our vast diversity, not from any single, perfect principle.” (SoM §30.8, www.aurellem.org/society-of-mind/som-30.8.html) For an orchestrator this is a live design choice. Either scale one strong generalist with redundancy and checks (Gershenfeld), or compose many differently specialised agencies that “constantly challenge one another” (Minsky, same section).
Tension 2: explanation versus testing. Gershenfeld dismisses the demand to explain how a chess program plays: it plays chess, and we trust by testing interfaces (PDF p.154). Society of Mind is an attempt at exactly the explanation he waves away, how competence arises from the arrangement of agents (SoM §1.1, www.aurellem.org/society-of-mind/som-1.1.html). Minsky's B-brain also watches the A-brain's internal process, not its outputs: “connect it so that the A-brain is the B-brain's world!” (SoM §6.4, www.aurellem.org/society-of-mind/som-6.4.html) That is a supervisor that inspects internals. Gershenfeld would look only at external interfaces. An orchestrator needs both: output tests to decide correctness, and process monitoring (loops, stalls, thrash) to decide when to intervene.
Tension 3: the communication bottleneck. Minsky assumes that “the smaller an agency is, the harder it will be for other agencies to comprehend its tiny “language”” (SoM §6.12, www.aurellem.org/society-of-mind/som-6.12.html). Gershenfeld's history starts from Shannon removing the communication bottleneck by digital coding (PDF p.152). LLM agents sharing one natural language sit on Gershenfeld's side of this split, but natural language is an analog channel in his sense (§3.4). The bottleneck has moved from comprehension to fidelity.
6. Seeds for open questions
- Is there a measurable threshold in cross-checked coding swarms? For a fixed task class, plot consensus error against the number of heterogeneous workers (1, 2, 3, 5) and against per-worker error rate. Does it show threshold behaviour (exponential improvement below some per-worker error rate, degradation above it), and how far does inter-model error correlation move the threshold?
- Does re-digitizing at handoffs bound error accumulation with chain length? Compare an “analog” pipeline (free-text handoffs between agents) with a “digital” one (each handoff must pass an executable or schema check) as the number of sequential agents grows. Does the digital pipeline turn linear or compounding error growth into a flat curve, as von Neumann's result predicts for computation?
- Concatenated orchestration for swarms of swarms. If each orchestration level verifies the outputs of the level below, does end-to-end error fall like a concatenated code as depth grows? What is the cost-optimal choice among “build” (do the task), “reproduce” (more peer workers) and “evolve” (spawn a sub-orchestrator), borrowing the planner from the 2022 swarms paper?
- Digital-material work units. Do task decompositions that satisfy the four digital-material properties (discrete types, reversible joins, fixed relative interfaces, locally checkable correctness) reduce integration failures compared with free-form decompositions of the same tasks?
7. Sources
Book: John Brockman (ed.), Possible Minds: Twenty-Five Ways of Looking at AI (Penguin Press, 2019), PDF pp.149–157 (local extract in resources/book-text/).
All URLs accessed 2026-10-03.
Society of Mind sections (Minsky 1986, via the reference pack)
- www.aurellem.org/society-of-mind/som-1.1.html (M1, §1.1, minds built from mindless agents)
- www.aurellem.org/society-of-mind/som-3.3.html (M4, §3.3, managers that do no physical work)
- www.aurellem.org/society-of-mind/som-6.4.html (M7, §6.4, B-brains)
- www.aurellem.org/society-of-mind/som-6.12.html (M13, §6.12, small agents cannot share languages)
- www.aurellem.org/society-of-mind/som-7.8.html (M9, §7.8, difference-engines)
- www.aurellem.org/society-of-mind/som-10.4.html (M11, §10.4, Papert's Principle)
- www.aurellem.org/society-of-mind/som-18.9.html (M15, §18.9, robustness through duplication)
- www.aurellem.org/society-of-mind/som-30.8.html (M16, §30.8, intelligence as diversity)
Cybernetics (via the reference pack)
- pespmc1.vub.ac.be/REQVAR.html (W4, Ashby's law of requisite variety, Principia Cybernetica)
- www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf (W2, Wiener 1960, on two control operators at different time scales)
- cepa.info/fulltexts/1707.pdf (W6, von Foerster 1979, “Cybernetics of Cybernetics”)
Gershenfeld's work (Center for Bits and Atoms)
- www.nature.com/articles/s44172-022-00034-3 (Self-replicating hierarchical modular robotic swarms, 2022)
- doi.org/10.1126/sciadv.abc9943 (Discretely assembled mechanical metamaterials, 2020)
- arxiv.org/abs/2008.11925v1 (Algorithmic Approaches to Reconfigurable Assembly Systems, 2020)
- arxiv.org/abs/2510.13686v1 (Hierarchical Discrete Lattice Assembly, 2025)
- arxiv.org/abs/2409.18390v7 (Speech to Reality, 2024)
Other research (threshold and error-correction reasoning applied to AI systems)
- arxiv.org/abs/2402.05120v2 (More Agents Is All You Need, 2024)
- proceedings.neurips.cc/paper_files/paper/2024/hash/51173cf34c5faac9796a47dc2fdd3a71-Abstract-Conference.html (Are More LLM Calls All You Need?, NeurIPS 2024)
- proceedings.mlr.press/v267/kim25e.html (Correlated Errors in Large Language Models, ICML 2025)
- arxiv.org/abs/2608.06940v1 (Blind to the Pivotal Vote, 2026; caveat in §3.1)