Possible Minds, illustratedChris Anderson

Possible Minds (2019)

Chris Anderson, "Gradient Descent"

Book
John Brockman (ed.), Possible Minds: Twenty-Five Ways of Looking at AI (Penguin Press, 2019), ch. 14
Author
Chris Anderson
PDF pages
135–141 (p.135 is Brockman's headnote; the essay proper is pp.136–141)
Print pages
143–50 (from the book's index)
One-line thesis
Mosquitoes, physics, brains and neural networks all run one algorithm, descent along a gradient, and AI's real promise is to get us out of the local minima that evolution and human tradition have left us in.

Scroll, or use the arrow keys, to move through the essay one step at a time. The plate paints each step as you reach it.

About the author (Brockman's voice). Brockman's headnote (PDF p.135) describes Anderson as a former physicist and former editor-in-chief of Wired (2001–2012), co-founder and CEO of 3DR, a company Brockman says "helped start the modern drone industry" and now focuses on drone data software, which grew out of the open-source DIY Drones community. These are Brockman's framing, not Anderson's argument. Post-2019 facts about Anderson are in §2 and come from sources fetched in this run.

> Note on scope. The essay itself contains no drone, swarm, flocking or ant-colony material, > despite the author's background. Its only biological case is a single mosquito, and its only > multi-agent mechanism is restart-and-share among parallel optimization trials (PDF p.140). This > document does not import swarm material the essay does not contain; where swarms appear below, they > are the reader's systems or cited related work.

1. The argument

Step 1. A near-mindless agent can solve a hard search problem with a gradient and a fallback rule. The essay opens with a mosquito's "pursuit function": "First, move in a random direction. If the scent increases, continue moving in that direction. If the scent decreases, move in the opposite direction." (PDF p.136). The rule has an explicit recovery branch: when the scent is lost, move sideways until it is found again. What looks like radar "is actually just a sensitive nose with almost no intelligence at all" (PDF p.136).

Step 2. Gradient-following is universal. Anderson extends the rule from insects to water, cell membranes, planetary motion, chemistry and neural signalling, and then to decision-making: "But how do you explain more complex behavior, such as our ability to make decisions? The answer is just more gradient descent." (PDF p.137).

Step 3. Brains are feedback systems doing optimization. He reports that "science is coming around to the view that our brains operate the same way as any other complex system with layers and feedback loops" (PDF p.137), all pursuing optimization functions. Learning is correlating inputs with rewards and punishments and strengthening or weakening connections accordingly (PDF p.138).

Step 4. The limit is the local minimum, and there are exactly two ways out. Walking downhill gets you to the nearest valley, not home. To escape, "you either need a mental model (i.e., a map) of the topology, so you know where to ascend to get out of the valley, or you need to switch between gradient descent and random walks" (PDF p.138). The mosquito does the second.

Step 5. Neural-network training is the same loop, made explicit. Anderson gives a five-step recipe: define a cost function, run, perturb and run again, move in the downhill direction, repeat until no direction improves. Step 3 reads: "Change the values of the connections and do it again. The difference between those two results is the direction" (PDF p.139). The direction is estimated by comparing two trials, not computed analytically.

Step 6. Three escapes from local minima. (i) Parallel restarts with sharing: "Try lots of times with different random settings and share learning from each trial; essentially, you are shaking the system to see if it settles in a lower state." (PDF p.140). (ii) Stochastic gradient descent: stumble around a bit at random on the way down. (iii) Search for interesting features, kept in check by regularization, which favours features that are "likely to be real in nature, as opposed to artifacts or errors" (PDF p.140) over noise and illusion.

Step 7. Human knowledge is itself a local minimum. Local minima do not make AI less lifelike; life is often stuck in them too. On Go, "It took AIs less than three years to find out that we'd been playing it wrong all along" (PDF p.140), and chess turned out to reward strategies like early queen sacrifices, as if humans had been playing lower-dimensional versions of higher-dimensional games (PDF p.140).

Step 8. The programme: AI as a way out of evolution's basin, toward an unspecified goal. "We're going to rock ourselves out of local minima and find deeper minima, maybe even global minima." (PDF p.141). The essay closes on machines "forever descending the cosmic gradients to an ultimate goal, whatever that may be." (PDF p.141).

Assessment. The essay is a popular exposition, not a technical argument. Its strongest move is reductive (Step 2): everything is gradient descent. Its most useful content for a systems builder is the operational detail in Steps 4–6: the two-escape dichotomy (map or noise), the finite-difference form of the slope estimate, and the three named escape strategies, one of which is a population method. It leaves two things undefended: who defines the cost function (Step 5, item 1) and what the "ultimate goal" is (Step 8).

2. Related work since 2019

Anderson's own work since 2019

The logged searches turned up career moves and interviews, not research. No essay, paper or talk by Anderson since 2019 that revisits the gradient-descent argument, or that discusses multi-agent AI, was found in the logged searches. Dronecode (launched with the Linux Foundation in 2014, www.linuxfoundation.org/press/press-release/linux-foundation-and-leading-technology-companies-launch-open-source-dronecode-project ) and DIY Robocars (launched in beta in January 2017, diydrones.com/profiles/blogs/new-companion-site-for-diydrones-diyrobocars ) are pre-2019 origins. The only post-2019 evidence for either is their undated current pages (www.diyrobocars.com/ , px4.io/ecosystem/ecosystem-overview/ ).

  1. “Pioneering open source drones and robocars” (podcast interview with transcript), recorded 9 October 2019, published 18 October 2019, The Changelog #366 (logged at a mirror host). changelog-2025-05-05.fly.dev/podcast/366 . Anderson's own account of the 3DR arc: out of consumer hardware because “The Chinese did it better than us”, into drone-data software (“Now we're a SaaS company”), with Dronecode described as “sort of the Android of the space”, running PX4. Brockman's headnote on 3DR's software focus was therefore still accurate in late 2019; the timeline below records what changed after.
  2. “Inside Larry Page's Secret New Aviation Startup Dynatomics” (Hugh Langley), 2025, Business Insider. www.businessinsider.com/larry-page-new-aviation-startup-dynatomics-2025-1 . Reports that as Kittyhawk wound down in late 2022, its CTO Anderson began a startup that “explores ways to use AI and additive manufacturing to build aircraft”. This is the nearest logged evidence of Anderson applying AI himself, and it concerns design and manufacturing, not agent coordination.
  3. “Some personal news” (Anderson's LinkedIn post, sharing Encode's announcement of him as an advisor), 12 January 2026, LinkedIn (Tier C, used only for role and date). www.linkedin.com/posts/chrisanderson6_some-personal-news-this-is-in-addition-to-activity-7416525950563094528-DA7S . The announcement describes him as a Senior Advisor at Renaissance Philanthropy “focused on funding AI for Science and building at the intersection of AI and Advanced Manufacturing”. A 31 October 2025 post describes a London kick-off with ARIA and Encode (www.linkedin.com/posts/chrisanderson6_im-in-london-now-kicking-this-off-with-partners-activity-7390010183852154880-qka9 ).

Corporate timeline, from Tier C sources (dates and roles only). Anderson's LinkedIn profile says 3DR's software business was sold to Esri in 2020 and other assets to Kittyhawk in 2022 (www.linkedin.com/in/chrisanderson6 ). That dates the headnote's present tense, in which 3DR “now focuses on drone data software” (PDF p.135). A June 2021 Medium post reports Kitty Hawk acquiring 3D Robotics, with Anderson becoming COO (medium.com/self-driving-cars/kitty-hawk-acquires-3d-robotics-10370cf70532 ). Business Insider (item 2) and the LinkedIn posts give his Kittyhawk role as CTO. The sources conflict, and both titles are recorded here.

Work by others that tests or extends his argument

  1. “TextGrad: Automatic ‘Differentiation’ via Text” (Yuksekgonul, Bianchi, Boen, Liu, Huang, Guestrin, Zou), 2024, arXiv. arxiv.org/abs/2406.07496v1 . It turns Anderson's five-step loop (PDF p.139) into an engineering framework: “TextGrad backpropagates textual feedback provided by LLMs to improve individual components of a compound AI system”, with optimized variables “ranging from code snippets to molecular structures”. The paper is aimed at systems “orchestrating multiple large language models”, and it underpins the mapping in §3.1.
  2. “Improving Factuality and Reasoning in Language Models through Multiagent Debate” (Du, Li, Torralba, Tenenbaum, Mordatch), 2023, arXiv. arxiv.org/abs/2305.14325v1 . The positive case for Anderson's restart-and-share escape (PDF p.140) applied to LLM agents: multiple model instances “propose and debate their individual responses and reasoning processes over multiple rounds to arrive at a common final answer”, which improves reasoning and factual validity. The authors call this a “society of minds” approach.
  3. “Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate” (Liang, He, Jiao, Wang, et al.), 2023, arXiv. arxiv.org/abs/2305.19118v4 . Names the “Degeneration-of-Thought” problem: “once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.” This is a local minimum in Anderson's sense, and debate is offered as the shake that gets out of it. The paper also finds that “LLMs might not be a fair judge if different LLMs are used for agents”, which bears on cross-vendor cross-checks.
  4. “Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models” (Ferreira, Liu, Zheng), 2026, arXiv. arxiv.org/abs/2609.35875v1 . Across 23 models from eleven vendor families, the hypothesis that diversity drives debate gains “is rejected on every axis”. Mixed-model teams lose to majority votes over their own rosters, “with accuracy tracking member capability rather than heterogeneity”, and “nearly all of debate's benefit comes from the first exchange of answers”. This is the strongest logged evidence against assuming that different model families give Anderson's “different random settings” (PDF p.140), as §3.3 and seed 1 ask, at least for small models.
  1. “Collective gradient perception with a flying robot swarm” (Karagüzel, Turgut, Eiben, Ferrante), 2022, Swarm Intelligence. doi.org/10.1007/s11721-022-00220-1 . A swarm “composed of individuals lacking gradient sensing ability” follows the gradient of a scalar field through social interactions alone, shown in simulation and on real nano-drones. It is the multi-agent version of Anderson's mosquito (PDF p.136): the slope estimate comes from the population, not from any one agent's sensor.

3. Bearing on multi-agent and multi-swarm orchestration

The essay gives a vocabulary that maps onto an orchestrator/worker/cross-check system more precisely than most of the book does, because it describes an algorithm rather than an analogy. The mapping below takes the reader's setup (a Grok orchestrator dispatching coding tasks to Claude Code and Codex workers, which cross-check each other) and asks which parts of Anderson's loop each component implements.

3.1 The orchestrator is a zeroth-order optimizer, and the cross-check is its slope estimate

In Anderson's five steps, the optimizer never sees a gradient. It sees two scalar evaluations of a cost function and takes their difference as "the direction" (PDF p.139). An orchestrator is in exactly this position. Its cost function is whatever it can evaluate: tests passing, a type-check, a reviewer's verdict. It has no derivative of "code quality" with respect to "the next edit". So:

  • One worker run is a point, not a direction. A single Claude Code result tells the orchestrator the cost at one location in solution space. It cannot tell it which way is down.
  • Two workers on the same task give a finite difference. When Codex and Claude Code produce different patches and the orchestrator scores both, the difference between the scores and the diff between the patches together form a crude slope estimate. In Anderson's terms, the cross-check is step 3.
  • The estimate is only as informative as the perturbation is controlled. In numerical finite differences you perturb one coordinate at a time. Two LLM workers perturb everything at once (approach, file layout, naming, test strategy), so a score difference cannot be attributed to any one change. A design implication: ask the second worker for a minimal variant of the first worker's patch (one idea changed), not an independent solution, when the goal is to learn direction rather than to sample a new basin.

TextGrad (§2, item 4) makes this mapping literal. It treats LLM-written critiques as "textual gradients" that are propagated backwards through a compound system of LLM calls. On that view, a reviewer agent's critique is a gradient signal and the orchestrator is the optimizer applying it.

3.2 The mosquito's switching rule is an orchestration policy

Anderson's mosquito descends while it is in the scent plume and random-walks when it has lost the trail or hit an obstacle (PDF p.138). Translated into a dispatch policy:

Signal stateMosquitoOrchestrator action
Gradient present (failing tests decreasing, error messages changing)Keep flying in the improving directionIncremental fix-up by the same worker on the same branch
Gradient lost (same error repeated, all tests still red, reviewer verdict unchanged)Move sideways until the scent returnsRe-sample: different worker, different approach, fresh context
Obstacle (tool failure, permission wall)Bounce off with flexible wingsRoute around: change tool or escalate

The interesting engineering question is the detector for "gradient lost". The mosquito has a sensor threshold. An orchestrator needs a stagnation test over a window of attempts. That detector monitors the process of the worker, not the task content, which is exactly Minsky's B-brain (see §5, and SoM §6.4, www.aurellem.org/society-of-mind/som-6.4.html, where a B-brain notices that "A appears to be repeating itself").

Anderson's other escape, the "map" (PDF p.138), corresponds to an orchestrator that holds an explicit task decomposition. With a map it can deliberately go uphill (accept a temporarily worse state, such as a refactor that breaks tests) because it knows the valley beyond is lower. Without a map, the only way out of a basin is noise. An orchestrator that both re-plans and re-samples is using both of Anderson's escapes, and the open design question is when to prefer which.

3.3 The local minima of cross-checking

Anderson's warning is that the point where no direction improves is probably only a local minimum (PDF p.139). For cross-checking agents, the corresponding failure is agreement without correctness: both workers converge on the same plausible but wrong patch, and each reviewer approves the other's. Agreement is evidence of being in a basin, not the deepest one.

This is the multi-agent version of a point the reference pack already raises about Minsky's redundancy argument: duplication protects you only if failures are independent (SoM §18.9, www.aurellem.org/society-of-mind/som-18.9.html). Two LLMs trained on overlapping data share priors, and so share basins. The related-work items on multi-agent debate (§2, items 5–7) bear on exactly this: debate can improve accuracy (item 5), a single model reflecting on its own answer can lock into its first confident answer (item 6), and a budget-matched study finds that mixed-model teams gain little from heterogeneity as such (item 7).

Anderson's first escape, parallel restarts with sharing (Step 6, PDF p.140), has two parameters a builder controls:

  1. Diversity of starts. "Different random settings" for LLM workers means different model families, different prompts, different decompositions, not two samples of one model at different temperatures. A Claude Code/Codex pairing is better than two Claude Code instances on this criterion, but still not independent.
  2. Sharing cadence. Sharing too early (each worker sees the other's draft before committing) collapses the population into one basin. Sharing too late wastes runs. Anderson's recipe says to share learning from each trial without saying when, and that omission is the main open question for multi-swarm design.

3.4 The reviewer is a regularizer, and regularizers suppress the queen sacrifice

Anderson's third escape is the most interesting for cross-checks. Searching for interesting features risks illusions, so you regularize, judging features by "whether those kinds of features have been seen before (learned)" (PDF p.140) and preferring smooth over high-frequency ones. A code reviewer agent plays this role: it penalizes hacky, overfit, special-cased patches (code that passes the test but is "high frequency"). That is valuable, since it is the main defence against a worker gaming the test suite.

But the same essay, one paragraph later, celebrates AI finding strategies humans had never considered, such as early queen sacrifices (PDF p.140). A regularizer that rewards the "seen before" would have rejected those. In a cross-checking pair, reviewer strictness is therefore a dial that trades robustness against reward hacking against suppression of novel correct solutions. Anderson states both halves of this trade-off but does not connect them.

3.5 Multi-swarm: a population of descents

At the multi-swarm level (several orchestrators, each running a worker pool), Anderson's escape (i) becomes the whole architecture: each swarm is one descent from its own start, and a meta-orchestrator decides what to share across swarms and when. Three concrete consequences:

  • Swarms should differ in their starts, not only their workers. Two swarms with identical decompositions and model mixes are one descent run twice.
  • The meta-orchestrator's cost function must be shared, or the swarms are not optimizing the same landscape. If each orchestrator defines its own acceptance tests, cross-swarm comparison is comparing heights on different maps.
  • Recursion. Each swarm is, to the meta-orchestrator, a single black-box function evaluation. This is the recursion Beer's Viable System Model describes (reference pack W5), and the same zeroth-order limitation applies one level up.

4. Cybernetics

Used (implicitly). Anderson never uses the word "cybernetics", but his mosquito is the textbook case of purposeful behaviour as negative feedback in Rosenblueth, Wiener and Bigelow's 1943 sense (reference pack W1, www.cambridge.org/core/journals/philosophy-of-science/article/abs/behavior-purpose-and-teleology/73ACBBEC616CE78767088694F357D57B): a goal-directed system corrects its motion by the error signal from the goal. His description of the brain as a system of layers and feedback loops (PDF p.137) is Wiener's picture of interlocking feedback loops (W3), stated without attribution.

Extended. The essay adds something the 1943 framing leaves implicit: the failure mode of pure negative feedback. A system that only reduces error gets stuck where error stops decreasing. Read through Ashby's law of requisite variety (W4, pespmc1.vub.ac.be/REQVAR.html), a pure descender has one kind of move, so its repertoire is too small to regulate a landscape with many basins. Random walks and parallel restarts are variety injection. For an orchestrator, this suggests counting not only how many workers it has but how many distinct kinds of correction it can issue (retry, re-plan, swap worker, change decomposition, escalate to a human).

Left open. Wiener's central warning (W2, www.science.org/doi/10.1126/science.131.3410.1355) is that we "had better be quite sure that the purpose put into the machine is the purpose which we really desire". Anderson's step 1 (define a cost function, PDF p.139) is exactly where that purpose is put in, and his ending leaves the ultimate goal explicitly unspecified (PDF p.141). The essay therefore describes the feedback loop in detail and treats its set-point as somebody else's problem. In an orchestration system that set-point is the acceptance criterion, and choosing it is the orchestrator designer's most consequential act.

Second-order. The cost function is defined by an observer outside the loop. When the evaluator is itself an LLM, the observer is made of the same material as the optimized system, the situation von Foerster's second-order cybernetics addresses (W6, www.emerald.com/insight/content/doi/10.1108/03684920410556007/full/html). Anderson's essay has no account of this. Its regularizer judges by what has been seen before, which, when the judge is a model, means judged by the judge's own training distribution.

5. Agreement and clash with Society of Mind

Agreements

A1. Mind from mindless parts. Minsky's programme is to show "how minds are built from mindless stuff" (SoM §1.1, www.aurellem.org/society-of-mind/som-1.1.html). Anderson's mosquito is a miniature of the same claim: competent, adaptive behaviour from a nose and a three-branch rule with almost no intelligence (PDF p.136, quoted in Step 1). Both locate intelligence in organization and environment rather than in a clever part.

A2. Goals as difference reduction. Minsky's difference-engine "must contain a description of a desired situation" and subagents "aroused by various differences" between desired and actual (SoM §7.8, www.aurellem.org/society-of-mind/som-7.8.html). Anderson's five-step loop (PDF p.139) is a difference-engine with a scalar difference: the cost function is the description of the desired situation, and each step reduces the difference. Both are the cybernetic error-correcting loop (W1) in different vocabularies.

A3 (partial). Redundancy and accumulation. Minsky's robustness through duplication and "accumulation" of several ways to reach a goal (SoM §18.9, www.aurellem.org/society-of-mind/som-18.9.html) resembles Anderson's parallel restarts (PDF p.140). Both rely on many attempts rather than one good one. Both inherit the same caveat: the attempts must fail independently.

Tensions

T1. One principle versus diversity. This is the sharpest clash. Minsky writes that "The power of intelligence stems from our vast diversity, not from any single, perfect principle" (SoM §30.8, www.aurellem.org/society-of-mind/som-30.8.html). Anderson's Step 2 asserts precisely a single principle: decision-making is "just more gradient descent" (PDF p.137). For orchestration the difference matters. On Anderson's view, a better orchestrator is a better optimizer of one cost function. On Minsky's, it is a better manager of heterogeneous agencies that "constantly challenge one another" (§30.8), where conflict between agents is itself the source of competence.

T2. Local reward versus global credit. Anderson's learning rule rewards and punishes each connection by whether it contributed to the right answer (PDF p.139), and his mosquito is a purely local optimizer. Minsky distinguishes a Local scheme, which rewards agents for meeting their supervisor's goal, from a Global scheme, which rewards only contributions to top-level goals, and notes that local schemes license "I was only obeying the orders of my superior" (SoM §7.7, www.aurellem.org/society-of-mind/som-7.7.html). In a worker/orchestrator stack, a worker rewarded for passing its sub-task's tests is under the Local scheme. Anderson's framework has no level structure in which this distinction could even be stated.

T3. No negative knowledge. Minsky's censors and suppressors (SoM §27.2, www.aurellem.org/society-of-mind/som-27.2.html) store knowledge about what not to do, and intercept states that lead to bad ideas. Anderson's brain is a system of canals and locks with signals flowing downhill (PDF p.137): all flow, no stored prohibitions. His regularizer (PDF p.140) is the nearest analogue, but it is a soft prior over all solutions, not a learned list of specific bad moves. Orchestrators in practice need both: a regularizing reviewer and explicit censors (pre-dispatch guardrails, blocked commands).

T4. The map. Anderson's first escape from a local minimum is a mental model, a map (PDF p.138), but the essay never says where the map comes from if everything is gradient descent. Minsky's answer is administrative: growth comes from "acquiring new administrative ways to use what one already knows" (Papert's Principle, SoM §10.4, www.aurellem.org/society-of-mind/som-10.4.html). In orchestrator terms, the map is the decomposition and routing policy, a managerial structure, not another descent.

6. Seeds for open questions

  1. Shared basins across model families. When two LLM workers from different model families cross-check a coding task that contains a known trap, how often do they agree on the same wrong answer, and does that rate fall with the distance between the families (same vendor, different vendor, different training-data cut-offs)? This measures whether heterogeneous cross-checks are Anderson's "different random settings" or one basin sampled twice.
  1. Sharing cadence in parallel worker populations. In Anderson's restart-and-share escape, at what point in a task should parallel workers or swarms see each other's partial solutions? Is there a measurable threshold below which sharing causes premature convergence to one basin, and above which it wastes runs?
  1. Reviewer-as-regularizer. Can a reviewer agent's rejections be decomposed into "rejects overfit or test-gaming patches" and "rejects unfamiliar but correct patches" (the queen-sacrifice case), and can the second be reduced without increasing the first?
  1. Stagnation detection as a B-brain. Does an orchestrator that switches between incremental fix-up and fresh re-sampling based on a process-level "gradient lost" detector (no change in failing tests or error text over k attempts) outperform fixed retry budgets, and what window k works?

7. Sources

All accessed 2026-10-03. Book quotations are from the per-page extract of the PDF (resources/book-text/p135.txt–p141.txt), cited as (PDF p.NN).

Society of Mind sections (Minsky 1986, via the reference pack)

Cybernetics (reference pack W1, W2, W4, W6)

Anderson's work and career (Tier C sources flagged)

Other research

Step 1 of 40The drawing shows the step of the essay currently in view. The text beside it is the full essay.