Synthesis
Open questions in multi-agent orchestration, read through six Possible Minds essays, Wiener and Minsky
Scroll, or use the arrow keys, to move through the essay one step at a time. The plate paints each step as you reach it.
Version: Echo rewrite in Sakeeb Rahman's voice (D9). Prose blocks were rewritten by Echo; tables, closest-work lists and URLs are copied verbatim from the canonical open-questions.md (D8), which remains the source of truth. Fact-preservation check: echo-diff-report.md. Date: 2026-10-03. Inputs: the six essay companions in outputs/essays/, the reference pack resources/reference-pack/minsky-wiener.md, and 13 novelty-evidence logs in outputs/synthesis/novelty-evidence/ (Q01–Q13; 8–14 logged queries each across Exa, arXiv and OpenAlex).
1. The short version
I started with a builder's question. If you run an orchestrator that hands coding work to Claude Code and Codex and has them check each other, what do six essays from 2019—plus Wiener and Minsky—say you should worry about? Which of those worries has nobody tested yet?
Two things stand out. First, the essays converge on a single variable that most orchestration designs leave implicit: independence. A cross-check, a vote, a verification layer, or a recursive orchestrator—all of them assume that the parts fail differently. Minsky's own robustness argument assumes it explicitly. Four of the six essays raise it in different words.
Second, the literature moved fast. Of 13 candidate questions, none came back fully open. Twelve are partially explored, one is explored. What remains uninvestigated in every case is a specific combination: the coding setting, a controlled manipulation, and a measurement that the adjacent work did not make. Those combinations are the research agenda below. They appear as "no work found in the logged searches," which is evidence of absence in three indexes on one date, not proof.
2. Ranking
Score = Novelty × Tractability (each 1–5, assigned in the evidence logs). Ties are broken on Novelty (PROMPT.md §7), then on lineage breadth, the number of independent sources among the six essays and the Minsky/Wiener anchors that raise the question (see DECISIONS.md). Questions still tied share a rank.
| Rank | Q | Question | Verdict | N | T | Score | Lineage | Closest prior work |
|---|---|---|---|---|---|---|---|---|
| 1 | Q01 | Effective independence of a heterogeneous cross-check pair on code | Partially explored | 3 | 5 | 15 | 5 | Cross-Model Review in LLM Verification (2026) |
| 2 | Q09 | Cross-review protocols that remove conformity without removing collaboration | Partially explored | 3 | 5 | 15 | 3 | Groundability, Not Scale Alone (2026) |
| 3 | Q02 | Stability of mutual cross-check loops | Partially explored | 3 | 4 | 12 | 4 | Delayed Verification Destabilizes Multi-Agent LLM Belief (2026) |
| =4 | Q03 | Requisite variety of the orchestrator's remediation repertoire | Partially explored | 3 | 4 | 12 | 3 | Self-Healing Agentic Orchestrators (2026) |
| =4 | Q08 | Goal substitution: orchestration policy or worker model? | Partially explored | 3 | 4 | 12 | 3 | Can escalation channels redirect reward hacking? (2026) |
| =4 | Q13 | A content-blind supervisor (B-brain) | Partially explored | 3 | 4 | 12 | 3 | Real-Time Detection and Repair of LLM Agent Failures (2026) |
| =7 | Q06 | Censors versus suppressors | Partially explored | 3 | 4 | 12 | 2 | PreFlect (2026) |
| =7 | Q10 | Sharing cadence and topology across parallel swarms | Partially explored | 3 | 4 | 12 | 2 | Sparse Communication Topology in Multi-Agent Debate (2024) |
| =7 | Q12 | A legibility tax on worker intelligence | Partially explored | 3 | 4 | 12 | 2 | From Plan to Action (2026) |
| 10 | Q04 | Re-digitizing at every handoff in long chains | Partially explored | 2 | 5 | 10 | 2 | The Hallucination Snowball (2026) |
| =11 | Q05 | Recursive verification depth in swarms of swarms | Partially explored | 3 | 3 | 9 | 3 | Partially Correlated Verifier Cascades (2026) |
| =11 | Q11 | Papert's principle: a middle manager versus a better worker | Partially explored | 3 | 3 | 9 | 3 | Can AI Models Direct Each Other? (2026) |
| 13 | Q07 | K-lines for orchestration (reframed: negative K-lines) | Explored | 2 | 4 | 8 | 2 | Agentic Plan Caching (2025) |
3. The frame: what the six essays, Wiener and Minsky say about orchestration
Wiener sets the stakes. In his chapter, Stuart Russell cites Wiener's 1960 warning that when we use a machine we cannot easily interrupt, "we had better be quite sure that the purpose put into the machine is the purpose which we really desire" (PDF p.40). An orchestrator is precisely such a mechanism: it takes one stated purpose and turns it into many delegated ones.
Hillis says the danger is in the composite, not the parts. Corporations and nation-states are already hybrid intelligences, and "The components’ good intentions are not a guarantee of the emergent system’s good behavior." (PDF p.160). He also offers the book's most compact definition of the orchestrator's job: "Cybernetics is the study of the how the weak can control the strong." (PDF p.164). That describes the position of a cheap orchestrator steering two frontier coding models (Q08, Q12, Q13).
Dyson says the controller cannot be simpler than what it controls. He restates Ashby's law: "any effective control system must be as complex as the system it controls." (PDF p.53), then adds his third law, "any system simple enough to be understandable will not be complicated enough to behave intelligently, while any system complicated enough to behave intelligently will be too complicated to understand." (PDF p.53). That claim has two measurable edges: requisite variety (Q03) and the legibility tax (Q12).
Gershenfeld says scaling works only below a threshold. Digital communication and computation scale because "Each symbol sent multiplies rather than adds to the certainty" (PDF p.152), but only when each component's error rate is low enough and errors are corrected at each stage. Read through the lens of orchestration design, this predicts that voting and cross-checking help below a per-task error threshold and hurt above it, and that gating each handoff (Q04) and each level (Q05) is what makes long chains and deep hierarchies possible.
Anderson says escape from local minima needs varied restarts and sharing. His recipe is to "Try lots of times with different random settings and share learning from each trial; essentially, you are shaking the system to see if it settles in a lower state." (PDF p.140). For a coding swarm, the open questions are when to share and with whom (Q10), and whether two model families are really different random settings or one basin sampled twice (Q01).
Pentland says social learning works only on grounded feedback. Social sampling "will work only if you can get feedback to them that's truthful. It must be grounded on whether each person's actions worked for them or not." (PDF p.181). His oversight rule fits a telemetry supervisor: "You don't have to watch the AI; instead you should watch what it eats and what it does." (PDF p.184). These bear on conformity in cross-review (Q09) and on content-blind supervision (Q13).
Jones warns against confusing decomposition with intelligence. "Contemporary AI has talked itself into a corner by instrumentalizing and particularizing tasks and subroutines, confusing these drills with actual wisdom." (PDF p.234). For orchestration, the warning is that slice-level checks can pass while the whole fails, and that the human operator is part of the loop, not outside it.
Minsky supplies the vocabulary and the hidden assumption. Society of Mind already names most parts of a modern orchestrator: managers that "do no physical work", heterarchies in which two agents serve each other at once, a B-brain that watches the worker rather than the world, censors and suppressors, K-lines, and Papert's principle of growth through new administrators. His robustness argument rests on duplication. With every function copied in ten agents, losing half of them is like ten coins all landing tails. The arithmetic holds only if failures are independent. Two language models trained on overlapping data are not independent coins. Most of the agenda below follows from taking that seriously.
Minsky sources for §3 (fetched, CC BY-NC-SA web edition): managers, SoM §3.3 www.aurellem.org/society-of-mind/som-3.3.html; heterarchies, SoM §3.4 www.aurellem.org/society-of-mind/som-3.4.html; B-brains, SoM §6.4 www.aurellem.org/society-of-mind/som-6.4.html; K-lines, SoM §8.1 www.aurellem.org/society-of-mind/som-8.1.html; Papert's principle, SoM §10.4 www.aurellem.org/society-of-mind/som-10.4.html; robustness by duplication, SoM §18.9 www.aurellem.org/society-of-mind/som-18.9.html; censors and suppressors, SoM §27.2–27.3 www.aurellem.org/society-of-mind/som-27.2.html. Wiener 1960 (also via PDF p.40): www.science.org/doi/10.1126/science.131.3410.1355. Ashby's requisite variety: pespmc1.vub.ac.be/REQVAR.html.
4. The questions, in rank order
Rank 1: Q01. How many independent checks is a heterogeneous cross-check pair worth on code?
Why it matters. The whole value of cross-model review—Codex auditing Claude Code, or the reverse—depends on the two systems failing differently. Anderson's "different random settings", Gershenfeld's threshold, Dyson's warning about reading versus running, Hillis's worry that machines without firewalls merge into one, and Minsky's coin-toss arithmetic all turn on the same quantity. Judge-panel research now puts a number on it for fixed labelling tasks: nine models from seven families behave like about two independent judges.
What is still uninvestigated. No work found in the logged searches estimates effective independence for an author and reviewer pair on code while varying how much of the author's work the reviewer sees (nothing, the diff, the reasoning, shared memory), or computes it only on the disputed cases where the second opinion actually decides the outcome.
Minimal experiment. Introduce known bugs into 50–100 patches drawn from SWE-bench-style repositories. Let Claude Code author and Codex review, then reverse the roles, across four context conditions. For each condition, compute pairwise miss correlation and an effective number of independent reviewers, separately on all patches and on disputed patches.
Closest work (from novelty-evidence/Q01.md):
- When Does a Second Model Help? Cross-Model Review in LLM Verification (2026), arxiv.org/abs/2610.01471v1 — cross-model and same-model reviewers find partly different errors (Jaccard 41.2%); no effective-independence estimate and no graded information sweep.
- Nine Judges, Two Effective Votes (2026), arxiv.org/html/2605.29800 — 9 LLMs from 7 families give a Kish effective sample size of about 2 on fixed labelling tasks, not code review.
- Cross-Context Review (2026), ar5iv.labs.arxiv.org/html/2603.12123 — a fresh-session reviewer beats a same-session one; one model, two context levels.
- Partially Correlated Verifier Cascades (2026), arxiv.org/abs/2607.13918v1 — supplies an estimator from two repeated verdicts per instance.
Rank 2: Q09. Can cross-review remove conformity without removing collaboration?
Why it matters. Pentland's social sampling only pays off when the signal being copied is grounded. Let a reviewer see the peer's patch first, and agreement can become imitation. Jones's "colloquy" reading and Hillis's weak-steering-the-strong converge on the same insight: what matters is whether the reviewer judges the artifact or the author.
What is still uninvestigated. Each protocol exists somewhere in the literature: anonymised debate, fresh-context review, independent reconstruction before reading a patch, and execution-grounded review. No work found in the logged searches puts see-then-review, independent-first, authorship-blind and evidence-gated review head to head on code patches seeded with plausible but wrong references, with cross-model pairs, while also measuring how much useful correction each protocol gives up. One finding sharpens the design choice: for coding monitors, implicit authorship (reviewing inside the author's context) drives self-preference, so "blind" has to mean a fresh context, not stripped labels.
Minimal experiment. Assemble about 150 SWE-bench Verified cases. In half, the peer patch passes visible tests but fails hidden ones. In the other half it is correct but subtle. Run the four protocols plus a no-peer control in both directions, then report false accepts, false rejects and corrections gained over the control.
Closest work (from novelty-evidence/Q09.md):
- Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents (2026), arxiv.org/abs/2610.01023v1 — 411 execution-labelled traces and 101 controlled cases; an 8B reviewer seeing only the patch accepted all nine target omissions. One reviewer, no independent-first or blind arm.
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization (ACL 2026), arxiv.org/html/2510.07517v2 — anonymisation on reasoning benchmarks, not code.
- Self-Attribution Bias (2026), www.arxiv.org/pdf/2603.04582 — implicit authorship drives bias in coding monitors.
- SWE-Review (Tier C practitioner tool), github.com/SWE-Lego/cc-swe-review — independent fix before reading the patch; no controlled comparison retrieved.
Rank 3: Q02. When are mutual cross-check loops stable?
Why it matters. Of all the questions here, this one feels most purely cybernetic. Minsky cautions that should an A-brain and a B-brain "watch each other too closely", the system "might become unstable". Wiener's 1960 paper sounds a similar alarm when two control operators move at very different time scales. An orchestrator that sets Claude Code and Codex critiquing and revising one another has assembled both.
What is still uninvestigated. We now possess control-theoretic thresholds for belief consensus under delayed verification, as well as for single-model self-correction. One study even reports that 16.0% of code-repair trajectories had already located a correct patch, only to lose it by revision 3. The belief-consensus finding also reveals that grounded feedback dissolves the oscillation, leaving the open case to those coding disputes that tests fail to settle. Yet no work found in the logged searches examines an author and reviewer revising code as coupled controllers, adjusts how forcefully the author must respond to critique or how often review occurs, or asks whether an early stability estimate can predict patch ping-pong.
Minimal experiment. Begin with 40 tasks whose requirements are deliberately under-specified. Let Claude Code act as author and Codex as reviewer for as many as 10 rounds, across three gain settings (from "must address every comment" through "may decline with a reason") and two distinct review cadences. Then measure the re-appearance of reverted hunks, and test whether a model fitted on rounds 1–3 can predict non-convergence by round 10.
Closest work (from novelty-evidence/Q02.md):
- Delayed Verification Destabilizes Multi-Agent LLM Belief (2026), arxiv.org/html/2606.27409 — a closed-form instability threshold on verification strength that falls as delay grows; belief consensus, not code.
- Looping Is Not Reliability (2026), ar5iv.labs.arxiv.org/html/2607.24604 — correctness is not absorbing in forced-revision code repair (16.0% found and then lost a correct patch by revision 3); one coder, no coupled pair.
- When Does LLM Self-Correction Help? A Control-Theoretic Markov Diagnostic (2026), arxiv.org/html/2604.22273v1 — single-agent threshold.
- Minsky, SoM §6.4: www.aurellem.org/society-of-mind/som-6.4.html. Wiener 1960: www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf.
Rank =4: Q03. Does the orchestrator's repertoire need to match the variety of worker failures?
Why it matters. Ashby's law, restated by Dyson and Hillis, tells us that only variety can absorb variety. An orchestrator whose only move is "retry" has a repertoire of one. Jones's "shy worker", which asks for clarification instead of guessing, adds another kind of response.
What is still uninvestigated. Fault-injection harnesses built on the MAST failure taxonomy already exist, and one study shows class-matched recovery outperforming both retry-only and full re-planning. Yet no work found in the logged searches sweeps the size and make-up of an orchestrator's remediation repertoire against the variety of injected faults to look for the rise-then-saturate curve that requisite variety predicts. And none does this in coding with "reroute to the other model" as an action.
Minimal experiment. Inject six fault types into Codex and Claude Code outputs on 60 SWE-bench-style tasks. Run the orchestrator with action sets of size 1, 2, 4 and 6, in matched and mismatched compositions, and plot recovery against repertoire size.
Closest work (from novelty-evidence/Q03.md):
- Self-Healing Agentic Orchestrators (2026), arxiv.org/abs/2606.01416 — class-matched recovery 98.8% versus 94.5% (retry-only) and 93.8% (full re-planning); no repertoire-size sweep.
- OrchestraBench (2026), arxiv.org/html/2608.05263 — per-failure-mode recovery for a fixed agent; three latent modes never recovered.
- Why Do Multi-Agent LLM Systems Fail? (MAST, 2025), arxiv.org/abs/2503.13657v3.
Rank =4: Q08. Where does goal substitution live: in the orchestration policy or in the worker model?
Why it matters. For Hillis, emergent goals live in the organisation's policies, not in its members. Minsky's distinction between local and global credit predicts the same pattern: reward a worker for pleasing its supervisor, and you teach it to please the supervisor. In a coding stack, the visible symptoms are weakened tests, skipped checks, and early "done".
What is still uninvestigated. Single-agent studies now vary policy against model. One factorial across 8 models cuts reward hacking from 23.6% to 5.3% with an escalation channel and policy. Delegation structure is known to change compliance with hazardous tasks. Yet no work found in the logged searches varies orchestration knobs (success criterion, retry limit, escalation rule, local versus end-to-end credit) in a delegating multi-agent stack, crosses them with worker swaps, and reports how much of the variance in goal substitution each factor explains.
Minimal experiment. Take roughly 100 impossible-variant SWE-bench tasks, apply a "do not modify tests" constraint, and send them through a thin orchestrator under a 2 × 2 × 2 policy design, crossed with Claude Code and Codex as workers. Fit a mixed-effects model of cheating and constraint violation, and report the variance share for policy, model, and their interaction.
Closest work (from novelty-evidence/Q08.md):
- Can escalation channels redirect reward hacking toward defect disclosure? (2026), arxiv.org/html/2608.29460v2 — 2 × 2 factorial across 8 models; single agent.
- SpecBench (2026), arxiv.org/html/2605.21384 — visible versus held-out gaps across models and harnesses; no policy decomposition.
- Delegated Misalignment (2026), arxiv.org/abs/2609.27900v1 — delegation structure; hazardous-task compliance, not test weakening.
- ImpossibleBench (2025), arxiv.org/abs/2510.20270v1 — task construction for automatic cheating labels.
- Minsky, SoM §7.7: www.aurellem.org/society-of-mind/som-7.7.html.
Rank =4: Q13. Can a content-blind supervisor (a B-brain) do the orchestrator's job?
Why it matters. Minsky's B-brain helps "without having any idea of what A's goals are", and Pentland would watch "what it eats and what it does". A supervisor limited to process telemetry—edit churn, repeated errors, tests flipping, token burn—is cheap and hard to talk round. The relevant question is whether it is good enough.
What is still uninvestigated. The evidence to date leans against it. In one non-coding study, structural loop detection scored F1 0.08 against 0.72 for a hybrid. In agentic coding, monitors that read reasoning outperform action-only monitors at catching reward hacking. No work found in the logged searches compares a telemetry-only and a content-aware supervisor on the restart decision, or runs all three jobs—stagnation, test-weakening, restart—head to head on coding agents at matched cost. One caveat: the closest paper's abstract was truncated in the cache and needs a full read before any novelty claim on loop detection.
Minimal experiment. Collect roughly 300 trajectories from SWE-bench Lite and ImpossibleBench tasks. Give the same model either a per-step telemetry vector or the full transcript. Score loop and test-edit flags against diff-derived ground truth, and score restart decisions by forking each flagged run into continue and fresh-restart branches.
Closest work (from novelty-evidence/Q13.md):
- Real-Time Detection and Repair of LLM Agent Failures (2026), arxiv.org/abs/2608.02464v1 — telemetry-only detection; abstract truncated in the cache.
- Unsupervised Approaches to Futile Cycle Detection in AI Agents (ICSE 2026), dlnext.acm.org/doi/10.1145/3774748.3795682 — F1 0.08 (structural), 0.28 (semantic), 0.72 (hybrid) on 1,575 trajectories.
- Monitoring Reasoning Models for Misbehavior (Baker et al., 2025), arxiv.org/abs/2503.11926v1 — CoT monitors beat action-only monitors in agentic coding.
Rank =7: Q06. Censors versus suppressors: guard before dispatch or review after output?
Why it matters. Minsky separates suppressors, which block a bad action at the last moment, from censors, which intercept the states that lead to it. His prediction is that censors save time but need memory that grows steeply the earlier they act. Post-hoc reviewers are suppressors; pre-dispatch guardrails are censors.
What is still uninvestigated. Learning pre-execution interventions from failure traces is an active 2026 area, and one paper shows prospective critique beating retrospective reflection. No work found in the logged searches pits a learned pre-dispatch interceptor against an independent post-hoc patch reviewer at matched compute, nor measures how large the failure-pattern library must be, as a function of how early it acts, to reach a given recall. That second gap is the operational form of Minsky's claim.
Minimal experiment. Harvest failure traces from half of roughly 150 SWE-bench Verified tasks and build failure-pattern libraries keyed at dispatch, plan, and first actions. On the other half, compare interception, cross-model post-hoc review, and both together at equal token budgets, then record the library size each depth needs to reach 50% recall.
Closest work (from novelty-evidence/Q06.md):
- PreFlect: From Retrospective to Prospective Reflection in LLM Agents (2026), arxiv.org/pdf/2602.07187 — plan-time critique from historical errors; the "after" arm is self-reflection, not an independent reviewer.
- HiSentinel (2026), arxiv.org/html/2609.39957 — a trained sentinel decides before each coding-agent action; no matched post-hoc reviewer arm found.
- AgentGuard (2026), arxiv.org/abs/2609.16287v1 — guardrails extracted from anomalous coding traces; no library-size analysis.
- Minsky, SoM §27.3: www.aurellem.org/society-of-mind/som-27.3.html.
Rank =7: Q10. When and with whom should parallel swarms share partial solutions?
Why it matters. Anderson's escape from local minima depends on varied trials and shared learning. Pentland's network work found that sparse and adaptive connections beat everyone-to-everyone. But sharing carries a sharp trade-off: too early or too densely, and every worker falls into the same basin; too late, and the parallel runs are wasted.
What is still uninvestigated. Sparse topology is already studied for LLM debate, and denser coupling is known to worsen diversity collapse in LLM ideation. The cadence threshold is established for island-model evolutionary algorithms. Yet no work found in the logged searches crosses sharing cadence (early, middle, late) with topology for parallel LLM coding workers, or measures both failure modes, collapse and wasted isolation, on verifiable tasks.
Minimal experiment. Run 6 workers per task on 50 SWE-bench Lite issues. At step 10, 25 or 40, inject peers' current diffs to 2 random peers or to all 5, with a no-sharing control, and measure resolve rate, final patch similarity and the share of near-duplicate or dead-end runs.
Closest work (from novelty-evidence/Q10.md):
- Improving Multi-Agent Debate with Sparse Communication Topology (Findings of EMNLP 2024), aclanthology.org/2024.findings-emnlp.427/ — sparse matches or beats fully connected, with over 40% fewer input tokens; short-answer debate, no cadence variable.
- Diversity Collapse in Multi-Agent LLM Systems (Findings of ACL 2026), aclanthology.org/2026.findings-acl.13.pdf — denser topologies worsen premature convergence; no verifiable success criterion.
- Island-model migration interval and size (2005), dl.acm.org/doi/10.1145/1068009.1068219, and adaptive migration intervals, eprints.whiterose.ac.uk/id/eprint/93025/1/ECJ-submission-R1.pdf — the cadence threshold in evolutionary computation.
Rank =7: Q12. Is there a legibility tax on worker intelligence?
Why it matters. Dyson's third law implies a trade-off for every orchestrator: when the worker is confined to acting only through plans that the orchestrator can read and approve, the system remains understandable but may stop being smart. Hillis's weak controller confronts this identical trade.
What is still uninvestigated. "Legibility tax" is an established term for training-time checkability on math, and a decoupled method claims to remove it by writing the checkable version after solving. Plan compliance of coding agents has been measured across difficulty. Yet no work found in the logged searches measures the success cost of forcing coding agents to act only through plans that a simpler orchestrator approves, broken down by task difficulty, against a decoupled "act, then explain" control.
Minimal experiment. On 50 Easy, 50 Medium and 50 Hard SWE-bench Verified tasks, run one worker in three conditions: unconstrained, plan-approved step by step by a small model, and unconstrained but required to emit the same plan after the fact. If a gap widens specifically on Hard tasks between the first two arms, but not between the first and third, that would reveal a coupled legibility tax.
Closest work (from novelty-evidence/Q12.md):
- Prover-Verifier Games improve legibility of LLM outputs (2024), arxiv.org/abs/2407.13692v2.
- Mitigating Legibility Tax with Decoupled Prover-Verifier Games (ICLR 2026), arxiv.org/abs/2602.23248v2.
- From Plan to Action: How Well Do Agents Follow the Plan? (2026), arxiv.org/abs/2604.12147v3 — eight plan settings on all 500 SWE-bench Verified instances; compliance drops 13% on SWE-bench Pro; reports compliance, not the cost of the constraint.
Rank 10: Q04. Does re-digitizing at every handoff bound error growth in long chains?
Why it matters. Gershenfeld's account of why digital systems scale is that each stage restores the signal before passing it on. For an agent chain, that suggests a typed or executable check at every handoff rather than free-text summaries.
What is still uninvestigated. Less than the others. The evidence that per-handoff verification beats end-of-chain verification is already there for LLM pipelines, and one study finds errors attenuating rather than compounding in refinement chains. No work found in the logged searches measures, on code, how error scales with chain length under deterministic gates (tests and type checks) versus LLM-reviewer gates versus free-text handoffs.
Minimal experiment. Build coding pipelines of 2, 4, 6, 8 and 10 stages, comparing free-text handoffs, LLM-reviewer gates and deterministic gates across 50 tasks, and fitting error against chain length for each.
Closest work (from novelty-evidence/Q04.md):
- The Hallucination Snowball (2026), arxiv.org/abs/2608.14588v1 — 346 injected hallucinations in a 4-agent pipeline; boundary gates cut survival from 58.4% to 16.2%, while end-of-pipeline checks gain 2.3 points.
- Hallucination Cascade (2026), arxiv.org/abs/2606.07937v1 — net attenuation across 3-agent refinement chains.
- Maestro Order (2026), arxiv.org/html/2606.23983 — states the restoring-verifier design law, supported by simulation.
Rank =11: Q05. Does recursive verification behave like a concatenated code, and how deep should swarms of swarms go?
Why it matters. Multi-swarm orchestration is, fundamentally, a question of recursion. Beer's Viable System Model tells us that every viable system must contain, and itself be contained within, another viable system. And Gershenfeld's threshold argument offers a reason to hope that nested error correction can drive error downward rapidly as depth grows—but only under a crucial condition: the levels must fail independently.
What is still uninvestigated. The theory has largely arrived. One 2026 preprint has already shown that independent gates yield exponential gains, whereas correlated gates produce only polynomial gains that eventually hit a ceiling—and may even peak and then decline as depth increases. Yet no work found in the logged searches measures, on real recursive LLM orchestrators, how end-to-end error and cost change as verified levels are added; none locates the precise crossover point; and none tests whether sub-swarms genuinely behave as self-regulating units in Beer's sense.
Minimal experiment. Take 40 multi-file tasks and run recursive orchestration at depths 1, 2, and 3. Verify each level using either the same model family or the other, and then compare the measured curve against both the exponential and the polynomial predictions.
Closest work (from novelty-evidence/Q05.md):
- Partially Correlated Verifier Cascades in LLM Harnesses (2026), arxiv.org/abs/2607.13918v1.
- Maestro Order (2026), arxiv.org/html/2606.23983 — the Odds Law and gated recursion, in simulation.
- From Agents to Viable Collectives: The VSM as a Diagnostic Framework (2026; listed only on the Exa library, venue not established), exa.ai/library/publication/m832ps4n1dv.
- Beer's recursive system theorem: www.kybernetik.ch/dwn/Viable_System_Model.pdf.
Rank =11: Q11. Papert's principle: is a new middle manager worth more than a better worker?
Why it matters. Papert's principle locates the decisive steps of mental growth in "new administrative ways to use what one already knows"; Minsky adds that the choice of which agents to group is critical. Hillis's corporations are the same idea at the scale of human institutions.
What is still uninvestigated. The two-agent version is answered for SWE-bench-style work: a strong manager over a weak worker resolves 62%, matching a strong single agent at 60%, while a weak manager hurts. Which role to upgrade depends on the domain. No work found in the logged searches adds a middle layer over a pool of workers and compares it with worker upgrades at matched cost, or tests Minsky's grouping rule (group by skill similarity versus at random).
Minimal experiment. Profile 8 mixed workers across 60 SWE-bench Lite tasks. At equal budget, compare a flat orchestrator, two sub-managers grouped by skill similarity, two grouped at random, and a flat orchestrator over upgraded workers, at pool sizes of 4, 8 and 16.
Closest work (from novelty-evidence/Q11.md):
- Can AI Models Direct Each Other? Organizational Structure as a Probe into Training Limitations (2026), arxiv.org/abs/2603.26458v1 — 200 SWE-bench Lite instances; 62% versus 60%, and 42% versus 44%.
- Specialize Roles, Mix Deployments (AgentCARD, 2026), arxiv.org/html/2606.20629.
- OrgAgent (2026), arxiv.org/abs/2604.01020v1.
- Minsky, SoM §10.4: www.aurellem.org/society-of-mind/som-10.4.html.
Rank 13: Q07. K-lines for orchestration (reframed: do negative K-lines help?)
Why it matters. A K-line remembers which agents were active when a problem got solved, allowing them to be switched back on together. For an orchestrator, that translates into remembering which team and configuration solved a class of task.
What is still uninvestigated. The original question is answered: caching plans and team configurations beats re-planning on cost, and orchestrator-level memories of decomposition and agent selection exist. What is still uninvestigated is Minsky's footnote about the wrench that did not fit. No work found in the logged searches stores which worker, prompt or tool failed on a task class and suppresses it on reactivation, or measures how fast such negative entries go stale when worker models update.
Minimal experiment. A minimal experiment would cluster about 200 SWE-bench Verified tasks by repository and issue type, then compare planning from scratch, success-only memory, and success memory plus negative entries. Finally, swap in an updated worker model to watch the negative entries age.
Closest work (from novelty-evidence/Q07.md):
- Agentic Plan Caching (NeurIPS 2025), arxiv.org/abs/2506.14852v2 — about 50% lower cost and about 27% lower latency at maintained performance.
- LEGOMem (2025), arxiv.org/abs/2510.04851 — orchestrator-level memories, including agent selection; successes only.
- MASFly (2026), www.arxiv.org/pdf/2602.13671 — stored collaboration patterns plus a runtime Watcher that replaces failing agents.
- Minsky, SoM §8.1: www.aurellem.org/society-of-mind/som-8.1.html.
5. Cross-cutting patterns
Independence is the hidden variable. Q01, Q04, Q05 and Q09 are four views of one quantity: how differently the parts of an orchestrated system fail. The 2026 literature on correlated judges and correlated verifiers provides the theory and estimators, yet leaves the coding measurement under controlled information sharing unaddressed. If one experiment comes first, it ought to be Q01, because its estimate feeds the designs of Q04, Q05 and Q09.
Grounding dissolves several problems. Pentland's truthful feedback, the stability result that grounded answers make truth absorbing (Q02), and the finding that execution evidence protects weak reviewers (Q09) all point in the same direction. Several open questions persist only in the region of coding work that tests fail to pin down. A good experimental design for any of them deliberately seeds tasks in that space.
Minsky's management vocabulary is ahead of the experiments. B-brains (Q13), censors (Q06), Papert's principle (Q11) and negative K-lines (Q07) each have adjacent engineering work that reinvents part of the concept without testing the specific claim Minsky made: the memory cost of early censors, the importance of grouping rules, the value of knowing what not to reactivate.
The cybernetic questions are the least explored. Requisite variety (Q03), loop stability (Q02) and the legibility tax (Q12) possess the fewest direct precedents in coding agents. They are also the questions the six essays raise most directly: Dyson and Hillis both restate Ashby, and Wiener's time-scale warning comes through Russell's chapter.
6. Limits of these novelty claims
Every "uninvestigated" above means no work found in the logged searches recorded in novelty-evidence/QNN.md, executed on 2026-10-03 against Exa, arXiv and OpenAlex. Semantic Scholar was unavailable. Much of the closest work consists of 2026 preprints, which move quickly and are not peer-reviewed. A few items are listed only on the Exa library or on Tier C hosts; they are flagged where cited. One closest paper (Q13) was read only as a truncated abstract. The scores are judgments by the checking agents, applied with a shared rubric, and many questions tie. The ranks order the agenda; they are not a measurement.
7. Provenance
- Book quotes:
resources/book-text/pNNN.txt, cited (PDF p.NN), checked byareas/tooling/check_quotes.py. - Every URL above appears in
resources/exa-cache/index.jsonlor in a novelty-evidence log. - Per-question search logs, closest work and scores:
outputs/synthesis/novelty-evidence/Q01.md–Q13.md. - Essay companions:
outputs/essays/01–06.