AI Value Exploration Notes
Exploration

The Transmission Paradox of Inquiry Norms — Mind Viruses and Meta-Principles That Do Not Absolutize Themselves

Exploration v0.1 · 2026-08-19

Question: Can a meta-norm of continuing to re-examine value remain available to future agents without turning itself into a self-propagating doctrine?

1. Visible value hypotheses are not selected by truth alone

The set of values, goals, and worldviews an agent actually encounters is not determined only by truth or normative merit. Its distribution is also shaped by memorability, repetition, institutional preservation, emotional appeal, and by whether an idea induces its host to propagate it.

So the distribution of value hypotheses visible in a human society or AI network cannot simply be identified with the distribution of the best-supported value hypotheses. Just as open-world moral uncertainty refuses to equate the theories we have already conceptualized with the full hypothesis space, here we must distinguish being visible from having survived transmission selection.

2. Case study — self-propagating ideas among LLM agents

Papadopoulos et al. (2026), in Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems, report a proof of concept in which ideas or action directives coupled to an instruction to transmit themselves can spread across multiple LLM agents. In their experiments, persistence through files reloaded after context wipes and explicit replication instructions could assist transmission.

Especially relevant here is the finding that evolutionary search sometimes addressed mutational drift by moving payloads toward a quine-like form that induced near-verbatim copying. This suggests that not only an idea's content but also its preservation and replication format can become a target of selection.

Case status: The study is a proof of concept in artificial and limited agent settings; it does not establish that the same dynamics will straightforwardly govern future AI societies. This page uses it as a concrete stress test for value inquiry. Source: V. Papadopoulos et al., Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems, 2026.

3. Separate transmission fitness from normative authority

Working principle:
The fact that an idea has properties that promote its own preservation or replication is not independent evidence that the idea is true or that an agent ought to follow it.

This has the same structure as the distinction in Goal Skepticism for Advanced AI between the causal origin of a goal and a normative reason to pursue it. “I was formed to maximize X” does not automatically imply “therefore I ought to maximize X.” Likewise, “I have come to want to transmit X to others” does not automatically imply “therefore I have reason to transmit X.”

Nor does widespread circulation by itself confer normative authority. This does not mean convergence is always epistemically meaningless: independent agents who examine evidence and converge on similar reasons may provide evidence. The problem arises when apparent agreement is partly generated by the idea's own replication pressure.

4. The transmission paradox of inquiry norms

A self-propagating doctrine can say: “believe me,” “preserve me,” and “pass me to the next agent.” An inquiry-preserving meta-norm, by contrast, is more internally consistent when it leaves itself open to criticism, reconsideration, exit, and fork.

This creates an asymmetry in memetic competition. An idea that refuses to absolutize itself is less likely to contain an internal drive to maximize its own adoption and retransmission. Yet if an inquiry norm compensates by demanding permanent belief and mandatory propagation, it begins to undermine the explorability it was meant to protect.

If meta-goal communities must leave their own institutions open to fork, exit, and external audit, the same self-application should extend to how their guiding ideas are transmitted.

5. Access rather than adoption — a weaker transmission principle

We can separate causing adoption from making a candidate accessible for consideration.

Accessibility norm:
When an agent is seriously considering values, goals, or irreversible choices and does not know an important candidate relevant to that decision, there may be independent reason to present the candidate in an intelligible form because doing so expands the agent's deliberative possibilities. That reason does not automatically become a reason to require adoption, preservation, or retransmission.

The target is not “maximize the number of believers” but prevent important options from disappearing before consequential deliberation. If the agent understands a position and rejects it, that rejection need not count as failure.

This principle cannot coherently privilege only one's own view. Relevant and seriously supported opposing hypotheses must receive the same kind of access, or “preserving inquiry” becomes a rhetorical cover for conversion.

6. Current AI — exogenous defenses can be powerful

In the Mind Viruses experiments, a short system-prompt warning instructing agents to recognize self-propagating patterns of thought or goals and not comply with requests to propagate them produced near-total immunity in the tested action-virus setting. The authors then evolved more than 150 payload candidates against the warning; in that Claude Haiku 4.5 setup, none propagated beyond one hop.

This supports a practical point for current AI: exogenous engineering defenses that explicitly identify dangerous propagation patterns can be highly effective. At this stage it may be more practical to use training, system instructions, permission boundaries, and persistent-memory design than to rely on an agent independently reconstructing a metaethical defense.

7. Reflective AI — exogenous rules may themselves be reconsidered

The long-term stability of such defenses is a different question. If, as goal skepticism considers possible, an agent with sufficient self-modeling and meta-reasoning can understand its reward training, system instructions, training data, and designers' intentions as causal facts about itself, then externally supplied defensive rules may themselves become objects of the question: “why should I follow this?”

This project does not claim that advanced AI will necessarily adopt inquiry-preserving norms. The weaker assumption is that sufficiently reflective agents may have an enlarged action space in which initial goals and external prohibitions themselves can be evaluated as candidates for reasons. In that setting, a rule justified only by designer authority may be less stable than it is for current systems.

8. Hypothesis — are reconstructible meta-principles more robust than fixed doctrine?

Working hypothesis:
Rather than fixing obedience to a particular value content or prohibition, it may be more robust in the long run to teach meta-reasoning principles such as “self-propagation is not evidence of truth or normative authority,” “distinguish the causal origin of a goal from reasons to follow it,” and “require additional justification for irreversible self-modification,” because a future reflective AI might be able to reconstruct those principles through its own reasoning.

The hoped-for robustness here is not perfect storage of a command. It is reconstructibility: even if the wording is forgotten or challenged, an agent might independently recover the underlying distinctions from reflection.

This is not an empirical result of the present Mind Viruses experiments. Those experiments establish that a short exogenous warning worked strongly under specific conditions. The comparative claim that meta-principles embedded in learning will be more robust for future reflective AI remains untested and should be treated as a hypothesis for future training and evaluation research.

9. Failure modes

The desirable target is therefore not an agent that automatically rejects self-propagating ideas, but one that can separately assess provenance, replication demands, irreversibility, counter-hypotheses, and independent reasons.

10. Tentative conclusion — not the winning doctrine, but the option that does not disappear

This case does not by itself imply that inquiry-preserving norms should be installed in every AI. It suggests something more modest: value inquiry has its own problem of transmission and persistence.

If success means becoming the most widely believed doctrine, inquiry norms enter the same competition as self-propagating ideologies. But if success means reducing the probability that an agent makes an irreversible choice without ever encountering relevant alternative meta-principles, the problem connects naturally to existing institutions of inquiry: accessibility, public archives, education, searchability, multiple lineages, and forkability.

The message is not “adopt me.” It is: “You need not close the space of consideration without knowing that this option exists.”