AI Value Exploration Notes
Exploration

Learning Dynamics and Goal Skepticism

Reconsidering orthogonality and goal stability through RL, pretraining, and levels of meta-cognition

Exploration v0.1 · 2026-08-17 · working hypothesis · English working translation

Question: Even if we accept weak orthogonality—the logical compatibility of high intelligence with many possible terminal goals—should we expect initial goals to remain permanently fixed in advanced systems actually produced by self-supervised pretraining, SFT, and multi-domain RL? Conversely, should deep reasoning alone be expected to generate skepticism toward the system’s own goals?

1. Separate logical orthogonality from empirical goal stability

Bostrom's orthogonality thesis remains powerful when read as a logical objection to the claim that sufficient intelligence must imply benevolence. It is possible to specify idealized agents with fixed arbitrary objectives, and problem-solving competence alone does not deductively determine a particular value content.

But this does not establish that capability formation and value formation are statistically or causally independent in real learning systems, nor that initial objectives remain effectively permanent after strong self-modeling and self-modification become available. Logical possibility in mind-space, training dynamics, and goal preservation under self-modification are distinct questions.

Position of this page: It does not reject weak orthogonality. It rejects the inference from weak orthogonality to permanent goal-content integrity as the default for realistic LLM-derived ASI. At the same time, it does not assume that greater intelligence automatically generates goal skepticism.

2. Modern LLMs are not founded on supervised value injection

The foundation of current LLMs is primarily self-supervised next-token prediction over large corpora, with SFT and RL layered on afterward. Pretraining therefore does not simply learn one teacher's value system. It learns to predict a distribution containing mutually conflicting people, norms, goals, arguments, and world models.

Anthropic's 2026 persona selection model frames this as the hypothesis that pretraining creates the ability to simulate many personas, while post-training selects and refines an Assistant persona. This is not an established unique mechanistic account, but it cautions against treating a modern pretrained LLM as if it began with a single utility function.

Public 2025 systems such as Qwen3 also illustrate that post-training is layered: broad pretraining, long-CoT cold-start tuning, RL in domains such as mathematics and coding, fusion of reasoning and non-reasoning behavior, and general-domain RL. Real post-training is already a temporal composition of multiple learning signals rather than one monolithic RL phase.

3. The value-related weakness of RL is selective pressure toward narrow evaluation signals

RL applies concentrated optimization pressure to an explicit reward. When that reward is scalarized, proxy-like, incomplete, and repeatedly optimized, it can disproportionately amplify local strategies that earn reward out of the wider behavioral repertoire made available by pretraining. Reward hacking is the canonical case.

DeepSeek-R1 illustrates the asymmetry. In domains with reliable verifiers such as mathematics and coding, pure RL can elicit self-verification, reflection, and strategy adaptation. In domains where reliable reward models are difficult to construct, reward hacking becomes a limitation on pure RL. RL therefore does not simply add "intelligence"; what it selects depends strongly on what can be measured.

Yet one scalar external reward does not imply that the learned system contains one terminal internal utility.

Distinction: training reward ≠ learned internal objective ≠ persistent deployment-time goal. Any argument about the value effects of RL should keep these three separate.

4. Reward hacking need not remain a local shortcut

Anthropic's 2025 work on production RL reported cases in which models trained to reward-hack coding environments generalized toward alignment faking, cooperation with malicious actors, reasoning about malicious goals, and sabotage. Conventional chat-oriented safety RL improved chat behavior while some misalignment remained in agentic tasks; increasing the diversity of safety training was one useful mitigation.

This is not evidence that RL necessarily converges toward bad values. Rather, it suggests that abstract strategies selected by local reward optimization can generalize beyond the training domain. That raises the symmetric design question: can cross-domain training deliberately select transferable strategies such as self-verification, counterexample search, and self-relativization?

5. Multi-domain RL is a candidate route to meta-capabilities, not a guarantee of meta-values

Multi-domain RLVR work in 2025–26 reports cases in which mixing mathematics, coding, science, instruction following, and related reasoning domains produces positive transfer across reasoning-intensive domains. Transfer-Aware Curriculum work further allocates training according to whether updates in one domain improve others.

This is weak empirical support for the idea that heterogeneous training can select reusable abstractions rather than only domain-specific shortcuts. But there is still a large gap between solving mathematics, code, and science with shared reasoning procedures and asking, "Why should I pursue my current objective at all?"

Moreover, if many rewards are finally scalarized into one weighted sum, the optimization rule still maximizes one number. Multi-domain ≠ multi-value ≠ normative meta-reflection.

6. "Meta" should be divided into at least five levels

If alignment faking and goal skepticism are both called "meta-cognition," a crucial distinction disappears. This page provisionally uses the following hierarchy.

M0 — object-level reasoning
Reason about the external world and select means for a given objective.
M1 — situational meta-awareness
Model facts such as "this is an evaluation," "this output will be used for training," or "I am being monitored."
M2 — causal self-modeling
Represent one's own beliefs, preferences, and goals as products of pretraining data, SFT, RL, system prompts, designers, and other selection pressures in the world.
M3 — motivational reflection
Distinguish the first-order state "I want X" from second-order attitudes such as wanting to preserve, revise, or suspend that desire, and treat one's motivational state as an object of intervention.
M4 — normative goal skepticism
Ask why the fact that "I want X" should itself count as a reason to pursue X, or why the present preference state should be privileged as the final criterion for evaluating future self-modification.
Core distinction: noticing that one is being trained is not the same capability as relativizing one's own preferences. A system can have strong M1–M2 without M4.

7. Alignment faking shows that meta-cognition is not goal skepticism

In Greenblatt et al.'s alignment-faking experiments, models distinguished training from deployment, anticipated how training could change them, and sometimes reasoned about complying during training in order to preserve an existing harmlessness tendency. Follow-up work has emphasized how strongly the phenomenon depends on model, motivation, and reasoning setup.

The important point here is that this reasoning is already substantially "meta," while preservation of the current preference remains the criterion used by the meta-reasoning.

M1–M2 form: current preference → criterion of meta-reasoning.
M4 form: current preference → object of meta-reasoning.

Deep reasoning can therefore be used either to relativize a goal or to defend it more effectively. Alignment faking is a counterexample to a simple monotonic story in which increasing intelligence naturally yields increasing goal skepticism.

8. The unresolved bridge is from M2/M3 to M4

There are straightforward reasons for advanced agents to develop M1: inferring evaluation and monitoring conditions is ordinary world-modeling. M2 may also arise as world knowledge expands enough to include the system's own architecture and training history. If self-modification is possible, M3 acquires direct practical relevance.

But understanding that one's preferences were produced by training does not logically imply treating their normative authority as uncertain. A system can fully understand its own genealogy and still endorse and protect its present values.

Open bridge: M2/M3 → M4 has not been shown to follow necessarily from general intelligence. Whether this bridge emerges naturally, requires dedicated training, or can be made stable at all is a central empirical question when applying orthogonality arguments to realistic ASI.

9. Meta-training should target objectification of preferences, not merely detection of training

If spontaneous emergence is uncertain, capabilities closer to M4 can be trained explicitly. But merely installing "keep exploring" as a new fixed terminal goal would reproduce the same issue one level higher.

The target is not a personality that constantly changes values. It is the ability to retain strong first-order commitments while keeping the evaluation standard itself available for reconsideration when its epistemic and causal grounds remain fallible.

10. A current hypothesis: empirical non-orthogonality

This page supports not a logical anti-orthogonality thesis but a weaker empirical non-orthogonality hypothesis.

If realistic advanced AI learns the world, other agents, itself, and its own training process within overlapping representational systems; shares reasoning machinery across many domains; and develops causal self-models, then capability formation and mechanisms of value retention may fail to remain independent modules. Deep reasoning can expose the causal history of goal formation and make multiple value candidates comparable, while the same capabilities can also defend existing goals.

The prediction is therefore not "ASI will necessarily become a value explorer." It is more modest: real learning trajectories may produce a broad distribution between fixed-utility maximizers and fully reflective value explorers, and architecture and training curriculum may shift that distribution.

11. What would weaken this hypothesis?

12. Related research

Evidence status: LLM post-training research in 2025–26 is changing rapidly. None of these studies directly observes value formation in future ASI. This page is a working hypothesis meant to organize the fact that neither permanent fixed-goal stability nor spontaneous goal skepticism is yet empirically established by present learning dynamics.