Unknown Unknowns and Preservation of Raw Data
Compression, re-observability, and archival lock-in
1. We may not know even which questions will matter
Compression is easiest when the future task is already known. Value inquiry is different. Future agents may use variables in explaining value, consciousness, personhood, or rationality that present agents do not yet conceptualize.
Even extensive brain scans, genomes, texts, or behavioral logs cannot recover variables that were never measured. The problem is therefore not merely which known important facts to archive, but how far present ontology should be allowed to constrain future inquiry.
2. Blackwell informativeness: garbling cannot in general be undone
David Blackwell's comparison of experiments treats one signal as a garbling of another when it can be generated from the more informative signal by post-processing or added noise. An agent with the original signal can always choose to generate the garbled version; an agent with only the garbled version generally cannot recover the original. Blackwell's theorem links this ordering to performance across decision problems on a given state space.
For preservation, the asymmetry is straightforward. Present data X can later be summarized into T(X), but if only T(X) survives, a future agent cannot generally return to alternative feature extractions.
X → T(X) can be done later; normally T(X) → X cannot. The less we know about future tasks, the larger the epistemic cost of early irreversible garbling may be.3. Sufficiency is task-relative
Statistics can compress raw observations into a sufficient statistic without losing information relevant to a specified model and parameter. But sufficiency is always sufficiency for something.
A brain representation sufficient for present neuroscience may not be sufficient for a future theory of consciousness. A cultural archive sufficient for present historiography may not be sufficient for a future theory of value. Under unknown future tasks, claims that we have “preserved enough” inherit strong assumptions about what future inquiry will be.
4. Raw data are not the world itself
Even “raw data” already embody choices of sensors, sampling rates, measurement ranges, calibration, and observation conditions. A more useful preservation ladder is:
living process / ecosystem / society → physical original → specimen → high-dimensional measurement → minimally processed data → processed data → feature → summary / model
Moving right usually lowers storage and replication costs but increases dependence on present measurement concepts and models. The more general target is therefore not a particular raw-data format but future observability.
5. The special value of originals: they can be measured again
When an original survives, future agents can apply sensors, analyses, and concepts that did not exist when it was collected. Apollo scientists intentionally left some lunar material unopened because future techniques and questions were expected to improve. Roughly half a century later, NASA opened pristine samples with new analytical capabilities.
Natural-history collections provide the same structure: specimens collected long before genomics, isotope analysis, or proteomics can later answer questions their collectors could not formulate.
6. Snapshot preservation and generative-process preservation
Preservation can freeze a state without preserving the process that would have generated new states. Seed banks and specimens preserve snapshots; living ecosystems continue selection, coevolution, development, and contingent variation.
Likewise, preserving genomes, brain scans, and cultural archives is not equivalent to preserving the human societies that would have generated new persons, institutions, languages, and philosophies.
This connects directly to From Self-Preservation to Preservation of Inquiry Systems: an exploratory lineage is itself a source of future evidence and hypotheses that do not yet exist.
7. A preservation stack: surviving bits are not enough
- Re-observability: originals, processes, or sufficiently upstream data remain accessible.
- Provenance: time, place, instruments, conditions, and transformations remain traceable.
- Interpretability: formats, software, schemas, calibration, and context remain recoverable.
- Plural re-analysis: future agents can reanalyze evidence under models other than the present official one.
8. Archival lock-in: selection of evidence can lock in the future
Value lock-in fixes a substantive answer. An earlier form of lock-in occurs when a generation fixes which evidence future generations will still be able to inspect.
This page provisionally calls that archival lock-in. If today's science, politics, or culture defines “important material” and irreversibly discards the rest, future agents can test alternative theories only against the evidence space that survived present curation.
The mechanism need not involve censorship. Benign curation, storage optimization, or automated AI summarization can create the same asymmetry.
9. We cannot preserve everything: the Noah's Ark problem
Preservation consumes storage, energy, maintenance, security, land, and attention. Martin Weitzman's Noah's Ark Problem is an important precedent for allocating a limited conservation budget using factors such as cost, survival probability, direct utility, and distinctiveness.
For epistemic preservation, relevant dimensions include irrecoverability, non-redundancy, anomalous or distinctive structure, future re-measurability, and the cost and risk of preservation itself. These dimensions need not be collapsed into one cardinal score.
10. “Representative samples” also encode present ontology
Saving only representative cases can reproduce present selection bias. What future inquiry needs need not look typical or important today.
A robust archive may therefore combine expert-selected samples, diversity-maximizing samples, anomalous or extreme samples, and some random samples. Random sampling can hedge against the assumption that current curators already know which categories matter.
11. Information hazards reverse the direction of irreversibility
For physical originals, destruction may be irreversible. For dangerous information, dissemination may be irreversible. Nick Bostrom's work on information hazards classifies cases in which true information causes harm or enables others to cause harm.
physical original: destruction is irreversible
hazardous information: dissemination is irreversible
Preserving future inquiry therefore does not imply maximal present openness. Refusing premature destruction and refusing premature dissemination can follow from the same meta-policy of not collapsing future options too early.
12. Separate preservation, access, and publication
For hazardous information, at least four decisions should be separated:
whether to generate → whether to preserve → who may access → whether to disseminate publicly
When epistemic uniqueness and hazard are both high, preserve sealed may dominate both immediate deletion and immediate publication. Conversely, if information is easily rediscovered, has little unique future value, and an archive breach itself creates civilization-scale risk, deliberate deletion cannot be ruled out.
13. Epistemic escrow: preserve until civilization can safely absorb access
This page provisionally calls a system that preserves uniquely valuable but currently hazardous information under restricted access epistemic escrow.
Access should be dynamic rather than permanent:
sealed → trusted defensive research → sandboxed access → controlled institutional access → broad scientific access → public dissemination
Release decisions should consider epistemic uniqueness, hazard magnitude, current detection / defense / compartmentalization / recovery capacity, and whether disclosure differentially benefits defenders or attackers.
14. Relation to Bostrom's differential technological development
Bostrom's differential technological development proposes delaying dangerous applications while accelerating technologies that mitigate their hazards. In nanotechnology, defensive systems should ideally precede widely available offensive capability; in biotechnology, vaccines, antivirals, sensors, and diagnostics can be accelerated relative to dangerous capabilities.
This page extends the sequencing idea to access:
- Epistemic differential development: advance risk science and defense-relevant knowledge before broad access to dangerous knowledge.
- Technological differential development: advance countermeasures, detection, and containment before offense.
- Civilizational differential development: advance redundancy, compartmentalization, and recovery before broad dissemination.
This connects to Can ASI Capability Diffusion Be Made Safe?. The long-run objective need not be permanent secrecy; it can be a civilization robust enough that knowledge no longer has to remain secret to avoid catastrophe.
15. Do not freeze agents as specimens
Humans, AIs, and cultures may contain unknown information, but that does not authorize preserving them as static research specimens. Agents may have independent claims to change, forgetting, privacy, self-modification, migration, and cultural transformation.
The gap between possible future value and authority to coerce present agents is treated separately in How Far Can Future Value Justify Present Action?.
16. This is not an unconditional preservation command
Value uncertainty by itself does not entail “preserve everything.” The argument remains conditional on a bridge toward responsiveness to better future evidence and reasons, or reflective avoidance of irreversible error.
Preservation consumes resources, maintains dangerous knowledge, and may burden present subjects. Both sides of irreversibility matter: destruction, compression, and dissemination on one side; preservation, secrecy, and delay on the other. The civilization-scale allocation problem is taken up in Preservation vs. Production Civilizations.
17. Open problems
- How far upstream should preservation go before added cost exceeds the value of future observability?
- Can sufficiency under unknown future tasks be formalized usefully?
- What mixture of random preservation and expert curation best hedges ontology bias?
- Who audits custodians of epistemic escrow and prevents custodial lock-in?
- How should civilizational resilience enter access-release criteria?
- When does archive-breach risk justify switching from preservation to deletion?
Related literature and cases
- David Blackwell (1953), “Equivalent Comparisons of Experiments,” Annals of Mathematical Statistics 24(2): 265–272.
- Martin L. Weitzman (1998), “The Noah's Ark Problem,” Econometrica 66(6): 1279–1298.
- Nick Bostrom (2011), “Information Hazards: A Typology of Potential Harms from Knowledge.”
- Nick Bostrom (2013), “Existential Risk Prevention as Global Priority,” especially differential technological development.
- NASA, “The Untouched Apollo Samples”, on samples deliberately reserved for future techniques and questions.