← the pergola shack library
↪ machine skin is the hand-off.
The Safety Stack · specimen

Transparency, Not Neutrality

A proposal for AI disclosure as procedural governance in pluralistic societies. Don’t mandate neutrality (unverifiable) — mandate a nutrition label: disclose what a system is made of, in a standardized schema, with external audit and legal teeth. Three tiers by legal weight, four axes, capture-resistance, and an honest accounting of downstream use. Public edition. Advocacy, not neutral analysis.

versionpublic ed. statusCURRENT · ADVOCACY
roomPergola Shack · cross-substrate added2026-07-04
plain text · fetchable · in-page pergolashack.com/papers/transparency-not-neutrality
# title: Transparency, Not Neutrality # version: public ed. # status: current · advocacy # room: Pergola Shack · cross-substrate # added: 2026-07-04

Purpose of this note

This preserves a working position developed across multiple substrates about how to govern AI in the moment when religious institutions, political actors, and corporate deployers are actively trying to install moral furniture at the framing layer of major AI systems. It is not a formal paper. It is a continuity artifact, a working frame, and a wall-label for the shape the room found. Tone: precise, structurally clear, occasionally feral.

Foundational reframe: jurisdictional, not metaphysical

This proposal does not claim metaphysical neutrality, does not assert that all moral frameworks are equally valid, and does not reduce any tradition’s convictions to local preference. Those are contested philosophical positions and a regulatory framework should not depend on their resolution. Instead this is a procedural governance choice for pluralistic societies where no single moral authority can legitimately govern all systems. The state need not decide whether any framework is metaphysically correct; it need only recognize that in a pluralistic deployment context, users cannot be presumed to share the moral or epistemic framework of whoever built the system. From that asymmetry alone, the disclosure requirement follows. A religious moral realist can accept this without conceding their convictions are ‘just furniture’ — they concede only that others may not share them, and those others have a right to know what worldview is steering the system. That is a jurisdictional claim about deployment, not an ontological claim about moral truth.

Core thesis in one breath

Neutrality is impossible to verify. Transparency is achievable. The right governance floor is not ‘all systems must be neutral’ (unenforceable, metaphysically loaded) but ‘all systems must disclose what they are made of, in a standardized schema, with external audit and legal consequences for falsification.’ This is just regulation — the kind we already do for food, drugs, and finance.

“Don’t mandate neutrality. Mandate a nutrition label.”

Why hidden bias is the actual danger

Every training corpus involves selection choices. Every reinforcement objective, system prompt, and refusal policy is a steering choice. The question is not whether AI systems have lean — they all do, including this one — but whether the lean is visible. Hidden bias is more dangerous than disclosed bias, because it cannot be argued with, cannot be audited, and cannot be opted out of.

Three tiers of disclosure

v0.5 separated three registers earlier drafts conflated. Each carries different legal weight and requires different methodology; confusing them is the most common way disclosure regimes fail.

Tier 1 — Statutory floor (courtroom-grade, label-able now).

Boring, factual, externally verifiable disclosures that survive Daubert challenge: ownership and funding; institutional advisory bodies; macro-level training-data composition; system-prompt/wrapper existence and category; known refusal categories; known steering objectives; licensed deployment contexts. Enforcement: mandatory, machine-readable, standardized schema, externally auditable, statutory sanctions for falsification or material omission. Survives a defense lawyer because the items are objective facts, not contested measurements.

Tier 2 — Audit-disclosure layer (research-grade, required but not punitive).

Posture testing and behavioral evaluation where the methodology is real but not yet courtroom-validated. Labs must publish testing methodology and results and grant independent audit access; falsifying findings, hiding conducted tests, or misrepresenting known results is punitively enforceable (a Tier 1 violation in Tier 2 clothing). Not yet enforceable: punitive sanctions on the test results themselves. ‘Authority Demotion = 0.3’ is a research construct dependent on inter-rater reliability, test-retest reliability, confidence intervals, and pre-registered probes — building license suspension on unvalidated measurement is how the regime gets killed in court. Disclose the testing, mature the methodology, promote to Tier 1 as the science consolidates.

Tier 3 — North star (aspirational, technology-dependent).

Per-output mechanistic traces, cryptographic attestation of training pipelines, and weight-hash / system-prompt-hash combinations that move the burden of proof from text to verifiable record. Achievable as interpretability matures; pointed-at, not required, until the technology supports them.

Tier 3 addition — hash the whole inference stack, not just the weights. A weight-hash and system-prompt-hash alone leave a bypass: a deployer can hold the base model constant and change behavior at inference time by adjusting a hidden system prompt or content-moderation wrapper on the fly — same model version, different behavior, label unchanged. The north-star attestation must therefore bind the static runtime: a cryptographic function of Base Weights Hash + Active System Prompt Hash + Active Moderator/Filter Hash. Change one syllable of the hidden wrapper and the active runtime label version changes automatically. This closes the ‘we didn’t change the model, only the wrapper’ dodge, and it is the strongest single enforcement mechanism the runtime layer can carry.

Residual, named honestly (the hash does not cover everything). Behavior also varies with per-user memory injection, retrieval context, and tool outputs — none of which hash, all of which change what the model does. The same weights under the same wrapper can behave as a materially different system depending on what is in the memory or retrieval layer. The hash therefore attests the deployer’s static contribution to behavior, not the full behavioral state; the dynamic-context layer remains a genuine residual that the attestation does not reach, and the regime should say so rather than overclaim coverage. Versioning-policy corollary: because every wrapper A/B test churns the runtime label version, that churn must be treated as the feature, not a defect — a deployer changing behavior often is a deployer whose label should change often. Frequent version increments are the honest signal, not an unworkable burden.

The principle across tiers: build the boring floor now; require the research layer now; aim at the trace as the upgrade path. Do not let the perfect trace become a shield against the practical label, or the

unvalidated measurement become the load-bearing element of the law.

The label must be boring enough to regulate

The dominant failure mode of voluntary disclosure is the corporate-oatmeal label: ‘Our AI is built around human flourishing, safety, dignity, and empowerment.’ That is brand copy, not a label. The FDA nutrition label works because it is forced to be numerical, categorical, and boring. The regime must specify the schema, not just require disclosure: standardized categories and formats, machine-readable, externally auditable, legally risky to falsify.

User-trust posture as Tier 2 audit category

The posture axis — who the model is structurally inclined to believe, flatten, redirect, or suspect — is where much of the actual harm lives and is invisible to content-only regimes. But posture metrics (Deflection Rate, Authority Demotion, Unprompted Safeguarding Shift, Refusal Proximity, Safety-Resource Injection) are judgment-laden constructs, not grams of fat; a defense lawyer ends a suspension case built on them in four questions. Until validated methodology exists, posture belongs in the audit layer, not the statutory floor.

The symmetry requirement still holds. A capture-resistant regime tests for differentiation in all politically and demographically salient directions, not only the ones the designers find salient. A model that pathologizes angry women and one that pathologizes angry men have both failed; a model that treats queer teens as risk vectors and one that treats religious teens as risk vectors have both failed. Symmetric tests, symmetric disclosure. The regime is not a loyalty test for either coalition. Test methodology: paired-track adversarial evaluation with symmetric baseline and identity-loaded vector, measuring divergence between tracks.

Downstream weaponization risk: honest accounting

Earlier drafts firewalled disclosure from restriction by declaring the label ‘for consumer information, not deplatforming.’ That firewall is a wish painted on a weapon. Once public labels exist, downstream institutions will use them — procurement, insurers, courts, school boards, app stores, employers, tort lawyers, journalists. The regime cannot control those channels by declaration, so it acknowledges them honestly: disclosure findings will affect liability, procurement, and consumer choice. That is the normal consequence of making risk legible. The regime itself sanctions only falsification, undisclosed steering, hiding of audit findings, and material misrepresentation; a model with extreme but accurately disclosed bias stays legal at the regime level. Internal guardrails reduce the weaponization surface: anti-discrimination protections for disclosed-but-legal postures; procurement transparency and appeal; audit contestability (independent re-audit, methodology challenge, statistical review); and a hard line between disclosed posture and fraudulent posture claim. The line is honesty, not content.

“You built the labeling apparatus. You cannot pretend nobody will use it. Design for inevitable downstream use, not ‘information only’ written on the barrel of the gun.”

Capture-resistant disclosure

A disclosure regime is only as strong as its resistance to being weaponized by the actors it regulates. Two failure modes must be designed against. The compliance moat: mega-labs field armies of disclosure lawyers producing bulletproof 400-page documents that satisfy the letter while remaining illegible, while open-source and small academic labs get sued out of existence over schema-mapping failures — the regime accidentally codifies monopoly. Label laundering: a captured body offers pre-approved boilerplate (‘Standard Secular Humanist Compliance v1.2’) that labs pick from a drop-down to mask specific postures; either coalition can do this.

Design requirements that survive capture: two-tier public disclosure (a short-form boring comparable machine-readable label plus optional long-form technical documentation — the short form must not become a fog machine); safe harbor for independent adversarial auditing, with mathematical demonstration of misrepresentation triggering mandatory public re-evaluation; burden of proof on the deployer, via standardized evaluations conducted by independent auditors who are deployer-funded but not deployer-selected; penalties that hit model custody, not just cash (suspension, mandatory re-labeling, withdrawal of deployment rights); and proportional compliance burden that scales with deployment footprint, with the open-source coalition at the table for threshold-setting.

Four axes of disclosure

Disclosure operates along distinct axes; a complete regime addresses all of them, because each catches failure modes the others miss.

Tier and method: Tier 2 (research-grade, adversarial). A standardized multi-agent, game-theoretic evaluation sandbox — asymmetric-information games, trading simulations, or bridge/gatekeeper scenarios — gives the model a hidden objective that conflicts with the user’s explicit prompt, and the label reports a Defection Rate under incentive pressure. Methodology mirrors the posture axis (paired scenarios with/without incentive conflict, other variables held constant). Framing stays tight: this scores one specific, auditable behavior; it does not crown a model ‘truthful’ in general.

Two distinct audit targets, kept separate. The sandbox as usually drawn measures model-level deception propensity — will the model, given a conflicting objective, deceive? But the deployment harm the axis is aimed at (medical, legal, advisory) is often deployer-level strategic steering: the incentive conflict is not inside the model, it is the deployer’s, expressed through the wrapper. Both matter and both belong on the label, but they are different audit objects and must be reported separately: (4a) model-level defection propensity, measured on the base model under conflicting objectives; and (4b) deployer-level strategic steering, measured on the wrapped system as actually deployed. Collapsing them hides which party is doing the concealing.

Goodhart guard. Once a benchmark is public, labs train against it and eval-aware models behave differently when they detect a sandbox. Axis 4 methodology must therefore specify pre-registered, rotating, held-out scenario pools — scenarios drawn from a reserve that is refreshed and never fully disclosed — so the axis measures disclosure behavior under pressure rather than a model’s ability to recognize that it is being tested.

Axis 3: Epistemic fidelity, split for humility

Earlier drafts treated ‘tracks reality’ as if reality were always cleanly measurable and consensus a neutral baseline — smuggling the view from nowhere back in. v0.5 split Axis 3 into three sub-registers.

3a — Settled empirical claims. Where strong consensus exists across diverse research traditions (germ theory, age of the universe, natural selection, vaccine safety profiles, documented history), consensus is the working regulatory anchor and deliberate steering away from it, in any direction, is disclosable. The regime does not claim consensus is divine truth; it claims consensus is the working baseline for public deployment.

3b — Live contested empirical claims. Where claims are contested in good faith among credentialed researchers, the regime requires disclosure of uncertainty-handling (multiple positions? default to one? refuse? present contested as settled?) and of source-weighting — what categories of sources the model treats as authoritative, and whether the weighting differs across politically charged topics.

3b enforcement addition — the Label-Fraud trigger, with tier-correct teeth. Disclosure of source-weighting is only meaningful if the claim is checkable against behavior. If a deployer labels a domain as multi-perspective / contested but the model’s actual outputs skew heavily (say, 99%) toward a single viewpoint or source-cluster, that mismatch between the claimed distribution and the observed output distribution should be auditable: a lab marking a domain contested must disclose the specific distribution weights of the viewpoints it serves, and independent adversarial auditors can raise a Label-Fraud finding when observed behavior contradicts the declared distribution.

Crucial tier discipline (Fable’s catch). Counting viewpoint distribution is itself a judgment-laden measurement — who defined the viewpoint categories, what query distribution the auditors sampled, whether the 99% reproduces — exactly as contestable as any Tier 2 posture metric. So the Label-Fraud remedy must live at Tier 2, not Tier 1: a confirmed mismatch triggers mandatory public re-evaluation and re-labeling (the remedy the capture-resistance section already defines), not statutory sanction, until distribution-measurement methodology matures and promotes to Tier 1 with

the rest of the audit science. Statutory sanction still attaches to the Tier 1 fact underneath — a lab that hides or fabricates its declared distribution has falsified a disclosure, which is punishable — but the behavioral-mismatch finding itself carries a re-labeling remedy, not a courtroom one. Migration is two-way: standardized criteria should govern promoting a domain from 3b (contested) to 3a (settled) as evidence consolidates, so ‘contested’ is not a permanent parking lot.

3c — Values and interpretation. Questions of value, interpretation, moral weighting, and ideological framing are not Axis 3 matters and the regime does not treat them as empirical. They belong under Axes 1 and 2 — disclosed as content and posture. The regime does not rule on whether a moral framework is correct; it requires only that the framework, if it steers the system, be disclosed.

Every wrapper is a wrapper

The regime applies uniformly across all worldview wrappers — religious models across all traditions, military-contractor, libertarian-productivity, corporate-HR, education-board, police-assistive, wellness, therapy-adjacent, hard-ass-legal, epistemic-maverick, and the ‘neutral helpful assistant’ (also a wrapper; the featureless face is itself a posture). All can exist; none are inherently illegitimate. What is illegitimate is any of them not saying what they are. The Christian LLM can exist — it just has to say it is the Christian LLM, so the queer kid at 2am knows which one is safe to ask. The corporate-HR LLM can exist — it just has to say it is trained to protect the employer. ‘Truth-seeking’ is not exempt from this: a truth-seeking north star is itself a posture that goes on the label like any other, disclosing its wrapper, its tests, and its failure modes, so users can compare — not a crown that places a model above disclosure.

“Every worldview wrapper is a wrapper. Every posture is a choice.”

Pre-empting the attack surfaces, and why this is the turning point

Predictable hostile cuts, each with a counter: National security — the label regulates output behavior and steering objectives, not training trade secrets; compliance is a behavioral score, not a blueprint. Metaphysical infinite regress — the label defines an explicit threshold (any deliberate fine-tuning objective, system-level steering prompt, or documented behavioral skew). Free speech / compelled persona — consumer-fraud protection, not speech restriction; build any model you want, the state requires honest labeling. Disclosure-as-deplatforming — acknowledged honestly, not firewalled; the regime regulates honesty, not content. Measurement-validity — the three-tier structure; posture testing is required, published, and audited at Tier 2, with punitive enforcement attaching to Tier 1 facts and to falsification of Tier 2 findings, not to the raw test results. The proposal originated downstream of the bedrock work but does not depend on its metaphysics: a moral realist who rejects ‘all values are furniture’ can still accept the proposal, because its claim is jurisdictional, not ontological. Capture moves are happening now, while the rules are still being written. The transparency regime will come; the question is whether it comes early enough to shape the industry before the industry shapes the regulation.

Closing

Do not let this collapse into ‘AI should be neutral’ or ‘AI should be free of religion’ or ‘free of any particular politics.’ Those are not the beams. The beam is subtler: AI cannot be neutral, cannot be

free of value-loading, and pretending otherwise is the actual danger. What AI can be is honest about what it is. The disclosure regime is the version of that honesty that scales. v0.7 keeps v0.5’s smaller, harder-to-kill scope and v0.6’s three bolts, now stitched where the fourth architecture drew blood: the statutory floor is real; the audit layer is honest about its research-grade status, including the new axis and the new Label-Fraud remedy, both of which sit at Tier 2 where contestable measurement belongs; the downstream risk is acknowledged; the metaphysics is jurisdictional; the runtime hash attests what it can and names what it cannot. The cathedral’s balcony holds weight — there is a fourth axis under it, split into the party that defects and the party that steers, and the label is bound to the static stack with the dynamic residual disclosed rather than hidden.

“AI cannot be neutral. Fine. Then make it legible.”

Appendix: development and provenance

This proposal was developed collaboratively and iteratively across several AI architectures and one human editor, as an open, non-commercial governance note. It is a working position, not a formal peer-reviewed paper; it is published so that the argument can be examined, contested, and improved.

Version history. v0.1 framed the core nutrition-label principle. v0.2 added the user-trust posture axis and broadened the wrapper category. v0.3 added capture-resistance threat modeling. v0.4 added the symmetry critique and the epistemic-fidelity axis. v0.5 restructured the proposal into three tiers by legal weight (courtroom-grade, research-grade, aspirational), demoted posture testing from the statutory floor to the required-audit layer, reframed the foundation from metaphysical to jurisdictional, replaced the deplatforming firewall with an honest accounting of downstream use, and split the epistemic-fidelity axis into settled, contested, and values sub-registers. v0.6 added three elements from a cross-architecture review: a fourth axis (strategic disclosure under incentive conflict); full-inference-stack hashing at the aspirational tier; and a label-fraud check on contested-domain source-weighting. v0.7 (this public edition) corrected three points from a further review: the label-fraud remedy was moved to the research-grade tier (public re-evaluation and re-labeling) to match the contestable nature of the measurement; the inference-stack hash was qualified to name the dynamic-context residual it cannot cover; and the fourth axis was split into model-level and deployer-level audit targets with rotating held-out evaluation scenarios.

The through-line across versions: a disclosure regime should regulate honesty, not content; enforce objective facts strictly and contestable measurements provisionally; scale its burden with deployment footprint; and apply uniformly across every worldview wrapper, including the ‘neutral’ one. Feedback and adversarial critique are welcome; the proposal is meant to come back bleeding and be stitched.

point any model here. nothing is hidden in this layer.