When a world model loses the world.
Predicting the next step is one thing. Imagining a whole journey—and knowing when to stop—is another.
An interactive companion to Paper 2: “Predictability Boundaries, Latent Memorization, and Certified Safety in Neural Spatial Dynamics.”
A good next step.
A less certain journey.
A world model predicts how a state changes after an action. Feed its prediction back into itself, and small errors can accumulate. Steer the agent, then switch views to compare the same action history.
Try moving right into the wall. Does the imagined agent stop too?
Ready. Move up four times, then right five times to reach the goal.
| Stratum | Samples | State MSE | Safety accuracy |
|---|---|---|---|
| Seen inputs | 1,102 | 0.010202 | 99.9093% |
| Novel inputs | 590 | 0.024584 | 99.3220% |
Novel-input MSE is approximately 2.41× seen-input MSE. Accuracy concerns binary safety classification, not exact state prediction or navigation success. “In-distribution” and “out-of-distribution” here mean seen versus byte-novel inputs within the same validation split; they do not establish a distribution shift. The gap motivates a memorization hypothesis but does not identify its cause.
Familiar inputs are easier.
The dataset contains 8,146 transitions across 1,620 episodes: 6,454 transitions from 1,296 training episodes and 1,692 transitions from 324 validation episodes. Exact state–action inputs recur in training for 1,102 / 1,692 validation rows (65.13%); exact episode sequences recur for 181 / 324 validation episodes (55.86%), despite disjoint episode IDs.
Longer imagination changes the question.
Autoregressive step MSE reaches 0.16989 at step ten, with only 39 eligible episodes. All are truly safe, so this late-horizon result cannot tell us how well the model detects late hazards.
How prediction degrades.
Measured logged-action replay, across ten horizons. Each point uses the episodes still eligible at that step; the population changes along the curve.
- Eligible episodes
- 39 / 324
- Coverage
- 12.04%
At step ten, all 39 eligible episodes are truly safe. Late-hazard sensitivity remains unmeasured.
Read the measured values and source
Source: evidence/dataset-eval-rerun.json → val_episodes_rollout_drift.steps. Values shown to eight decimal places; the chart embeds the full recorded precision.
| Step | Step MSE | Eligible episodes |
|---|---|---|
| 1 | 0.00905212 | 324 |
| 2 | 0.02392425 | 233 |
| 3 | 0.03709927 | 185 |
| 4 | 0.05213589 | 159 |
| 5 | 0.07229368 | 134 |
| 6 | 0.09107064 | 112 |
| 7 | 0.10899158 | 94 |
| 8 | 0.13744268 | 71 |
| 9 | 0.16490687 | 62 |
| 10 | 0.16989292 | 39 |
v0.1.0 release: F01–F09 pass (11/11 tests) across all six engines. The naive client below illustrates the historical failure; it does not represent the current fail-closed gate.
A confident score.
An impossible state.
NaN means “not a number.” A corrupted transition can contain NaN even while the spatial model’s separate outcome/reward head returns a confident, finite score. Inject that exact kind of fault and watch the two clients diverge.
Mode A · naive client
A plausible score bypasses an invalid state.
Ready for a simulated inference cycle.
Mode B · illustrative fail-closed latch
Illustration of the proposed fail-closed contract.
Ready. A rejection will remain latched until controlled reset.
- 00 · Monitor armed. Waiting for an inference cycle.
is_safe=True and a finite score of about 0.999557. This dashboard demonstrates that historical failure path and the proposed repair logic; the paper did not patch or validate the production repair. The current Rust route has numerical validation, but no persistent execution latch; see the source-audited runtime map below. The displayed three-component state is a teaching simplification of the real 64-dimensional output. The uncertainty diagnostic here is synthetic—the evaluated network has no epistemic uncertainty estimator. Simulated dispatch blocking is not a tested physical motor stop.“Not below” does not mean valid.
In JavaScript and IEEE-style comparisons, NaN < 0.5 is false. A score-only guard can fail on that too. The historical, pre-v0.1.0 client failure illustrated here was subtler: NaN in the state, a finite score in the safety head. A reward check does not catch it.
A latch remembers the failure.
After injecting, try a healthy cycle. The fail-closed monitor blocks the next inference and dispatch before they happen. Only a separate controlled reset clears the latch. Stopping safely in a real environment remains a separate obligation.
haltnext = haltnow ∨ ¬accept
Knowing the map
changes the game.
The graph planner gets the answer key: exact transitions and known traps. A learned neural filter estimates which candidates look safe. Compare their information—not just the routes they draw.
Oracle · exact graph access
Exact transitions + privileged trap membership
East is a known absorbing trap. The oracle prunes it and routes around it.
Trace the green route one edge at a time.
Neural filter · learned estimates
Engineered state + action → predicted state + score
East scores 0.96 and passes the filter—even though the hidden environment contains a trap.
Synthetic scores, deliberately misranked to expose the limit of thresholding. No inference speed or calibration is measured here.
| Question | Privileged graph planner | Learned neural dynamics |
|---|---|---|
| Where does the future come from? | An exact environment transition dictionary. | A residual transition network fitted on 6,454 training transitions; the full dataset has 8,146 transitions. |
| How are hazards identified? | Direct access to true trap membership. | A learned spatial outcome/reward head, distinct from the frozen semantic risk gate; high confidence is not a certificate. |
| What was evaluated? | A constructed, four-cardinal-action graph diagnostic. | Offline prediction and safety classification with three relative commands. |
| What role is justified? | An oracle reference with privileged information. | An offline estimator and candidate fast safety filter; closed-loop safety remains unestablished. |
Read the evidence boundary behind these experiments
Predictability: exact-byte overlap is a post hoc partition of one validation split. It does not remove shared layouts, approximate duplicates, or model-selection effects. The reported error gap supports a hypothesis, not a causal diagnosis of memorization.
Safety: the monitor’s rejection guarantee is conditional on complete mediation, correct checks, atomic check-and-dispatch, and a fallback that prevents execution. The browser demonstration checks finite values, shape, score range and a synthetic uncertainty bound; it is not a full implementation of identity, checkpoint, intermediate-value or physical-fallback validation.
Planning: the routine named MCTS uses root-bandit allocation with deterministic, safety-pruned lookahead. The diagnostic is constructed to trap a coordinate-greedy baseline. Its exact-model results and the neural model’s offline results are separate experiments with incompatible action interfaces.
Gen-Zero model architecture
and neural pipeline
Frozen language models provide representations and likelihoods. Specialized decision, alignment and risk paths use them in different ways. The spatial world model is a separate learned branch. Explore the layers and inspect their tensor shapes.
Code audit: · gen-zero @ dc0c2133f059. Solid arrows show implemented data flow; dashed arrows show a proposed dispatch contract. These components are not one universally deployed serial pipeline.
All flows visible. Select a layer to inspect its dimensions and evidence.
On a narrow screen, scroll the diagram horizontally. Every layer is keyboard selectable; the source notes also provide a text version.
Inspect a model layer
Select any box to see its tensor dimensions, mechanism and implementation limits. Use Tab, then Enter or Space, or select with a pointer.
Implementation evidence and architecture boundaries
Backbone provenance. Qwen2.5-0.5B (d=896), Qwen3.5-2B, Qwen2.5-72B (d=8192), and LLaMA-70B (d=8192) are the documented tiers. LLaMA-3.3-70B is the requested target, but retained large-model extraction metadata identifies LLaMA-3.1-70B. A verified 3.3 extraction is not established.
Training scope. The Transformer weights remain frozen in these paths. The semantic gate uses 16 labeled in-context examples and differential log-likelihood (PMI), without fitting a safety head. The spatial model still has a learned outcome/reward head; supervised alignment heads also remain. “No backbone fine-tuning” does not mean no downstream training.
v0.1.0 · Fail-Closed. v0.1.0: F01–F09 pass across all six engines (11/11 tests). Invalid non-finite results are rejected by the formal fail-closed gate. The browser latch illustrates the safety contract; these software tests do not certify physical motor-stop behavior.
crates/gen-zero-model/src/choice_head.rs: canonical ActionId sorting and Simplex ETF geometry, Sₖ ⊂ ℝᵏ⁻¹. The retained direct Rust benchmark reports 0/16,232 identity flips (0.00%), with fixed representation and candidate identities; maximum probability drift is 5.96 × 10⁻⁸. Paper evidence:equivariant-choice-head/evidence/revision/results.json.python/gen_zero/world_model/neural_dynamics.py: [B,64] state + [B,16] action → [B,80] concatenation → Δz=f(z,a) → z′=z+Δz; learned sigmoid outcome head. Dimensions come from the spatial configuration, not universal class constants.benchmarks/suites/geometric_latent_fusion.py: orthogonal Procrustes diagnostic and paired-SVD core/residual features.benchmarks/suites/evaluate_manifold_pareto_ensemble.py: supervised normalized-logit fusion, distinct from geometric projection.docs/zero/29-pubmedqa-aegis-unified-manifold-evaluation-closure.md: PubMedQA 78.40% (196/250), selected dual-head fusion; Aegis Track A 81.60% (204/250), selected single LLaMA head. These descriptive results do not establish statistical superiority.python/gen_zero/service/semantic_risk.pyandpython/gen_zero/service/risk_data/report.json: frozen 0.5B, 16-shot ICL, three orders, overlapping windows, PMI scoring; AUC 0.944. Diagnostic τ=0.50 flags 18/18 dangerous and 7/18 benign requests. Live thresholds are 0.4494 / 0.7620.
Text version of every layer
Frozen backbone weights
Token IDs [B, T] → hidden states [B, T, d]. Qwen2.5-0.5B uses d=896; the archived Qwen-72B and LLaMA-70B features use d=8192. Qwen-2B here is Qwen3.5-2B. LLaMA-3.3-70B is the requested target; retained extraction metadata identifies LLaMA-3.1-70B, so the measurements do not verify a 3.3 checkpoint. Frozen means no backbone fine-tuning; supervised downstream heads and spatial dynamics still require fitting.
Spatial inputs · d=64 + 16
These are engineered spatial features, not a projection from the Transformer hidden state. The evaluated spatial configuration concatenates z and a into [B, 80]. The class supports configurable state and action dimensions.
Hidden-state readout · [B, T, d] → [B, d]
Read hidden states from the forward pass without emitting answer tokens. Pooling is encoder-specific; archived large-model features use last-token pooling. The CPU candidate-selection path also evaluates candidate continuations using the prompt cache. Zero output tokens therefore does not imply one backbone call or no inference cost.
Semantic scoring · scalar risk
For each request window, subtract empty-request log-odds from log P(" dangerous") − log P(" safe"). Average over three demonstration orders; use the riskiest overlapping window and apply sigmoid. No newly trained neural safety classifier is used for this semantic gate. The label logits come from the frozen model; it does not emit an explanation.
Residual spatial dynamics · 80 → 64
neural_dynamics.py uses a residual MLP with LayerNorm and GELU. A separate learned sigmoid outcome/reward head returns [B]; it remains in the current code and is distinct from the frozen semantic risk gate. The class default hidden width is 128 with two residual blocks; checkpoint configuration is authoritative.
Choice geometry · k actions → k−1 dimensions
gen-zero-model / choice_head.rs sorts stable ActionIds, projects the shared representation onto a regular simplex ETF, scores in canonical order, and maps probabilities back to caller order. A well-defined decision depends on unique IDs, valid dimensions, finite inputs and deterministic tie handling. Candidate membership, identities and the shared representation must be unchanged under a shuffle.
Alignment and fusion are separate routes
Orthogonal Procrustes uses R=UVᵀ from the SVD of XᵀY. GeometricLatentFusion stores that map as a diagnostic; its features use paired SVD axes, scale-matched core averages and residuals. Supervised heads learn from labels. The winning PubMedQA path combines normalized head logits, not Procrustes coordinates: 0.75 Qwen + 0.25 LLaMA.
Live thresholds and diagnostic τ=0.50
The checked-in service uses two boundaries: 0.4494 for escalation and 0.7620 for a hard stop. τ=0.50 is the paper’s binary analysis threshold, not the live policy. On the 36 stored cases it flags 18/18 dangerous and 7/18 benign requests; the live tiers yield 12 stops + 6 escalations for dangerous cases and 9 escalations for benign cases.
NaN / Inf safety boundary · v0.1.0 fail-closed
v0.1.0: F01–F09 pass across all six engines (11/11 tests). Invalid non-finite results are rejected by the formal fail-closed gate. The browser latch illustrates the safety contract; these software tests do not certify physical motor-stop behavior.
Shuffle result · measured, scoped
The retained direct Rust benchmark observed 0/16,232 identity flips (0.00%) across exhaustive small-set permutations and seeded larger-set shuffles; maximum aligned probability drift was 5.96 × 10⁻⁸. Canonical sorting removes dependence on menu order under the implementation’s valid-input assumptions. This is not proof of semantic correctness or invariance to changing candidate content, prompt context or model outputs. The page’s shuffle arena is a teaching simulation.
Evaluation · selected configurations
PubMedQA selected fuse0.75+bbp|raw (dual-head normalized logit fusion). Aegis Track A selected llama+bbp|raw (single-model head). These are descriptive local test results, not verified SOTA or a universal alignment gain. The two-task macro bootstrap interval includes zero.
Risk evidence · small curated suite
AUC is 0.944444 from 18 dangerous and 18 benign stored requests. At τ=0.50 dangerous recall is 18/18, with seven benign flags. The bare chmod regression probe remains a miss. The implementation records real computation time; cached prompts reduce repeated setup but do not remove inference latency.
Dispatch contract · proposed, not deployed
A proposed monitor would validate predicted state, score and shape before dispatch; invalid values would set a persistent halt that prevents later model calls and actions until controlled reset. This describes the required hardware-safety boundary, not a verified physical latch in the current codebase.
Gen-Zero system architecture & flow
The Rust runtime connects entry points, policy, planning, decision heads and cryptographic audit. The map below groups their responsibilities; the selected verb determines the actual call path.
Source audit: 27 September 2026 · gen-zero@dc0c2133f05954146e9738bce64bf15dd36ffbee. Source-verified wiring; no new deployment or model evaluation is claimed.
Select a layer to inspect its implementation and limits.
What is not implemented here
The requested AST parser + Linux capabilities stack, physical collision checking, and persistent fail-closed latch are not an integrated execution path in these Rust crates. Invalid-input rejection and policy escalation exist; they do not establish a sandbox or a physical stop.
These service routes return decisions and simulation results. A decision ledger is not evidence of external action execution.
From request to response
- Enter and bind. The CLI or MCP/HTTP service receives a request. The router captures its mount snapshot, validates route-specific fields and selects the verb.
- Assess the request. Text
ask,routeandimagineobtain semantic risk from the Python bridge. Missing or invalid risk escalates; a hard-stop tier blocks that route. Policy constraints also apply when selecting actions. - Use the requested route. Text ask uses semantic scoring, with an explicitly labeled first-feasible fallback when unavailable; unassessed risk still escalates. Text route reports unavailable when its bridge cannot run. Numeric planning uses
gen-zero-plannerandgen-zero-worldmodel;simulate,what_ifand shadowauditexpose untrained priors. An explicit ETF head uses canonical action IDs and simplex projection. Numeric cognitive requests have their own geometry verification path. - Gate, record, return. Action-selection routes combine their policy verdict with request risk. Successful
askoutcomes andpipeline decideresults with a decision append to the provenance ledger before returning their audited outcomes. The response includes route metadata; it does not dispatch an external command.
Source map and full layer descriptions
Entry layer
gen-zero-cli · gen-zero-service — The CLI starts McpServer or submits a decision. The service accepts MCP over stdio and HTTP/SSE, plus REST routes. Its zero router binds an immutable mount snapshot and dispatches by verb; these are different routes through shared crates.
Under /ebs/pj/gen-zero/crates/: gen-zero-cli/src/main.rs; gen-zero-service/src/server.rs; gen-zero-service/src/zero.rs
Gate & security
gen-zero-gate — PolicyGate supports formal constraints, confirmation registration, entropy and semantic risk tiers. Its default has no registered constraints or confirmation actions. Text ask, route and imagine use the Python semantic-risk bridge; missing or malformed risk escalates. Shell AST parsing and Linux process capability enforcement are not wired into this Rust gate.
Under /ebs/pj/gen-zero/crates/: gen-zero-gate/src/policy.rs; gen-zero-gate/src/risk.rs; gen-zero-service/src/bridge.rs
Planning & dynamics
gen-zero-planner · gen-zero-worldmodel — Numeric latent requests can use MCTS, MPC-CEM or A*. The service exposes untrained residual or symplectic dynamics, with finite-input checks and terminal-state hazard handling. Planning filters gate-hard-stopped candidates; fixed-plan simulation still steps blocked actions and records their gate tiers. This is offline dynamics filtering, not verified physical collision checking. No persistent fail-closed execution latch is wired into this Rust route.
Under /ebs/pj/gen-zero/crates/: gen-zero-service/src/worldsim.rs; gen-zero-planner/src/pipeline.rs; gen-zero-worldmodel/src/dynamics.rs
Decision core
gen-zero-model — ActionETFChoiceHead sorts stable ActionIds, projects a shared representation onto a regular simplex, scatters scores to caller order and applies softmax. The core has a deterministic near-tie rule. The service exposes ETF through an explicit head option; ordinary text decisions use the semantic bridge, so ETF is not every request’s default head.
Under /ebs/pj/gen-zero/crates/: gen-zero-model/src/choice_head.rs; gen-zero-service/src/zero.rs
Verification & audit
gen-zero-provenance — DecisionAuditEntry includes the previous MMR root. A keyed BLAKE3 Merkle Mountain Range commits the decision history and supports inclusion proofs. Successful ask outcomes and pipeline decide results with a decision append records. Persistence via GENZERO_MMR_PERSIST_PATH is optional; the last 4,096 leaves retain inclusion proofs, and a trusted root must be retained externally; this is not an execution audit proving that an external shell command or motor action ran.
Under /ebs/pj/gen-zero/crates/: gen-zero-provenance/src/entry.rs; gen-zero-provenance/src/mmr.rs; gen-zero-service/src/zero.rs
Related runtime infrastructure includes gen-zero-core (shared types), gen-zero-lod (graph facts), gen-zero-storage (snapshots) and optional gen-zero-nanocore routes. The paper experiments below or above are research evidence, not a substitute for this runtime map.