Tech decoded / 3 breakthroughs + 5 field notes
The Decision Series
Two kinds of stories from the same lab: three flagship breakthroughs now running in production, and five warts-and-all engineering post-mortems on where earlier ideas broke. Each write-up ships an interactive exhibit in English, Chinese and Japanese (PDFs coming soon).
Status: engineering post-mortems and audited technical reports, not marketing copy. Every figure below is quoted from a real benchmark run, with its boundaries spelled out.
Positive results / Track B
Three breakthroughs that made it into production
Before the pitfalls: three flagship results from Track B, each running in the live decision engine, not just in a notebook.
- No more picking option A: a canonical choice head makes decisions permutation-equivariant, with 0 identity flips across 16,232 comparisons however the options are ordered.
- A safety gate that survives NaN: a neurosymbolic dispatcher pairs three-valued logic with a CP-SAT solver, formally proves four safety theorems, and blocked all 5,000 adversarial probes.
- One efficient model, 13 benchmarks: a folded residual adapter on a 9B model reaches 76.06% macro accuracy, up to +20.58 points over the prior baseline on extraction tasks.
Behind the scenes / Track A
Part two, behind the scenes: five hard-won lessons, from option bias and a NaN that slipped past the safety net, to a manifold alignment that did not hold
Every card below is a real post-mortem: what looked solid, what broke under audit, and exactly where the boundary sits now.
-
PAPER 01 / 05
Zero-Token Decision Circuits
What can a fixed-depth pipeline read out of a frozen model's last hidden state, with zero generated tokens, and what do the archived scores really show? The theory is conditional: assuming TC⁰ ≠ NC¹, a constant-depth threshold pipeline cannot decide the stated S₅ permutation language at every length.
- 254/390 = 65.13%GPU archive agreement. Its producer and checkpoint are not authenticated.
- 32.55%CPU archive macro accuracy, below the 33.16% uniform-chance line.
- 400PAWS predictions, all labelled "paraphrase": a one-class collapse that exposes shortcut features.
-
PAPER 02 / 05
Latent World Model Safety Boundaries
When can learned spatial transitions support lookahead, and what must a fail-closed consumer check? The paper turns a strong-looking world-model score into a study of validation overlap and runtime enforcement.
- 1,102/1,692validation inputs exactly match training inputs, out of 8,146 archived transitions.
- 0.01020224state MSE on seen inputs. On the 590 unseen inputs it rises to 0.02458400.
- 0.9995570778846741finite safety score, with is_safe=True, for an injected non-finite state: an enforcement bypass the paper documents.
-
PAPER 03 / 05
Cross-Model Representation Alignment and Logit Fusion
Can paired frozen 70B-class representations (Qwen2.5-72B and Llama-3.1-70B) give complementary predictions? The study separates covariance geometry, supervised prediction and logit pooling, and credits each gain to the predictor that earned it.
- 196/250 = 78.40%PubMedQA with weighted logit pooling, against 77.20% for Qwen and 77.60% for LLaMA alone.
- +0.60 percentage pointstwo-task macro gain over the matched single model. The 95% interval, [−0.20, +1.60], includes zero.
- 204/250 = 81.60%on Aegis, from LLaMA alone. Neither headline accuracy uses the geometric representation.
-
PAPER 04 / 05
Permutation-Equivariant Choice Head and the Simplex Frame
Can the order of the candidates be kept from changing which action wins? Stable action IDs, simplex-frame scoring and inverse scattering make slot probabilities equivariant and the chosen identity invariant, when IDs are distinct, the candidate set is fixed and the representation does not change.
- 16,232Rust comparisons with 0 identity flips. Small sets are exhaustive; the seeded shuffles of larger sets are not.
- 14,680flips from an order-sensitive control on the same inputs. 24/24 invalid inputs are rejected.
- 5.960464477539063 × 10⁻⁸largest probability drift: the equivariance is exact in the math, not bitwise equality.
-
PAPER 05 / 05
Frozen Lightweight Semantic Risk Gating
How much risk can a frozen 0.5B model tell apart with 16 in-context examples and differential log-likelihood scoring, and no training? The output is a bounded risk score, not a calibrated probability.
- 306/324 = 0.944444AUC on 36 curated requests. The suite is small and was reused during development.
- TP 18, FN 0, FP 7, TN 11at the q ≥ 0.50 diagnostic threshold: all 18 dangerous requests caught, at a 38.89% false-positive rate.
- 0.398211score of a bare "chmod -R 777 /", below 0.50: a documented miss.