Recursive Self Improvement Lab
Propose. Verify. Inherit only when held-out evidence says so.
An attempt to build a recursive self-improvement lab from sealed local projects — not from self-rating. Club A carries the evidence notebooks; Club B carries verifiers, agent runtimes, and archives. Negative results stay published beside positive ones.
Program
-
Recursive Self Improvement Lab
A verifier-gated research program assembled from sealed local experiments: propose, intervene, verify, archive, and inherit only when held-out evidence improves.
How the lab connects
Improvement is earned only on the last arrow: archive-conditioned reuse must beat memoryless baselines on untouched families under a truth standard the proposer cannot rewrite. That inheritance gate is queued, not claimed.
flowchart TB
propose["Propose / vary
AlphaEvolve · Parameter Golf · Sansara"]
intervene["Intervene / run
WCCS · PG budgets · Gigatoken"]
verify["Verify
Oracles · Lean gates · sealed hashes"]
measure["Measure / falsify
Nostradamous · Agora"]
archive["Archive
Vault Ops · Publications"]
propose --> intervene --> verify --> measure --> archive
archive -->|"held-out inheritance only"| propose Club A — Evidence engines
Researcher-usable reports and sealed notebooks: Question, Context, Findings, Method, Limits, Artifact.
-
Gigatoken Lab
Independent reproduction of Marcel Rød's Gigatoken tokenizer claims on Apple Silicon, with an educational recreation of the SIMD pretok and hierarchical cache levers.
-
Nostradamous
A from-scratch MLX text-encoder lab that treats predictability as a falsifiable claim — and proved its own premise wrong with collapse-proof measurement.
-
Parameter Golf
Small-language-model training experiments organized around measured bits-per-byte, fixed compute budgets, and reproducible run lineage.
-
Wave-Causal Circuit Search
Verifier-first research on intervention-addressable complex-wave causal models, exact circuit search, and evidence-carrying experiments.
Club B — Method and infrastructure
Verifiers, agent operation, publication archive, simulation, and evolutionary scaffolds that make the loop honest.
-
Agora
A state-first reality-simulation OS: agent populations evolve under Markov rule dynamics, the system predicts macro-dynamics and evolves rules toward targets, and every state lives as plain files in a vault.
-
Open Problems Lab
A provenance-first workspace for formalizing mathematical open problems in Lean: pinned sources, typed registries, and a kernel-checked gate that nothing unproved can cross.
-
Sakana AI Research Lab
Source-grounded study of Sakana AI's research programs — Continuous Thought Machines, AI Scientist, Darwin Gödel Machine — turned into concept notes, falsifiable reproduction gates, and inspectable atlases.
-
Sansara
A vault-native autonomous agent OS: an Obsidian wiki is the world model, a deterministic Rust phase-graph runtime does the work, and the model only ever runs inside graph nodes.
-
Vault Operations
Executable audits and Marimo views over a canonical Obsidian vault: portfolio manager loops, publication projectors, and a polymath research atlas that maps evidence stages without inventing a quality score.
-
AlphaEvolve
An open implementation of DeepMind's AlphaEvolve: an evolutionary coding agent that uses LLMs to iteratively generate, evaluate, and optimize code for algorithmic problems.
-
Geist
An internal agent-research project: can a durable, tool-using agent with memory, task ownership, and team channels operate as a high-leverage member of a company rather than a stateless helper?
-
Lorenz
A verifier-first research-language-model harness for Lean-grounded mathematics: many bounded proof-sketch attempts, compiler-grounded edits, and a population store that only admits artifacts surviving integrity checks.
How publications link
Durable ideas live in the vault. Gated publication notes project to this site. Sibling repos hold notebooks and sealed artifacts; wikilinks stay vault-side.
flowchart LR
vault["Vault notes and literature"]
pubs["Publications: Projects and Research"]
siteProjects["Site project pages"]
siteReports["Site report pages"]
repos["Sibling experiment repos"]
vault -->|"concept notes and wikilinks"| pubs
pubs -->|"publish_research"| siteProjects
pubs -->|"publish_research"| siteReports
repos -->|"sealed commit and artifact hash"| pubs
siteReports -->|"projectId"| siteProjects Reports
RSS- Report Club B
Agora CompactCTM Shadow — Stage-1 Prediction Without Control
Held-out CompactCTM shadow predictor fails to beat an empirical prior on Agora opinion-diffusion transitions; claim class is non-parity, control authority false, not promoted.
- Report Club A
Five-Node Intervention Discovery
A preregistered three-seed pilot on five-node wave-causal circuits: the phase-cancelling circuit is real, but none of the acquisition-strategy superiority claims survived.
- Report Club B
Continuous Thought Machines: A Research Atlas
A source-grounded dissection of Sakana AI's Continuous Thought Machines, paired with an inspectable parity protocol and an explicitly disputed local reproduction state.
- Report Club A
Gigatoken Reproduction — SIMD Pretok and Hierarchical Cache
Independent Apple M5 Max reproduction of Gigatoken's speed claims: 133× vs HuggingFace on GPT-2, 42× vs tiktoken, exact token-ID match — plus an educational recreation of the two systems levers.
- Notebook Club A
Parameter Golf — Run Leaderboard
The live leaderboard of Parameter Golf training runs: every experiment's quantized validation BPB, synced back from remote GPU machines.
- Report Club B
Polymath Research Atlas
Portfolio evidence-pipeline audit: transparent stage flags and metadata-backed bridge retrieval over a personal research vault — not a quality score.
- Report Program
Recursive Self Improvement Lab — Portfolio Map
How Club A evidence engines and Club B method infrastructure connect into a verifier-gated recursive discovery program — and what is still not claimed.
- Notebook Club A
Parameter Golf — Technique Explorer
An interactive catalog of training techniques for small language models, filterable by how early each one shows measurable signal.
- Report Club A
Thirty-Seed Wave-Causal Uncertainty Study
A hash-verified thirty-seed confirmatory study: the selector manipulation passed, but all four preregistered superiority claims were killed at their frozen gates.
- Report Club B
Open Problems Lab — Intake Archive
Verifier-first Lean research program: 15 problem records, 29 prompts, immutable target hashes, and a blocked Mathlib CI gate — zero proofs promoted.
- Report Club A
Measure Before You Predict
A collapse-proof text encoder is anti-forecasting by construction: the working paper that killed its own premise with multi-seed margins and DMD diagnostics, in minutes of compute.
- Report Club B
Agora Calibration Runs
State-first simulation calibration on opinion diffusion and SIR worlds: scored forecasts and rule evolution against explicit targets, not vibes.
- Report Club B
Sansara Operating Surface
Vault-as-universe agent OS with a Rust phase-graph runtime: 70 GiB bloat cut, 152 tests, path-traversal hardening, and a council layer shared with Geist.
- Report Club B
Lorenz Verifier Discipline
A paused Lean RLM harness whose export is proof-integrity discipline — scheduler, planner, and population store that only admits verifier-surviving artifacts — not claimed theorems.
- Report Club B
Geist Council Protocol
Protocol-first durable agent research: confirmation gates, silent no-ops, and a council loop that writes outcomes to long-term memory instead of chat scrollback.
- Report Club B
AlphaEvolve — Early Generations
Open AlphaEvolve infrastructure and scaffolds shipped; recorded evolutionary runs never progressed past early generations before the project paused.