Recursive Self Improvement Lab

Propose. Verify. Inherit only when held-out evidence says so.

An attempt to build a recursive self-improvement lab from sealed local projects — not from self-rating. Club A carries the evidence notebooks; Club B carries verifiers, agent runtimes, and archives. Negative results stay published beside positive ones.

Program

  1. Recursive Self Improvement Lab

    A verifier-gated research program assembled from sealed local experiments: propose, intervene, verify, archive, and inherit only when held-out evidence improves.

    Active

How the lab connects

Improvement is earned only on the last arrow: archive-conditioned reuse must beat memoryless baselines on untouched families under a truth standard the proposer cannot rewrite. That inheritance gate is queued, not claimed.

flowchart TB
  propose["Propose / vary
AlphaEvolve · Parameter Golf · Sansara"] intervene["Intervene / run
WCCS · PG budgets · Gigatoken"] verify["Verify
Oracles · Lean gates · sealed hashes"] measure["Measure / falsify
Nostradamous · Agora"] archive["Archive
Vault Ops · Publications"] propose --> intervene --> verify --> measure --> archive archive -->|"held-out inheritance only"| propose

Club A — Evidence engines

Researcher-usable reports and sealed notebooks: Question, Context, Findings, Method, Limits, Artifact.

  1. Gigatoken Lab

    Independent reproduction of Marcel Rød's Gigatoken tokenizer claims on Apple Silicon, with an educational recreation of the SIMD pretok and hierarchical cache levers.

    Active Notebook
  2. Nostradamous

    A from-scratch MLX text-encoder lab that treats predictability as a falsifiable claim — and proved its own premise wrong with collapse-proof measurement.

    Paused Notebook
  3. Parameter Golf

    Small-language-model training experiments organized around measured bits-per-byte, fixed compute budgets, and reproducible run lineage.

    Active Notebook
  4. Wave-Causal Circuit Search

    Verifier-first research on intervention-addressable complex-wave causal models, exact circuit search, and evidence-carrying experiments.

    Active Notebook

Club B — Method and infrastructure

Verifiers, agent operation, publication archive, simulation, and evolutionary scaffolds that make the loop honest.

  1. Agora

    State-first agent-based simulation lab: Markov populations, scored forecasts against target distributions, and a verifier-first shadow brain that stays powerless until held-out gates clear.

    Active
  2. Open Problems Lab

    A provenance-first workspace for formalizing mathematical open problems in Lean: pinned sources, typed registries, and a kernel-checked gate that nothing unproved can cross.

    Active
  3. Sakana AI Research Lab

    Source-grounded study of Sakana AI's research programs — Continuous Thought Machines, AI Scientist, Darwin Gödel Machine — turned into concept notes, falsifiable reproduction gates, and inspectable atlases.

    Active
  4. Sansara

    A vault-native autonomous agent OS: an Obsidian wiki is the world model, a deterministic Rust phase-graph runtime does the work, and the model only ever runs inside graph nodes.

    Active
  5. Vault Operations

    Executable audits and Marimo views over a canonical Obsidian vault: portfolio manager loops, publication projectors, and a polymath research atlas that maps evidence stages without inventing a quality score.

    Active
  6. AlphaEvolve

    An open implementation of DeepMind's AlphaEvolve: an evolutionary coding agent that uses LLMs to iteratively generate, evaluate, and optimize code for algorithmic problems.

    Paused
  7. Geist

    An internal agent-research project: can a durable, tool-using agent with memory, task ownership, and team channels operate as a high-leverage member of a company rather than a stateless helper?

    Active
  8. Lorenz

    A verifier-first research-language-model harness for Lean-grounded mathematics: many bounded proof-sketch attempts, compiler-grounded edits, and a population store that only admits artifacts surviving integrity checks.

    Paused

Other projects

  1. Memory Store

    The cognitive memory architecture product I build at Julep AI — durable checkin / recall / record, living Briefs, and a workspace world model so agents compound instead of re-deriving context every session.

    Active
  2. phi9

    Physical AI lab and data company: high-fidelity motion capture, egocentric video, and VLA research aimed at world models that act in the physical world.

    Active

How publications link

Durable ideas live in the vault. Gated publication notes project to this site. Sibling repos hold notebooks and sealed artifacts; wikilinks stay vault-side.

flowchart LR
  vault["Vault notes and literature"]
  pubs["Publications: Projects and Research"]
  siteProjects["Site project pages"]
  siteReports["Site report pages"]
  repos["Sibling experiment repos"]

  vault -->|"concept notes and wikilinks"| pubs
  pubs -->|"publish_research"| siteProjects
  pubs -->|"publish_research"| siteReports
  repos -->|"sealed commit and artifact hash"| pubs
  siteReports -->|"projectId"| siteProjects

Reports

RSS
  1. Report Club B

    Agent Simulation Lab: What I Built and How

    Agora as a state-first agent-based simulation lab: Markov worlds, scored forecasts, verifier-first shadow brains, and a sealed CompactCTM negative — with authority boundaries that keep Memory Store out of tick-level control.

    Jul 2026 Agora Mixed
  2. Report Club B

    Agora CompactCTM Shadow — Stage-1 Prediction Without Control

    Held-out CompactCTM shadow predictor fails to beat an empirical prior on Agora opinion-diffusion transitions; claim class is non-parity, control authority false, not promoted.

    Jul 2026 Agora Negative
  3. Report Program

    Cognitive Systems Portfolio Thesis

    How Memory Store (cognitive memory at Julep), phi9, Sansara, Agora agent simulation, and the RSI Lab form one research program: durable world models, verifier-gated discovery, and agents that do not own the control plane.

  4. Report Club B

    Continuous Thought Machines: A Research Atlas

    A source-grounded dissection of Sakana AI's Continuous Thought Machines, paired with an inspectable parity protocol and an explicitly disputed local reproduction state.

  5. Report Club A

    Five-Node Intervention Discovery

    A preregistered three-seed pilot on five-node wave-causal circuits: the phase-cancelling circuit is real, but none of the acquisition-strategy superiority claims survived.

  6. Report Club A

    Gigatoken Reproduction — SIMD Pretok and Hierarchical Cache

    Independent Apple M5 Max reproduction of Gigatoken's speed claims: 133× vs HuggingFace on GPT-2, 42× vs tiktoken, exact token-ID match — plus an educational recreation of the two systems levers.

    Jul 2026 Gigatoken Lab Positive Notebook
  7. Notebook Club A

    Parameter Golf — Run Leaderboard

    The live leaderboard of Parameter Golf training runs: every experiment's quantized validation BPB, synced back from remote GPU machines.

    Jul 2026 Parameter Golf Not applicable Notebook
  8. Report Program

    Recursive Self Improvement Lab — Portfolio Map

    How Club A evidence engines and Club B method infrastructure connect into a verifier-gated recursive discovery program — and what is still not claimed.

  9. Notebook Club A

    Parameter Golf — Technique Explorer

    An interactive catalog of training techniques for small language models, filterable by how early each one shows measurable signal.

    Jul 2026 Parameter Golf Not applicable Notebook
  10. Report Club B

    Polymath Research Atlas

    Portfolio evidence-pipeline audit: transparent stage flags and metadata-backed bridge retrieval over a personal research vault — not a quality score.

    Jul 2026 Vault Operations Mixed
  11. Report Club A

    Thirty-Seed Wave-Causal Uncertainty Study

    A hash-verified thirty-seed confirmatory study: the selector manipulation passed, but all four preregistered superiority claims were killed at their frozen gates.

  12. Report Club B

    Open Problems Lab — Intake Archive

    Verifier-first Lean research program: 15 problem records, 29 prompts, immutable target hashes, and a blocked Mathlib CI gate — zero proofs promoted.

    Jul 2026 Open Problems Lab Inconclusive
  13. Report Club A

    Measure Before You Predict

    A collapse-proof text encoder is anti-forecasting by construction: the working paper that killed its own premise with multi-seed margins and DMD diagnostics, in minutes of compute.

  14. Report Club B

    Agora Calibration Runs

    State-first simulation calibration on opinion diffusion and SIR worlds: scored forecasts and rule evolution against explicit targets, not vibes.

    Jun 2026 Agora Mixed
  15. Report Club B

    Sansara Operating Surface

    Vault-as-universe agent OS with a Rust phase-graph runtime: 70 GiB bloat cut, 152 tests, path-traversal hardening, and a council layer shared with Geist.

    Jul 2026 Sansara Mixed
  16. Report Club B

    Lorenz Verifier Discipline

    A paused Lean RLM harness whose export is proof-integrity discipline — scheduler, planner, and population store that only admits verifier-surviving artifacts — not claimed theorems.

    May 2026 Lorenz Inconclusive
  17. Report Club B

    Geist Council Protocol

    Protocol-first durable agent research: confirmation gates, silent no-ops, and a council loop that writes outcomes to long-term memory instead of chat scrollback.

    May 2026 Geist Mixed
  18. Report Club B

    AlphaEvolve — Early Generations

    Open AlphaEvolve infrastructure and scaffolds shipped; recorded evolutionary runs never progressed past early generations before the project paused.

    Jul 2025 AlphaEvolve Negative