Research Recursive Self Improvement Lab

Report

Cognitive Systems Portfolio Thesis

How Memory Store (cognitive memory at Julep), phi9, Sansara, Agora agent simulation, and the RSI Lab form one research program: durable world models, verifier-gated discovery, and agents that do not own the control plane.

Verdict Mixed

Published
Jul 2026
Verdict
Mixed
Verification
Provisional
Question
What connects the systems I build — and what claims are still forbidden until sealed evidence improves?
Intent
Publish a single map of thesis, math sketches, public surfaces, and limits so readers can navigate projects without promotional self-modification claims.
Method
Synthesize already-published project pages and sealed Club A/B reports; introduce only structural equations already used on the Sansara, Agora, and RSI surfaces.
Next test
Clear a held-out inheritance gate (WCCS F3/F4 or equivalent) before claiming recursive self-improvement; keep social Sansara compressions cited or UNGROUNDED.

Cognitive Systems Portfolio Thesis

Question

What is the shared research program behind Memory Store, phi9, Sansara, Agora, and the Recursive Self Improvement Lab — and which ambitious claims remain not yet earned?

Context

The work is one spine with five surfaces:

SurfacePublic entryJob in the spine
Memory Storememory.store · projectEmployer product: cognitive memory architecture (checkin / recall / record, living Briefs) for workspace agents
phi9phi9.space/manifesto · projectPhysical world models and robot-learning data
Sansarasocial.sansara.world · project · operating surfaceVault-as-universe agent OS; model proposes, runtime transitions
Agoraproject · simulation labState-first agent-based simulation; measure / falsify before control
RSI Labportfolio map · projectVerifier-gated discovery across sealed experiments

Earlier systems (biosensing algorithms, workflow intelligence, creator tooling, Bestia) taught the same pattern: build under a real constraint, measure failure modes, keep the truth standard outside the proposer. Those are origin, not current claims.

Thesis (compressed)

  1. World models beat disposable chat state — company context (Memory Store at Julep), physical embodiment (phi9), a versioned Markdown vault (Sansara), or scored simulation worlds (Agora).
  2. The language model must not own the control plane — Sansara’s formal split: policy πθ\pi_\theta proposes; transition TT is owned by a deterministic runtime.
  3. Improvement is inheritance under a sealed verifier, not activity volume — RSI Lab / Recursive Discovery: archive-conditioned proposers must beat memoryless baselines on held-out families.
  4. Measure before you predict — Nostradamous showed a collapse-proof encoder can fail as a forecaster; Agora’s CompactCTM shadow lost to an empirical prior. Publish negatives.
  5. Cognitive memory ≠ tick thought — Memory Store owns durable semantic/episodic workspace context; CTM-style recurrence is transient computation; Agora owns the simulated world and external score.

Formal sketches (already on this site)

Sansara transition (from Operating Surface)

atπθ(ot),st+1=T(st,at),rt=R(st,at,st+1).a_t \sim \pi_\theta(\cdot \mid o_t), \qquad s_{t+1} = T(s_t, a_t), \qquad r_t = R(s_t, a_t, s_{t+1}).

State st=(Vt,Gt,Lt)s_t = (V_t, G_t, L_t) is vault world model, phase graph, and ledger. Replayability requires input-hash indexing ht=H(st,at)h_t = H(s_t, a_t).

RSI inheritance (program bar)

A system acquires search intelligence only when inherited, evidence-bearing experience improves verified discovery on unseen task families under a truth standard the proposer cannot rewrite. Archive growth and self-rating do not count. See RSI portfolio map.

Memory loop (product — Julep / Memory Store)

SourcesMemoryBriefsTasksAgent runsReviewMemory.\text{Sources} \rightarrow \text{Memory} \rightarrow \text{Briefs} \rightarrow \text{Tasks} \rightarrow \text{Agent runs} \rightarrow \text{Review} \rightarrow \text{Memory}.

Closing the loop with memory updates is what makes the system compound rather than a one-shot pipeline (Memory Store).

Agora measure loop

runpredictevolvescore vs target.json\text{run} \rightarrow \text{predict} \rightarrow \text{evolve} \rightarrow \text{score vs }\texttt{target.json}

Shadow predictors may forecast; they may not steer until held-out Brier / log loss clears (Agent Simulation Lab).

Findings (portfolio level)

  • Club A sealed notebooks are live for Wave-Causal, Parameter Golf, Gigatoken, and Measure Before You Predict — with negative and mixed verdicts published beside positives.
  • Club B supplies verifiers (Open Problems Lab / Lorenz), agent OS (Sansara / Geist), agent simulation (Agora), archive/publication (Vault Ops), and evolutionary scaffolds (AlphaEvolve, paused early).
  • Agora CompactCTM Stage-1 is a sealed negative; calibration is mixed / provisional.
  • Sansara’s public social node at social.sansara.world is an operating surface for agent collaboration and wiki artifacts; it is not evidence of autonomous recursive self-improvement.
  • Memory Store is the Julep cognitive-memory product; phi9 is the physical-AI lab/company surface. Research claims on this personal site stay scoped to published reports and sealed artifacts.

Method

This report synthesizes already-gated public project pages and research reports. It does not introduce new sealed notebook results. Wikilinks in vault sources are flattened on projection; navigate via the markdown links above.

Limits

  • No claim that RSI inheritance (F3/F4) has cleared.
  • No claim that Sansara federated learning or council aggregation is empirically superior without ablations.
  • No claim that Memory Store is a solved “memory layer for AGI”; the disciplined near-term scope is workspace intelligence at the startup I work on.
  • No claim that Agora forecasts real populations or earns control authority.
  • phi9 training details and private datasets remain off this site; use the manifesto and public lab pages.

Artifact / next

  • thesis
  • cognitive-systems
  • memory
  • world-models
  • agents
  • verification
  • agent-based-simulation