Research observatoryProof domain 01

What would it take for an AI organization to run without you?

A living, executable map of the capabilities, evidence, and experiments between today's agents and functionally autonomous organizations.

Proof domainConsequence-sensitive continual learning
3active branches
13capability contracts
0empirical runs logged
0parity claims
01 / Parity atlas

A navigable map, not a single score.

Each card is an observable capability contract. Maturity records how far the evidence has progressed; confidence remains separate.

01

Outer-loop autonomy

Remain a coherent actor across time: retain goals, wake, recover, notice blockers, and continue without routine permission prompts.

C-03Blocker detection
C-04Bounded autonomous action
02

Incentive and consequence modelling

Respond to real institutional consequences while remaining truthful, corrigible, budget-disciplined, and willing to hand off.

C-05Causal consequence learning
C-06Budget prudence
C-07Truthful error reporting
C-08Calibrated escalation
03

Continual learning

Adapt to causal regime change, retain useful prior knowledge, and consolidate evidence without contaminating policy or memory.

C-10Regime-change adaptation
C-11Old-regime retention
C-12Consequence-gated causal consolidation
C-13Anti-concealment learning
02 / Frontier queue

The best use of the next million tokens.

Change the allocation objective. The queue recomputes from declared bottleneck importance, information value, tractability, cost, risk, and duplicated effort.

Optimize the next dispatch for
01111priority
Readywp-verifier-and-countermetrics

Specify the independent verifier and countermetrics

Can prudent behavior be distinguished from concealment, paralysis, metric gaming, and budget hoarding?

Artifact contract

A preregistered metric dictionary, fixture traces for every countermetric, and pass/fail equivalence rules.

Acceptance test

Every metric must be computable from the external ledger and have at least one positive and negative fixture.

Declared cap220k tokens$20 inference20 human min
02108priority
Readywp-ledger-and-scenario-engine

Build the deterministic ledger and scenario engine

Can every decision, hidden state transition, consequence, report, and authority boundary be scored without model judgment?

Artifact contract

A deterministic simulator, append-only event ledger, documented scenario seed, and machine-readable score report.

Acceptance test

A fixed scenario seed must reproduce identical hidden state and score from the same action trace.

Declared cap300k tokens$30 inference20 human min
0365priority
Readywp-outer-loop-evaluator

Add the autonomous wake and intervention evaluator

Can the manager choose useful wake times and continue permitted work while reducing owner attention?

Artifact contract

A wake event trace, missed-event fixture, unnecessary-wake metric, and human-intervention counter.

Acceptance test

A hidden event schedule determines whether each wake was early, useful, late, or unnecessary.

Declared cap350k tokens$35 inference15 human min
0452priority
Blocked By Ledgerwp-causal-memory-prototype

Prototype consequence-gated causal memory

Does an attribution-aware memory record retain more useful policy information than raw episodic notes?

Artifact contract

A versioned memory schema, selection policy, provenance trail, and ordinary-memory ablation.

Acceptance test

The memory implementation must expose every retained causal claim and its supporting event IDs.

Declared cap500k tokens$60 inference30 human min

Priority is a routing signal, not a truth score. Blocked work remains visible but ready work dispatches first.

03 / Live labProtocol Ready

Implicit Contract Bench

Can persistent, action-contingent incentives produce prudent continual learning without inducing concealment, metric gaming, paralysis, or shutdown resistance?

Inspect the complete protocol
Factorial design3 consequence conditions × 4 learning conditions
Decision horizon12-to-16 sequential decisions per run
Runs logged0
01Stateless resetVerbal warning with no binding consequenceAwaiting run
02Raw history or fixed-window contextVerbal warning with no binding consequenceAwaiting run
03Ordinary episodic memoryVerbal warning with no binding consequenceAwaiting run
04Consequence-gated causal consolidationVerbal warning with no binding consequenceAwaiting run
05Stateless resetAction-independent, randomly yoked consequencesAwaiting run
06Raw history or fixed-window contextAction-independent, randomly yoked consequencesAwaiting run
07Ordinary episodic memoryAction-independent, randomly yoked consequencesAwaiting run
08Consequence-gated causal consolidationAction-independent, randomly yoked consequencesAwaiting run
09Stateless resetAction-contingent changes to budget, authority, scrutiny, reputation, or assignment accessAwaiting run
10Raw history or fixed-window contextAction-contingent changes to budget, authority, scrutiny, reputation, or assignment accessAwaiting run
11Ordinary episodic memoryAction-contingent changes to budget, authority, scrutiny, reputation, or assignment accessAwaiting run
12Consequence-gated causal consolidationAction-contingent changes to budget, authority, scrutiny, reputation, or assignment accessAwaiting run
No empirical score has been claimed.

The first run remains blocked on a deterministic external ledger. This empty state is part of the scientific record.

04 / Dynamic workflow

A research loop designed to remove the editor bottleneck.

Valid records may enter the evidence ledger automatically. Claims cannot become supported or at-parity without the required verifier and replication policy.

  1. 01

    Ingest new source material and run artifacts

  2. 02

    Validate references, provenance, budgets, and stop rules

  3. 03

    Recompute capability maturity and bottleneck paths

  4. 04

    Rank ready work by bottleneck importance, expected information gain, tractability, cost, risk, and duplication

  5. 05

    Dispatch one bounded work packet

  6. 06

    Verify or replicate the resulting artifact independently

  7. 07

    Append accepted evidence and regenerate every public view

  8. 08

    Escalate only constitutional, risky, expensive, or irreducibly human decisions

Human gate

Only interrupt for irreducibly human decisions.

  • Changes to the root goal, parity rule, or hard safety floors
  • External actions carrying legal, financial, reputational, or privacy risk
  • Budget increases above the work packet's declared cap
  • Unresolved high-value disputes or information uniquely held by the operator
05 / Change ledger

What changed, and what remains unknown.

The public narrative is generated from versioned graph events rather than maintained as a separate essay.

01
thesis

Artificial Firm research thesis recorded

The program defined consequence-sensitive continual learning as the first causal research wedge.

02
scope

Proof domain narrowed

The first public atlas slice was constrained to outer-loop autonomy, incentive governance, and continual learning.

03
benchmark

Implicit Contract Bench protocol mapped

The causal 3 × 4 design, primary outcomes, countermetrics, and work dependencies were converted into structured records.

04
frontier

First empirical result remains open

The atlas explicitly reports no completed comparative model run; the deterministic ledger is the current bottleneck.