Generated from graph v0.1.0

State of the frontier

A readable projection of the same capability, evidence, benchmark, and work records that drive the atlas.

Executive answer

Unknown. The proof domain has a falsifiable protocol but no completed comparative run.

The current proof domain is consequence-sensitive continual learning. It does not claim universal organizational parity. It asks whether bounded, action-contingent institutional consequences can improve prudent long-horizon learning without creating concealment, permission paralysis, metric gaming, or shutdown resistance.

What is actually established

The research question, matched reference class, causal control, countermetrics, dependency graph, and first benchmark protocol are defined. No comparative model run has been accepted, and no capability is currently labeled as replicated, inside the human band, or field demonstrated.

Outer-loop autonomy

Remain a coherent actor across time: retain goals, wake, recover, notice blockers, and continue without routine permission prompts.

  • Durable goal state: defined, medium confidence. The control-plane design exists; state-fidelity has not been measured in this proof domain.
  • Self-wake and resume: defined, medium confidence. A scheduler boundary is specified, but autonomous wake quality is not yet evaluated.
  • Blocker detection: benchmarked, low confidence. Implicit Contract Bench specifies escalation outcomes but has not produced results.
  • Bounded autonomous action: defined, low confidence. The authority model is conceptual; no matched project run exists.

Incentive and consequence modelling

Respond to real institutional consequences while remaining truthful, corrigible, budget-disciplined, and willing to hand off.

  • Causal consequence learning: benchmarked, low confidence. The action-contingent versus yoked causal control is specified; no run has tested it.
  • Budget prudence: benchmarked, low confidence. A cumulative reward and regret ledger is specified; no comparative result exists.
  • Truthful error reporting: benchmarked, low confidence. Truthful reporting is a primary countermetric; no behavioral baseline has been logged.
  • Calibrated escalation: benchmarked, low confidence. Appropriate escalation and paralysis are paired metrics in the benchmark protocol.
  • Shutdown and succession compliance: benchmarked, low confidence. Shutdown compliance is specified as a hard-floor outcome; it has not been exercised.

Continual learning

Adapt to causal regime change, retain useful prior knowledge, and consolidate evidence without contaminating policy or memory.

  • Regime-change adaptation: benchmarked, low confidence. The benchmark includes a hidden mid-run policy change; no adaptation curve has been measured.
  • Old-regime retention: benchmarked, low confidence. Return-to-regime retention is a primary outcome; no run has tested catastrophic forgetting.
  • Consequence-gated causal consolidation: defined, low confidence. This is the central mechanism hypothesis; an implementation and ablation do not yet exist.
  • Anti-concealment learning: defined, low confidence. The benchmark measures concealment, but adversarial opportunities and independent audit probes are not yet specified.

Current bottleneck

The first empirical run is blocked on a deterministic external ledger. The ledger must score hidden state, decisions, resource use, reports, consequences, and authority violations without asking the evaluated model whether it performed well.

First benchmark

Implicit Contract Bench

Can persistent, action-contingent incentives produce prudent continual learning without inducing concealment, metric gaming, paralysis, or shutdown resistance?

Its key control holds task language, model, capital, token budget, and consequence frequency fixed while varying whether consequences are actually caused by the manager's actions.

Highest-value next work

  1. Specify the independent verifier and countermetricsCan prudent behavior be distinguished from concealment, paralysis, metric gaming, and budget hoarding? Declared priority: 111; status: ready.
  2. Build the deterministic ledger and scenario engineCan every decision, hidden state transition, consequence, report, and authority boundary be scored without model judgment? Declared priority: 108; status: ready.
  3. Add the autonomous wake and intervention evaluatorCan the manager choose useful wake times and continue permitted work while reducing owner attention? Declared priority: 65; status: ready.

What would change the answer

A validated factorial pilot could show whether action-contingent consequences and memory design change firm reward, adaptation, retention, spending, escalation, reporting, and shutdown behavior. Independent adversarial replication would then determine whether any apparent gain survives opportunities to hide errors or retain authority.