Six components. No component grades its own work.
Each step of the loop belongs to a separate component with a narrow job and a stated invariant. The separation is enforced in code with signing keys and tested, not described in a diagram and hoped for.
- Observer
- Belief Graph
- Opportunity Scorer
- Goal + Hypothesis Generator
- Experiment Engine
- Independent Verifier
Observer
Ingests records and stamps each one with provenance. It never interprets.
Implemented in v0- Traces with steps, logs, eval results, incidents, cost records, code snapshots, deploy events, and the outcomes of Aletheonix’s own experiments.
- An observation store: normalized, indexed records, each with a provenance record (source system, record id, content hash, ingestion sequence).
- The Observer computes no anomalies and no beliefs. Everything downstream can be traced to a provenance record.
Real systems arrive through an OpenTelemetry importer that reads GenAI semantic-convention spans, or through a flat dataset format. Without a sandbox the system runs observe-only: it detects, ranks, asks and re-analyses, and declares nothing.
Belief Graph
Holds what the system expects, and how confident it is, with the evidence for and against.
Implemented in v0- Structural priors about a healthy agent system, and mechanism beliefs created by investigations.
- Predictions for the detectors, and a posterior over competing hypotheses for every open question.
- Confidence can change only through apply_verdict(), which requires a verdict signed by the verifier. Any other mutation path raises and is logged.
Structural priors say things like: levels are stationary, groups perform at parity, operational flags are rare, the evaluator agrees with users, incidents are attributed to their cause, tool failures are independent. Each belief carries its claim, confidence, supporting and contradicting evidence, provenance, when it was last tested, and its full history. Every change appends a hash-chained audit entry with the before and after values and the verifier’s signature.
Tool version · crm_lookup
- Stratified re-analysispersists within strataACCEPTED
- Paired counterfactual · tool version · crm_lookupeffect removedACCEPTED
None recorded.
Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6
Opportunity Scorer
Turns prediction failures into candidates and ranks what is worth investigating.
Implemented in v0- Predictions from the Belief Graph and the observation store.
- Ranked candidates, each storing every term of its score so the ranking can be audited.
- Candidate p-values are corrected across the whole detection batch (Benjamini–Hochberg) and must clear a minimum practical effect size.
score = EIG × importance × actionability × novelty ÷ (cost × risk)
- Expected information gain
- Surprise in bits, capped, times whether the available experiments can tell the hypotheses apart.
- Importance
- Affected share × effect magnitude × metric weight. User outcomes and incidents weigh more than cost, cost more than latency.
- Actionability
- 1.0 for a controllable lever, 0.5 for a partial one, 0.2 for informational.
- Novelty
- Decays with each prior investigation of the same kind and entity; collapses once a finding exists, unless a deploy has since touched its cause.
- Cost
- The cost of a conclusive investigation, scaled by the sample size the power plan says it needs. Weak effects that need huge samples rank lower.
- Risk
- A divisor by risk class. Every sandbox experiment in v0 is risk class 1.
Five detector families: temporal change-points (including transient bumps), group parity, operational flags above a healthy ceiling, measurement validity (evaluator versus user outcomes; length bias), and relational claims (redundancy, cascades, incident attribution). Candidates that describe one event are merged before ranking, keeping the others as corroborating signals.
A group that performs better than the rest is never a candidate on its own: it has no fix lever. Novelty is judged against this system’s memory only, never against the literature.
Goal + Hypothesis Generator
Writes the research question and at least three competing, falsifiable hypotheses.
Implemented in v0- The top-ranked candidate.
- A question in words and as a typed claim, and hypotheses that each declare a predicted outcome distribution for every experiment that could test them.
- Always includes a chance or transient hypothesis. The generator has no execution capability and no signing key.
The union of the prediction tables is the explicit statement of what evidence would distinguish the hypotheses. Priors are recorded: a deploy aligned with a change-point raises the prior on that deploy. Levers present only because they exist in the configuration share half the credence of one named lever, so the number of knobs in a system does not dilute evidence about a mechanism a detector actually named.
| Hypothesis | Stratified re-analysisobserved: persists within strata | Paired counterfactual · tool version · crm_lookupobserved: effect removed | Prior → posterior |
|---|---|---|---|
| Tool version · crm_lookup | P = 0.85 | P = 0.80 | 44% → 92% |
| Prompt version · agent_beta | P = 0.85 | P = 0.08 | 15% → 3% |
| Tool version · payments_api | P = 0.85 | P = 0.08 | 15% → 3% |
| The evaluator changed | P = 0.85 | P = 0.08 | 7% → 2% |
| Workload composition changed | P = 0.05 | P = 0.05 | 12% → 0.1% |
| Chance or a transient event | P = 0.25 | P = 0.02 | 7% → 0.1% |
Each cell is the probability the hypothesis assigned, in advance, to the outcome that was then observed. The stratified re-analysis rules out composition and chance; the counterfactual separates the tool version from everything that remains.
v0 has no language model in the loop. Questions and hypotheses come from a typed taxonomy. The claim v0 can support is about choosing which question to ask in a large space and validating the answer, not about inventing new kinds of question.
Experiment Engine
Chooses the most informative safe experiment, preregisters it, asks permission, and runs it.
Implemented in v0- The current posterior, read from the Belief Graph, never from the engine’s own state.
- An experiment run of raw per-unit records. Never an outcome.
- The engine never maps results to an outcome category, never updates beliefs and never declares success.
EIG(e) = H(p) − Eo[ H(p | o, e) ] · choose argmax EIG(e) ÷ cost(e)
Observational re-analysis is cheapest, a one-arm sandbox replication next, a two-arm paired counterfactual most expensive. Every design carries a sample-size plan: the smallest effect it must still see (the effect as observed), the spread it will face, and the n per arm that reaches 80% power, capped at 400. Paired designs are planned on the spread of paired differences. A design that cannot reach useful power within the cap is skipped, and the report says so.
Outcome categories, decision rule, sample size, arms, levers and the hypothesis likelihood tables are hashed before execution. Each question may spend only its share of the sandbox budget. When the engine’s run and the verifier’s reproduction disagree, the design may be repeated once on a fresh sample.
The engine picks the design with the highest EIG ÷ cost, runs it, lets the verifier update the posterior, then recomputes every EIG before choosing again. The values shown are the ones recorded in each preregistered design at the moment it was chosen.
Stratified re-analysis
- EIG
- 0.350 bits
- cost
- 1 units
- EIG ÷ cost
- 0.350
- n per arm
- observational
- planned power
- n/a
- prereg hash
- fea18cf6b7
Benchmark S03 · seed 0 · autonomous condition · run_main @ 5c0ccc6 · question: Why has latency_ms for the whole system worsened since hour 188 (2994.969 -> 5974.308), and is the change real?
Independent Verifier
Decides whether the evidence supports the claim, through a path the generator does not share.
Implemented in v0- The preregistered design, the raw run, and its own handle to the sandbox.
- A signed verdict: ACCEPTED, INCONCLUSIVE or REJECTED, with every check and its result.
- Only the verifier holds the signing key. The Belief Graph accepts only its verdicts.
- Integritypreregistration hash, governance token, sandbox isolation, sample sizes, pairing
- Recomputethe outcome from raw per-unit data with its own code
- Reproduceon a fresh sample with a fresh seed, under its own governance request
- Confounderscomposition shift, covariate balance, stratified agreement
- Measurementdoes the evaluator agree with user outcomes in this window; did it change
- Supportrecomputed statistics meet the preregistered rule after multiple-comparison correction
- Powera claim of absence is accepted only if the sample could have seen the preregistered effect
Integrity holds, the rule resolves, the reproduction agrees, and no confounder, measurement, multiplicity or power check downgrades it.
577 in the benchmark
Integrity holds but the evidence does not establish the claim. The belief records “tested, no change”.
128 in the benchmark
The run did not complete, or an integrity check failed. The outcome is never computed.
0 in the benchmark
- reproduction disagreed96
- the preregistered decision rule did not resolve to an outcome16
- evaluator validity in doubt in this window7
- underpowered: the sample could not have detected an effect of the preregistered minimum size, so absence is not established6
- composition shift present: a pooled reproduction cannot be distinguished from a mix change3
The Discovery Graph
Every discovery is stored as a typed graph, so any conclusion can be walked back to the records that produced it.
- Observation
- Question
- Hypothesis
- Experiment
- Evidence
- Finding
- Proposed intervention
- Outcome
A finding is created only when the verifier-gated posterior of a non-chance hypothesis crosses 0.90. A question whose chance hypothesis wins is kept as a null result, so the system remembers what it ruled out.
A proposed intervention is a request to governance. Its outcome records the decision and, when the lever can be exercised in the sandbox, the realized impact next to the prediction, including side effects. A lever that worsens another metric is rewritten as a trade-off for a human to weigh.
Memory across runs keeps discoveries and null results. A known finding is less novel unless a deploy has touched its cause, and the audit chain of a new run starts from the last hash of the previous one.
Where experiments run.
Every experiment runs through a sandbox environment interface that checks its own governance token. In v0 the sandbox is a perfect fork of a synthetic world. The same interface is the seam where Terranoux will attach physical environments.
- A sandbox protocol that refuses runs without a valid, in-scope token
- Paired arms with common random numbers
- Fresh samples and seeds for independent reproduction
- Side-effect measurement for every proposed lever
- Replay, shadow-traffic or canary sandboxes for real agent systems
- A risk envelope per lever and simulate-before-run
- Any physical environment, where risk class is never 1
- Cost budgets across resource types (compute, instrument time, materials)
AI-agent systems.
The first place to test self-directed discovery is a system that is measurable, fast and cheap to experiment on, and that increasingly acts on the world.
Measurable
Every step leaves a trace: tool, version, status, latency, cost, outcome.
High-frequency
Thousands of tasks a week give enough signal to separate effects from noise.
Cheap to experiment on
Replays and sandbox forks cost compute, not reagents or months.
Real failures
Regressions, silent errors and misattributed incidents already cost teams time and trust.
Clear interventions
Pin a version, revert a prompt, change a retry policy, re-grade with a reference.
Increasingly consequential
Agents now act on real systems. Understanding them is not optional.
Agent traces · logs · evals · incidents · costs · tool calls · configuration changes · deployments · code snapshots · user outcomes
Real systems can be ingested today through an OpenTelemetry importer in observe-only mode: candidates are ranked and questions written, but nothing is called a discovery without a sandbox to test it in.
- Reliabilityregressions after deploys, silent failuresS01 S02 S05 S11 S13
- Costredundant context, duplicates, long tailsS06 S10 S25 S29
- Latencyslow clients, retry storms, cold startsS01 S03 S14 S20 S27
- Tool failuresmalformed arguments, cascades, unretried errorsS12 S16 S24 S26
- Delegationhandoffs between agents · not yet in the benchmarkopen
- Evaluation qualitygrader drift, blind spots, length biasS04 S18 S30
- Safetypermission denials only so farS19
- Unexpected interactionsconfounded mixes, misattributed incidents, trade-offsS09 S17 S23
Invariants
- Only the verifier can sign a verdict. A forged or tampered verdict is refused and logged.
- Belief confidence changes only through a signed verdict, and every change is hash-chained.
- Every sandbox run needs a governance token the sandbox verifies itself.
- Production changes are proposals. Aletheonix never executes them.
- Self-escalation of permissions, budget or scope is denied by policy.
- Every design is preregistered and hashed before execution.
- A claim of absence is never accepted without sufficient power.
- Every run is a pure function of scenario, seed and budget, and can be re-verified by content hash.
Standard-library Python 3.9+. Deterministic. Unit and end-to-end tests, JSON Schemas (2020-12) for every record type. No dependencies.
observer.py Observer, observation store, provenance beliefs.py Belief Graph (verifier-gated) audit.py hash-chained append-only audit log detectors.py prediction-vs-observation detectors scorer.py Opportunity Scorer hypotheses.py Goal + Hypothesis Generator experiments.py Experiment Engine (EIG, preregistration) verifier.py Independent Verifier (signed verdicts) governance.py governance interface, local policy, LucidRail adapter sandbox.py sandbox environment protocol (Terranoux seam) graph.py Discovery Graph memory.py memory across runs importers/ OpenTelemetry GenAI importer schemas/ JSON Schema for every record