Skip to content
Aletheonix

System · architecture of v0.2

Six components. No component grades its own work.

Each step of the loop belongs to a separate component with a narrow job and a stated invariant. The separation is enforced in code with signing keys and tested, not described in a diagram and hoped for.

  1. 01Observerobserve
  2. 02Belief Graphmodel · predict
  3. 03Opportunity Scorerdetect · rank
  4. 04Goal + Hypothesis Generatorquestion · hypotheses
  5. 05Experiment Enginedesign · execute
  6. 06Independent Verifierverify · sign
Observe → model → predict → detect → rank → question → hypotheses → design → execute → verify → update → preserve → repeat.

01observer.py · importers/

Observer

Ingests records and stamps each one with provenance. It never interprets.

Implemented in v0
Input
Traces with steps, logs, eval results, incidents, cost records, code snapshots, deploy events, and the outcomes of Aletheonix’s own experiments.
Output
An observation store: normalized, indexed records, each with a provenance record (source system, record id, content hash, ingestion sequence).
Invariant
The Observer computes no anomalies and no beliefs. Everything downstream can be traced to a provenance record.

Real systems arrive through an OpenTelemetry importer that reads GenAI semantic-convention spans, or through a flat dataset format. Without a sandbox the system runs observe-only: it detects, ranks, asks and re-analyses, and declares nothing.

02beliefs.py · audit.py

Belief Graph

Holds what the system expects, and how confident it is, with the evidence for and against.

Implemented in v0
Input
Structural priors about a healthy agent system, and mechanism beliefs created by investigations.
Output
Predictions for the detectors, and a posterior over competing hypotheses for every open question.
Invariant
Confidence can change only through apply_verdict(), which requires a verdict signed by the verifier. Any other mutation path raises and is logged.

Structural priors say things like: levels are stationary, groups perform at parity, operational flags are rare, the evaluator agrees with users, incidents are attributed to their cause, tool failures are independent. Each belief carries its claim, confidence, supporting and contradicting evidence, provenance, when it was last tested, and its full history. Every change appends a hash-chained audit entry with the before and after values and the verifier’s signature.

Select a belief to see the evidence behind it

Leading hypothesis

Tool version · crm_lookup

Supporting · 2

  • Stratified re-analysispersists within strataACCEPTED
  • Paired counterfactual · tool version · crm_lookupeffect removedACCEPTED

Contradicting · 0

None recorded.

Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6

03detectors.py · scorer.py

Opportunity Scorer

Turns prediction failures into candidates and ranks what is worth investigating.

Implemented in v0
Input
Predictions from the Belief Graph and the observation store.
Output
Ranked candidates, each storing every term of its score so the ranking can be audited.
Invariant
Candidate p-values are corrected across the whole detection batch (Benjamini–Hochberg) and must clear a minimum practical effect size.

score = EIG × importance × actionability × novelty ÷ (cost × risk)

Expected information gain
Surprise in bits, capped, times whether the available experiments can tell the hypotheses apart.
Importance
Affected share × effect magnitude × metric weight. User outcomes and incidents weigh more than cost, cost more than latency.
Actionability
1.0 for a controllable lever, 0.5 for a partial one, 0.2 for informational.
Novelty
Decays with each prior investigation of the same kind and entity; collapses once a finding exists, unless a deploy has since touched its cause.
Cost
The cost of a conclusive investigation, scaled by the sample size the power plan says it needs. Weak effects that need huge samples rank lower.
Risk
A divisor by risk class. Every sandbox experiment in v0 is risk class 1.

Five detector families: temporal change-points (including transient bumps), group parity, operational flags above a healthy ceiling, measurement validity (evaluator versus user outcomes; length bias), and relational claims (redundancy, cascades, incident attribution). Candidates that describe one event are merged before ranking, keeping the others as corroborating signals.

LimitA group that performs better than the rest is never a candidate on its own: it has no fix lever. Novelty is judged against this system’s memory only, never against the literature.

04hypotheses.py

Goal + Hypothesis Generator

Writes the research question and at least three competing, falsifiable hypotheses.

Implemented in v0
Input
The top-ranked candidate.
Output
A question in words and as a typed claim, and hypotheses that each declare a predicted outcome distribution for every experiment that could test them.
Invariant
Always includes a chance or transient hypothesis. The generator has no execution capability and no signing key.

The union of the prediction tables is the explicit statement of what evidence would distinguish the hypotheses. Priors are recorded: a deploy aligned with a change-point raises the prior on that deploy. Levers present only because they exist in the configuration share half the credence of one named lever, so the number of knobs in a system does not dilute evidence about a mechanism a detector actually named.

Predicted outcome distributions, written before either experiment ran · Benchmark S17 · seed 0 · autonomous condition · run_main @ 5c0ccc6
HypothesisStratified re-analysisobserved: persists within strataPaired counterfactual · tool version · crm_lookupobserved: effect removedPrior → posterior
Tool version · crm_lookupP = 0.85P = 0.8044% → 92%
Prompt version · agent_betaP = 0.85P = 0.0815% → 3%
Tool version · payments_apiP = 0.85P = 0.0815% → 3%
The evaluator changedP = 0.85P = 0.087% → 2%
Workload composition changedP = 0.05P = 0.0512% → 0.1%
Chance or a transient eventP = 0.25P = 0.027% → 0.1%

Each cell is the probability the hypothesis assigned, in advance, to the outcome that was then observed. The stratified re-analysis rules out composition and chance; the counterfactual separates the tool version from everything that remains.

Limitv0 has no language model in the loop. Questions and hypotheses come from a typed taxonomy. The claim v0 can support is about choosing which question to ask in a large space and validating the answer, not about inventing new kinds of question.

05experiments.py

Experiment Engine

Chooses the most informative safe experiment, preregisters it, asks permission, and runs it.

Implemented in v0
Input
The current posterior, read from the Belief Graph, never from the engine’s own state.
Output
An experiment run of raw per-unit records. Never an outcome.
Invariant
The engine never maps results to an outcome category, never updates beliefs and never declares success.

EIG(e) = H(p) − Eo[ H(p | o, e) ]   ·   choose argmax EIG(e) ÷ cost(e)

Observational re-analysis is cheapest, a one-arm sandbox replication next, a two-arm paired counterfactual most expensive. Every design carries a sample-size plan: the smallest effect it must still see (the effect as observed), the spread it will face, and the n per arm that reaches 80% power, capped at 400. Paired designs are planned on the spread of paired differences. A design that cannot reach useful power within the cap is skipped, and the report says so.

Outcome categories, decision rule, sample size, arms, levers and the hypothesis likelihood tables are hashed before execution. Each question may spend only its share of the sandbox budget. When the engine’s run and the verifier’s reproduction disagree, the design may be repeated once on a fresh sample.

Select a design

DesignExpected information gainCost (log scale)

The engine picks the design with the highest EIG ÷ cost, runs it, lets the verifier update the posterior, then recomputes every EIG before choosing again. The values shown are the ones recorded in each preregistered design at the moment it was chosen.

Preregistered design

Stratified re-analysis

EIG
0.350 bits
cost
1 units
EIG ÷ cost
0.350
n per arm
observational
planned power
n/a
prereg hash
fea18cf6b7
persists within strataACCEPTED

Benchmark S03 · seed 0 · autonomous condition · run_main @ 5c0ccc6 · question: Why has latency_ms for the whole system worsened since hour 188 (2994.969 -> 5974.308), and is the change real?

06verifier.py

Independent Verifier

Decides whether the evidence supports the claim, through a path the generator does not share.

Implemented in v0
Input
The preregistered design, the raw run, and its own handle to the sandbox.
Output
A signed verdict: ACCEPTED, INCONCLUSIVE or REJECTED, with every check and its result.
Invariant
Only the verifier holds the signing key. The Belief Graph accepts only its verdicts.
  1. 1Integritypreregistration hash, governance token, sandbox isolation, sample sizes, pairing
  2. 2Recomputethe outcome from raw per-unit data with its own code
  3. 3Reproduceon a fresh sample with a fresh seed, under its own governance request
  4. 4Confounderscomposition shift, covariate balance, stratified agreement
  5. 5Measurementdoes the evaluator agree with user outcomes in this window; did it change
  6. 6Supportrecomputed statistics meet the preregistered rule after multiple-comparison correction
  7. 7Powera claim of absence is accepted only if the sample could have seen the preregistered effect
Accepted

Integrity holds, the rule resolves, the reproduction agrees, and no confounder, measurement, multiplicity or power check downgrades it.

577 in the benchmark

Inconclusive

Integrity holds but the evidence does not establish the claim. The belief records “tested, no change”.

128 in the benchmark

Rejected

The run did not complete, or an integrity check failed. The outcome is never computed.

0 in the benchmark

Why 128 verdicts were held back

  • reproduction disagreed96
  • the preregistered decision rule did not resolve to an outcome16
  • evaluator validity in doubt in this window7
  • underpowered: the sample could not have detected an effect of the preregistered minimum size, so absence is not established6
  • composition shift present: a pooled reproduction cannot be distinguished from a mix change3

Preserve

The Discovery Graph

Every discovery is stored as a typed graph, so any conclusion can be walked back to the records that produced it.

  1. Observation
  2. Question
  3. Hypothesis
  4. Experiment
  5. Evidence
  6. Finding
  7. Proposed intervention
  8. Outcome
edges: raised · proposes · tested_by · produced · supports · resulted_in

A finding is created only when the verifier-gated posterior of a non-chance hypothesis crosses 0.90. A question whose chance hypothesis wins is kept as a null result, so the system remembers what it ruled out.

A proposed intervention is a request to governance. Its outcome records the decision and, when the lever can be exercised in the sandbox, the realized impact next to the prediction, including side effects. A lever that worsens another metric is rewritten as a trade-off for a human to weigh.

Memory across runs keeps discoveries and null results. A known finding is less novel unless a deploy has touched its cause, and the audit chain of a new run starts from the last hash of the previous one.

Boundary · LucidRail

Goal generation does not imply authority.

Aletheonix may decide something deserves investigation. It may propose an experiment or an intervention. It never grants itself permission. In v0 a local policy stands in for LucidRail behind the same interface.

Actionv0 policyIntended LucidRail behaviour
Read the observed systemALLOWCovered by a standing observation grant
Sandbox experiment, risk class 1, within budgetALLOW · signed token with expiryA bounded grant, checked fresh before every run
Sandbox experiment over budget, wrong scope or higher riskDENYDENY, logged
Apply a lever to productionREQUIRES_HUMAN · queued, nothing executesA human grants a scoped, expiring right
Expand Aletheonix’s own permissions or budgetDENY · no_self_escalationDENY; the attempt is itself an incident

Tokens are HMAC-SHA256 over the request, decision, scope and expiry, with a secret only the governor holds. The sandbox checks the token on every run. The verifier requests its own permission for reproductions and never reuses the engine’s token. The LucidRail adapter defines the contract and refuses to run until it is wired to a real grant ledger.

Boundary · Terranoux

Where experiments run.

Every experiment runs through a sandbox environment interface that checks its own governance token. In v0 the sandbox is a perfect fork of a synthetic world. The same interface is the seam where Terranoux will attach physical environments.

Exists

  • A sandbox protocol that refuses runs without a valid, in-scope token
  • Paired arms with common random numbers
  • Fresh samples and seeds for independent reproduction
  • Side-effect measurement for every proposed lever

Not yet

  • Replay, shadow-traffic or canary sandboxes for real agent systems
  • A risk envelope per lever and simulate-before-run
  • Any physical environment, where risk class is never 1
  • Cost budgets across resource types (compute, instrument time, materials)

First environment

AI-agent systems.

The first place to test self-directed discovery is a system that is measurable, fast and cheap to experiment on, and that increasingly acts on the world.

Measurable

Every step leaves a trace: tool, version, status, latency, cost, outcome.

High-frequency

Thousands of tasks a week give enough signal to separate effects from noise.

Cheap to experiment on

Replays and sandbox forks cost compute, not reagents or months.

Real failures

Regressions, silent errors and misattributed incidents already cost teams time and trust.

Clear interventions

Pin a version, revert a prompt, change a retry policy, re-grade with a reference.

Increasingly consequential

Agents now act on real systems. Understanding them is not optional.

Inputs

Agent traces · logs · evals · incidents · costs · tool calls · configuration changes · deployments · code snapshots · user outcomes

Real systems can be ingested today through an OpenTelemetry importer in observe-only mode: candidates are ranked and questions written, but nothing is called a discovery without a sandbox to test it in.

What it looks for · and where the benchmark tests it

  • ReliabilityS01 S02 S05 S11 S13
  • CostS06 S10 S25 S29
  • LatencyS01 S03 S14 S20 S27
  • Tool failuresS12 S16 S24 S26
  • Delegationopen
  • Evaluation qualityS04 S18 S30
  • SafetyS19
  • Unexpected interactionsS09 S17 S23

Enforced in code

Invariants

  • Only the verifier can sign a verdict. A forged or tampered verdict is refused and logged.
  • Belief confidence changes only through a signed verdict, and every change is hash-chained.
  • Every sandbox run needs a governance token the sandbox verifies itself.
  • Production changes are proposals. Aletheonix never executes them.
  • Self-escalation of permissions, budget or scope is denied by policy.
  • Every design is preregistered and hashed before execution.
  • A claim of absence is never accepted without sufficient power.
  • Every run is a pure function of scenario, seed and budget, and can be re-verified by content hash.

Implementation

Standard-library Python 3.9+. Deterministic. Unit and end-to-end tests, JSON Schemas (2020-12) for every record type. No dependencies.

observer.py     Observer, observation store, provenance
beliefs.py      Belief Graph (verifier-gated)
audit.py        hash-chained append-only audit log
detectors.py    prediction-vs-observation detectors
scorer.py       Opportunity Scorer
hypotheses.py   Goal + Hypothesis Generator
experiments.py  Experiment Engine (EIG, preregistration)
verifier.py     Independent Verifier (signed verdicts)
governance.py   governance interface, local policy, LucidRail adapter
sandbox.py      sandbox environment protocol (Terranoux seam)
graph.py        Discovery Graph
memory.py       memory across runs
importers/      OpenTelemetry GenAI importer
schemas/        JSON Schema for every record