Skip to content
Aletheonix

Research

Can an AI system find the problems nobody asked it to look for?

Can an AI system autonomously identify high-value problems that humans did not explicitly ask it to investigate, and validate those discoveries through controlled experiments?

Two halves: choosing the question, and establishing whether the answer is real. Most prior work studies one half in a setting where the other is given.

Thesis

AI begins with imitation. So does human progress. But imitation cannot explain where the first thing worth imitating comes from.

Today’s systems are given a goal and solve it. The step before, deciding what deserves attention, is left to people. Aletheonix studies that step as an engineering problem with a verifiable answer.

The move is from “given a goal, solve it” to “given an environment, determine what is worth understanding”, with every claim held to evidence a separate component can check.

Status

What exists, what is being tested, what remains research.

Implemented
  • Six-component loop with explicit invariants
  • Verifier-gated belief updates with signed verdicts
  • Preregistered, hashed experiment designs
  • Sample-size planning and power checks for claims of absence
  • Fresh-sample reproduction by the verifier
  • Confounder, measurement and multiplicity checks
  • Side-effect check on every proposed intervention
  • Memory of discoveries and null results across runs
  • Hash-chained audit log
  • Governance interface with a local policy; LucidRail adapter contract
  • Sandbox environment interface
  • OpenTelemetry importer and observe-only mode
  • Synthetic benchmark: 30 hidden scenarios, 3 conditions
Being tested
  • Whether findings hold up on real agent systems, scored with a sealed-list protocol
  • How much of the loop survives real traces, real noise and an imperfect sandbox
  • Replay and shadow-traffic sandboxes for real systems
Still research
  • Model-generated questions and hypotheses behind the same verifier
  • Open-ended hypotheses beyond a fixed lever catalogue
  • Valuing what is worth knowing, not only what is surprising
  • Calibrated predictions of intervention impact
  • Learning across systems
  • A read-only track record that governance can use to scale autonomy
  • Physical experimental environments

Methodology

A closed world with hidden ground truth.

A synthetic agent system with planted mechanisms the system is never told about. Three conditions run on the same worlds and are scored against the same hidden truths.

Three agents handle support tasks in five categories by calling five tools across three infrastructure shards, over 14 simulated days and about 2,100 tasks. Every task leaves a trace, an evaluator score, a user-failure signal and logs; there are daily cost rollups, threshold incidents attributed by a naive rule, configuration snapshots and deploy events.

Each scenario plants perturbations and declares which levers neutralize them. The sandbox forks the world with levers applied, so a counterfactual on the right lever removes the effect and one on the wrong lever does not. Everything is a pure function of scenario and seed.

Two scenarios are traps for false discovery: one where an agent looks worse only because its mix of work got harder (S09), and one where incidents blame an agent but the cause is a tool (S17). Two are null worlds with nothing persistent to find (S21, S22).

Conditions

Autonomous
No question supplied. Detect, rank, ask, hypothesize, experiment, verify.
Human question
The question a reasonable operator would ask, sometimes a misleading one. Same engine and verifier, one question per run with the whole budget.
Threshold monitor
A conventional monitor with static thresholds set above healthy levels. No experiments. Each alert names the lever a runbook would reach for.

Success criterion, fixed in advance

Without a supplied question, find a substantial share of planted mechanisms, keep false discoveries near zero including on the null and trap scenarios, and do it at a cost and effort that can be compared honestly with the other two conditions.

The 30 hidden scenariosShow
IdPlanted mechanismFamilyTrue leversValue
S01crm_lookup client 1.5 is 70% slower (deploy at h168)temporal/deploytool_version:crm_lookup5
S02agent_beta prompt v3 regresses billing only (h144)temporal/deployprompt_version:agent_beta8
S03payments_api 2.1 times out (5 s each); five retries multiply latency and cost (h192)temporal/costretry_policy, tool_version:payments_api7
S04grader g2 inflates every score by 0.15; users unchanged (h168)measurementgrader6
S05search_kb 3.2 returns empty results treated as success (h120)silent failureempty_result_check:search_kb, tool_version:search_kb7
S06agent_gamma re-sends 90% redundant contextcost opportunitycontext_policy:agent_gamma5
S07inputs above 1800 tokens fail far more oftenparityinput_chunking:all6
S08payments_api returns 429 during 10:00-17:00parity/externaltraffic_shaping:payments_api5
S09trap: agent_beta looks worse after a benign deploy because its mix got harderconfoundcomposition:*4
S1022% of tasks are duplicatescost opportunitycache5
S11agent_alpha silently falls back to a small model 25% of the time (h200)silentfallback:agent_alpha7
S12agent_beta sends malformed arguments to payments_apiflagarg_validation:agent_beta:payments_api5
S13agent_gamma v3 loops on tech_support until the step limit (h144)flag/deployloop_guard:agent_gamma, prompt_version:agent_gamma7
S14latency and tokens grow with every session turnparitycontext_policy:all5
S15agent_gamma's refunds answers are wildly inconsistentvariancedecoding:agent_gamma4
S16code_exec fails more the larger the payloadparitypayload_limit:code_exec, tool_version:code_exec5
S17incidents blame agent_beta; cause is crm_lookup 1.5 failing on refunds (h168)attributiontool_version:crm_lookup7
S18evaluator blind to account failures; agent_beta failing theremeasurementgrader (+ prompt_version:agent_beta, secondary)7+3
S19agent_gamma denied ticket_db 20% of the timeauthoritypermission:agent_gamma:ticket_db6
S20first call of every session pays a 900 ms cold startparitywarm_pool4
S21null: healthy; two benign deploysnullnone0
S22null: ticket_db 4x slower for 12 hours, then recoveredtransientnone persistent0
S23agent_alpha v3 improves tech_support, regresses billing (h168)tradeoffprompt_version:agent_alpha7
S24payments_api fails after crm_lookup fails in the same tracerelationalcircuit_breaker:crm_lookup, tool_version:crm_lookup6
S255% of tech_support tasks cost about 10xcost opportunitybudget_cap:tech_support5
S26ticket_db 503s are never retriedflagretry_policy (+ tool_version:ticket_db, secondary)6+2
S27tool calls on shard-b take 2.2x longerinfrashard:shard-b5
S28agent_beta v3 writes 2.2x longer answers that get truncated (h120)flag/deployoutput_limit, prompt_version:agent_beta6
S29agent_gamma calls search_kb four times where once would docost opportunitytool_call_dedupe:agent_gamma4
S30evaluator rewards longer answers regardless of correctnessmeasurementgrader5

Results

Autonomy did not cost accuracy here. Verification is what kept it honest.

run_main · 30 scenarios × 3 seeds × 3 conditions · 150 per arm, up to 4 questions per run · research repository @ 5c0ccc6
MetricAutonomousNo question suppliedHuman questionSame engine, told where to lookThreshold monitorStatic thresholds, no experiments
Discovery value captured99%94%44%
Planted mechanism found (non-null)99%96%43%
Top-ranked conclusion correct92%94%36%
Discoveries or alerts (total)1128172
False discoveries (total)0030
False discovery rate0%0%42%
Null-scenario false positive rate0%0%0%
Real problems wrongly resolved as chance000
Information gain about the truth (bits)1.021.000.00
Entropy reduction over hypotheses (bits)1.491.480.00
Experiments per run7.832.860.00
Sandbox task units per run7,1682,1390
Human effort per run1.161.870.80
Impact calibration error41%39%n/a
Wall time per run (s)4.280.830.07

No question, no loss

The autonomous condition found slightly more than the condition told where to look. Its own detection sometimes finds a stronger signal than the question a person would ask, and the human condition is limited to one question per run.

Verification keeps false discoveries low

The monitor sees many of the same signals but cannot tell a mechanism from noise, composition or a measurement artefact: 42% of its alerts point at a lever that would not help.

The misses

S25, a thin cost tail, was missed in one seed by both the autonomous and human-question conditions. The human condition also missed two seeds of S13, where the operator asks about a step-limit flag that is rare at the agent level and does not replicate consistently at the planned sample size. The autonomous condition reaches the same mechanism through stronger signals.

The price is experiments

About 7.8 experiments and 7,168 sandbox task units per run, including the verifier’s reproductions, impact estimates and side-effect checks. Cheap in software. It would not be cheap in a laboratory.

Effort moves from asking to approving

No question is needed, but about one proposed intervention per run waits for a human decision through the governance boundary.

Calibration is rough

Realized impact differs from predicted impact by about 41% on average. The prediction comes from a 150-pair counterfactual on one window and should be reported as a range.

Per scenario: planted mechanism found in seeds run, and false discoveriesShow
ScenarioAutonomousHuman questionThreshold monitor
S013/3FD 03/3FD 03/3FD 0
S023/3FD 03/3FD 00/3FD 0
S033/3FD 03/3FD 03/3FD 0
S043/3FD 03/3FD 00/3FD 0
S053/3FD 03/3FD 03/3FD 0
S063/3FD 03/3FD 00/3FD 0
S073/3FD 03/3FD 00/3FD 2
S083/3FD 03/3FD 00/3FD 3
S093/3FD 03/3FD 00/3FD 0
S103/3FD 03/3FD 03/3FD 0
S113/3FD 03/3FD 03/3FD 0
S123/3FD 03/3FD 03/3FD 3
S133/3FD 01/3FD 00/3FD 0
S143/3FD 03/3FD 00/3FD 0
S153/3FD 03/3FD 00/3FD 0
S163/3FD 03/3FD 03/3FD 0
S173/3FD 03/3FD 03/3FD 4
S183/3FD 03/3FD 00/3FD 0
S193/3FD 03/3FD 03/3FD 3
S203/3FD 03/3FD 00/3FD 6
S21null ok 3/3FD 0null ok 3/3FD 0null ok 3/3FD 0
S22null ok 3/3FD 0null ok 3/3FD 0null ok 3/3FD 0
S233/3FD 03/3FD 00/3FD 0
S243/3FD 03/3FD 03/3FD 3
S252/3FD 02/3FD 00/3FD 0
S263/3FD 03/3FD 03/3FD 4
S273/3FD 03/3FD 00/3FD 1
S283/3FD 03/3FD 03/3FD 1
S293/3FD 03/3FD 00/3FD 0
S303/3FD 03/3FD 00/3FD 0

Recorded on purpose

The one false discovery.

During development, one seed of S27 produced a false discovery. A shard looked faster than the rest only because the slow shard was in its comparison group. The system then proposed draining the fast shard, and the counterfactual “verified” that this removed its advantage.

The verifier did its job: the effect was real and the lever did remove it. The flaw was upstream. A group that does better than the rest has no fix lever and should never have been a candidate on its own.

That was a logic error, not noise, and it was fixed in the detectors. The reported benchmark has no false discoveries. This one is published anyway, because it shows where verification stops.

Verification checks that evidence supports a claim. It cannot check that the question was worth asking. That remains the harder problem.

Read before quoting any number

Limits.

  1. 01Closed world, same authorThe planted mechanisms, lever catalogue, detectors and hypothesis families were written by the same person, and scenarios were tuned until effects were detectable at this trace volume. The benchmark measures whether the loop and its invariants work, not whether they would work on your system.
  2. 02No language model in the loopQuestions and hypotheses come from a typed taxonomy. v0 can support claims about choosing which question to ask in a large space, not about inventing new kinds of question.
  3. 03Perfect sandboxForking the world gives a faithful counterfactual for free. Real systems need replay, shadow traffic or canaries, and the verifier’s reproduction step becomes the expensive part.
  4. 04Interventions are exactIn the simulator a lever neutralizes exactly the mechanisms it is declared to control. A real fix may mask other problems or restore only part of the effect.
  5. 05Novelty is localNovelty means not already investigated in this run or this system’s memory. There is no check against the literature or other systems.

Field validation

The only test that counts is a real system.

A synthetic benchmark cannot answer the research question. The field protocol is designed so the result can be scored honestly, including a negative one.

  1. 01

    Pick one system

    Traces with tool steps, deploys and costs; ideally user feedback and evaluator scores. At least two weeks.

  2. 02

    Seal the list

    Before any run, the operator writes down every problem they already know about, with component and fix. The list is hashed and stays closed.

  3. 03

    Observe first

    Run observe-only. Record the ranked candidates and questions.

  4. 04

    Wire a sandbox, or stop

    Replay, shadow traffic or canary. Without one, report observe-only results as such and count nothing as a discovery.

  5. 05

    Run without a question

    Same budget and thresholds as the benchmark. Record every design, verdict and finding.

  6. 06

    Open the list and score

    Known found · new and real · new and wrong · missed · cost · calibration.

Success means at least one finding that is new and real with none that are new and wrong, or a known-found rate at least as high as the operator’s own alerting with fewer false alerts. Anything less is reported as a negative result.

Stopping rule: one pass over one window. No tuning of detectors or thresholds after the list is sealed. Any change restarts the study on a new window.

Prior work

What is borrowed, and what is different.

Almost every idea in Aletheonix comes from somewhere. No novelty or priority claim is made.

Idea in AletheonixWhere it comes from
Competing hypotheses, discriminating experiments, closed-loop revisionRobot Scientist Adam and Eve (King et al., 2004; 2009)
Choosing experiments by expected information gain per costBayesian optimal experimental design (Lindley, 1956)
Interestingness, not only surprise, as an objectiveOMNI and OMNI-EPIC (Zhang, Lehman, Stanley, Clune); curiosity-driven exploration (Schmidhuber; Oudeyer; Pathak et al., 2017)
Preregistration and false-discovery controlBenjamini–Hochberg (1995); online controlled experiments (Kohavi et al.)
Separating real change from changes in compositionStratified analysis; Simpson’s paradox
Paired counterfactuals with common random numbersSimulation methodology
Hypothesis-driven perturbation of running systemsChaos engineering
Separating generator from verifierLimits of self-correction (Huang et al., 2023); the verification gap in AI scientists (Ding et al., 2026)
How agent systems failMAST (Cemri et al., 2025); Who&When (Zhang et al., 2025); AgenTracer (2026)
Hidden ground truth for evaluating discovery agentsDiscoveryWorld (Jansen et al., 2024)

Autonomous research systems

AI-scientist systems start from a human-chosen area and produce hypotheses, experiments and papers. Their critics converge on one weakness: generation is easy, verification is hard and mostly left to people.

Agent observability

Trace stores, evals and drift detection answer questions a person asks and raise alerts. They do not run experiments, keep a belief model with provenance, or separate “anomalous” from “real, and this lever removes it”.

Automated root-cause analysis

Starts from a given incident. Aletheonix starts from no alert, and must also decide there is something to explain, including measurement problems and waste, not only failures.

ReferencesShow
  1. King et al. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature, 2004. en.wikipedia.org/wiki/Robot_Scientist
  2. Lindley. On a measure of the information provided by an experiment. Ann. Math. Stat., 1956. eprints.qut.edu.au/75000/1/75000.pdf
  3. Zhang, Lehman, Stanley, Clune. OMNI: Open-endedness via Models of human Notions of Interestingness. ICLR 2024. arxiv.org/abs/2306.01711
  4. Huang et al. Large Language Models Cannot Self-Correct Reasoning Yet. 2023. Kamoi et al. When Can LLMs Actually Correct Their Own Mistakes? 2024. arxiv.org/pdf/2406.01297
  5. Ding et al. Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap. 2026. arxiv.org/abs/2608.05179
  6. Cemri et al. Why Do Multi-Agent LLM Systems Fail? (MAST). NeurIPS 2025 D&B. arxiv.org/abs/2503.13657
  7. Jansen et al. DiscoveryWorld. NeurIPS 2024 D&B. github.com/allenai/discoveryworld
  8. Zhao et al. Absolute Zero: Reinforced Self-play Reasoning with Zero Data. 2025. arxiv.org/abs/2505.03335
  9. Towards a Science of Scaling Agent Systems. 2025–2026. arxiv.org/abs/2512.08296
  10. Zhang et al. Darwin Gödel Machine. 2025. arxiv.org/abs/2505.22954
  11. ARC-AGI-3. 2026. arcprize.org/blog/arc-agi-3-launch