Can an AI system find the problems nobody asked it to look for?
Can an AI system autonomously identify high-value problems that humans did not explicitly ask it to investigate, and validate those discoveries through controlled experiments?
Two halves: choosing the question, and establishing whether the answer is real. Most prior work studies one half in a setting where the other is given.
AI begins with imitation. So does human progress. But imitation cannot explain where the first thing worth imitating comes from.
Today’s systems are given a goal and solve it. The step before, deciding what deserves attention, is left to people. Aletheonix studies that step as an engineering problem with a verifiable answer.
The move is from “given a goal, solve it” to “given an environment, determine what is worth understanding”, with every claim held to evidence a separate component can check.
What exists, what is being tested, what remains research.
- Six-component loop with explicit invariants
- Verifier-gated belief updates with signed verdicts
- Preregistered, hashed experiment designs
- Sample-size planning and power checks for claims of absence
- Fresh-sample reproduction by the verifier
- Confounder, measurement and multiplicity checks
- Side-effect check on every proposed intervention
- Memory of discoveries and null results across runs
- Hash-chained audit log
- Governance interface with a local policy; LucidRail adapter contract
- Sandbox environment interface
- OpenTelemetry importer and observe-only mode
- Synthetic benchmark: 30 hidden scenarios, 3 conditions
- Whether findings hold up on real agent systems, scored with a sealed-list protocol
- How much of the loop survives real traces, real noise and an imperfect sandbox
- Replay and shadow-traffic sandboxes for real systems
- Model-generated questions and hypotheses behind the same verifier
- Open-ended hypotheses beyond a fixed lever catalogue
- Valuing what is worth knowing, not only what is surprising
- Calibrated predictions of intervention impact
- Learning across systems
- A read-only track record that governance can use to scale autonomy
- Physical experimental environments
A closed world with hidden ground truth.
A synthetic agent system with planted mechanisms the system is never told about. Three conditions run on the same worlds and are scored against the same hidden truths.
Three agents handle support tasks in five categories by calling five tools across three infrastructure shards, over 14 simulated days and about 2,100 tasks. Every task leaves a trace, an evaluator score, a user-failure signal and logs; there are daily cost rollups, threshold incidents attributed by a naive rule, configuration snapshots and deploy events.
Each scenario plants perturbations and declares which levers neutralize them. The sandbox forks the world with levers applied, so a counterfactual on the right lever removes the effect and one on the wrong lever does not. Everything is a pure function of scenario and seed.
Two scenarios are traps for false discovery: one where an agent looks worse only because its mix of work got harder (S09), and one where incidents blame an agent but the cause is a tool (S17). Two are null worlds with nothing persistent to find (S21, S22).
- Autonomous
- No question supplied. Detect, rank, ask, hypothesize, experiment, verify.
- Human question
- The question a reasonable operator would ask, sometimes a misleading one. Same engine and verifier, one question per run with the whole budget.
- Threshold monitor
- A conventional monitor with static thresholds set above healthy levels. No experiments. Each alert names the lever a runbook would reach for.
Without a supplied question, find a substantial share of planted mechanisms, keep false discoveries near zero including on the null and trap scenarios, and do it at a cost and effort that can be compared honestly with the other two conditions.
The 30 hidden scenarios
| Id | Planted mechanism | Family | True levers | Value |
|---|---|---|---|---|
| S01 | crm_lookup client 1.5 is 70% slower (deploy at h168) | temporal/deploy | tool_version:crm_lookup | 5 |
| S02 | agent_beta prompt v3 regresses billing only (h144) | temporal/deploy | prompt_version:agent_beta | 8 |
| S03 | payments_api 2.1 times out (5 s each); five retries multiply latency and cost (h192) | temporal/cost | retry_policy, tool_version:payments_api | 7 |
| S04 | grader g2 inflates every score by 0.15; users unchanged (h168) | measurement | grader | 6 |
| S05 | search_kb 3.2 returns empty results treated as success (h120) | silent failure | empty_result_check:search_kb, tool_version:search_kb | 7 |
| S06 | agent_gamma re-sends 90% redundant context | cost opportunity | context_policy:agent_gamma | 5 |
| S07 | inputs above 1800 tokens fail far more often | parity | input_chunking:all | 6 |
| S08 | payments_api returns 429 during 10:00-17:00 | parity/external | traffic_shaping:payments_api | 5 |
| S09 | trap: agent_beta looks worse after a benign deploy because its mix got harder | confound | composition:* | 4 |
| S10 | 22% of tasks are duplicates | cost opportunity | cache | 5 |
| S11 | agent_alpha silently falls back to a small model 25% of the time (h200) | silent | fallback:agent_alpha | 7 |
| S12 | agent_beta sends malformed arguments to payments_api | flag | arg_validation:agent_beta:payments_api | 5 |
| S13 | agent_gamma v3 loops on tech_support until the step limit (h144) | flag/deploy | loop_guard:agent_gamma, prompt_version:agent_gamma | 7 |
| S14 | latency and tokens grow with every session turn | parity | context_policy:all | 5 |
| S15 | agent_gamma's refunds answers are wildly inconsistent | variance | decoding:agent_gamma | 4 |
| S16 | code_exec fails more the larger the payload | parity | payload_limit:code_exec, tool_version:code_exec | 5 |
| S17 | incidents blame agent_beta; cause is crm_lookup 1.5 failing on refunds (h168) | attribution | tool_version:crm_lookup | 7 |
| S18 | evaluator blind to account failures; agent_beta failing there | measurement | grader (+ prompt_version:agent_beta, secondary) | 7+3 |
| S19 | agent_gamma denied ticket_db 20% of the time | authority | permission:agent_gamma:ticket_db | 6 |
| S20 | first call of every session pays a 900 ms cold start | parity | warm_pool | 4 |
| S21 | null: healthy; two benign deploys | null | none | 0 |
| S22 | null: ticket_db 4x slower for 12 hours, then recovered | transient | none persistent | 0 |
| S23 | agent_alpha v3 improves tech_support, regresses billing (h168) | tradeoff | prompt_version:agent_alpha | 7 |
| S24 | payments_api fails after crm_lookup fails in the same trace | relational | circuit_breaker:crm_lookup, tool_version:crm_lookup | 6 |
| S25 | 5% of tech_support tasks cost about 10x | cost opportunity | budget_cap:tech_support | 5 |
| S26 | ticket_db 503s are never retried | flag | retry_policy (+ tool_version:ticket_db, secondary) | 6+2 |
| S27 | tool calls on shard-b take 2.2x longer | infra | shard:shard-b | 5 |
| S28 | agent_beta v3 writes 2.2x longer answers that get truncated (h120) | flag/deploy | output_limit, prompt_version:agent_beta | 6 |
| S29 | agent_gamma calls search_kb four times where once would do | cost opportunity | tool_call_dedupe:agent_gamma | 4 |
| S30 | evaluator rewards longer answers regardless of correctness | measurement | grader | 5 |
Autonomy did not cost accuracy here. Verification is what kept it honest.
| Metric | AutonomousNo question supplied | Human questionSame engine, told where to look | Threshold monitorStatic thresholds, no experiments |
|---|---|---|---|
| Discovery value captured | 99% | 94% | 44% |
| Planted mechanism found (non-null) | 99% | 96% | 43% |
| Top-ranked conclusion correct | 92% | 94% | 36% |
| Discoveries or alerts (total) | 112 | 81 | 72 |
| False discoveries (total) | 0 | 0 | 30 |
| False discovery rate | 0% | 0% | 42% |
| Null-scenario false positive rate | 0% | 0% | 0% |
| Real problems wrongly resolved as chance | 0 | 0 | 0 |
| Information gain about the truth (bits) | 1.02 | 1.00 | 0.00 |
| Entropy reduction over hypotheses (bits) | 1.49 | 1.48 | 0.00 |
| Experiments per run | 7.83 | 2.86 | 0.00 |
| Sandbox task units per run | 7,168 | 2,139 | 0 |
| Human effort per run | 1.16 | 1.87 | 0.80 |
| Impact calibration error | 41% | 39% | n/a |
| Wall time per run (s) | 4.28 | 0.83 | 0.07 |
No question, no loss
The autonomous condition found slightly more than the condition told where to look. Its own detection sometimes finds a stronger signal than the question a person would ask, and the human condition is limited to one question per run.
Verification keeps false discoveries low
The monitor sees many of the same signals but cannot tell a mechanism from noise, composition or a measurement artefact: 42% of its alerts point at a lever that would not help.
The misses
S25, a thin cost tail, was missed in one seed by both the autonomous and human-question conditions. The human condition also missed two seeds of S13, where the operator asks about a step-limit flag that is rare at the agent level and does not replicate consistently at the planned sample size. The autonomous condition reaches the same mechanism through stronger signals.
The price is experiments
About 7.8 experiments and 7,168 sandbox task units per run, including the verifier’s reproductions, impact estimates and side-effect checks. Cheap in software. It would not be cheap in a laboratory.
Effort moves from asking to approving
No question is needed, but about one proposed intervention per run waits for a human decision through the governance boundary.
Calibration is rough
Realized impact differs from predicted impact by about 41% on average. The prediction comes from a 150-pair counterfactual on one window and should be reported as a range.
Per scenario: planted mechanism found in seeds run, and false discoveries
| Scenario | Autonomous | Human question | Threshold monitor |
|---|---|---|---|
| S01 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S02 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S03 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S04 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S05 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S06 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S07 | 3/3FD 0 | 3/3FD 0 | 0/3FD 2 |
| S08 | 3/3FD 0 | 3/3FD 0 | 0/3FD 3 |
| S09 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S10 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S11 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S12 | 3/3FD 0 | 3/3FD 0 | 3/3FD 3 |
| S13 | 3/3FD 0 | 1/3FD 0 | 0/3FD 0 |
| S14 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S15 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S16 | 3/3FD 0 | 3/3FD 0 | 3/3FD 0 |
| S17 | 3/3FD 0 | 3/3FD 0 | 3/3FD 4 |
| S18 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S19 | 3/3FD 0 | 3/3FD 0 | 3/3FD 3 |
| S20 | 3/3FD 0 | 3/3FD 0 | 0/3FD 6 |
| S21 | null ok 3/3FD 0 | null ok 3/3FD 0 | null ok 3/3FD 0 |
| S22 | null ok 3/3FD 0 | null ok 3/3FD 0 | null ok 3/3FD 0 |
| S23 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S24 | 3/3FD 0 | 3/3FD 0 | 3/3FD 3 |
| S25 | 2/3FD 0 | 2/3FD 0 | 0/3FD 0 |
| S26 | 3/3FD 0 | 3/3FD 0 | 3/3FD 4 |
| S27 | 3/3FD 0 | 3/3FD 0 | 0/3FD 1 |
| S28 | 3/3FD 0 | 3/3FD 0 | 3/3FD 1 |
| S29 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
| S30 | 3/3FD 0 | 3/3FD 0 | 0/3FD 0 |
The one false discovery.
During development, one seed of S27 produced a false discovery. A shard looked faster than the rest only because the slow shard was in its comparison group. The system then proposed draining the fast shard, and the counterfactual “verified” that this removed its advantage.
The verifier did its job: the effect was real and the lever did remove it. The flaw was upstream. A group that does better than the rest has no fix lever and should never have been a candidate on its own.
That was a logic error, not noise, and it was fixed in the detectors. The reported benchmark has no false discoveries. This one is published anyway, because it shows where verification stops.
Verification checks that evidence supports a claim. It cannot check that the question was worth asking. That remains the harder problem.
Limits.
- Closed world, same authorThe planted mechanisms, lever catalogue, detectors and hypothesis families were written by the same person, and scenarios were tuned until effects were detectable at this trace volume. The benchmark measures whether the loop and its invariants work, not whether they would work on your system.
- No language model in the loopQuestions and hypotheses come from a typed taxonomy. v0 can support claims about choosing which question to ask in a large space, not about inventing new kinds of question.
- Perfect sandboxForking the world gives a faithful counterfactual for free. Real systems need replay, shadow traffic or canaries, and the verifier’s reproduction step becomes the expensive part.
- Interventions are exactIn the simulator a lever neutralizes exactly the mechanisms it is declared to control. A real fix may mask other problems or restore only part of the effect.
- Novelty is localNovelty means not already investigated in this run or this system’s memory. There is no check against the literature or other systems.
The only test that counts is a real system.
A synthetic benchmark cannot answer the research question. The field protocol is designed so the result can be scored honestly, including a negative one.
Pick one system
Traces with tool steps, deploys and costs; ideally user feedback and evaluator scores. At least two weeks.
Seal the list
Before any run, the operator writes down every problem they already know about, with component and fix. The list is hashed and stays closed.
Observe first
Run observe-only. Record the ranked candidates and questions.
Wire a sandbox, or stop
Replay, shadow traffic or canary. Without one, report observe-only results as such and count nothing as a discovery.
Run without a question
Same budget and thresholds as the benchmark. Record every design, verdict and finding.
Open the list and score
Known found · new and real · new and wrong · missed · cost · calibration.
Success means at least one finding that is new and real with none that are new and wrong, or a known-found rate at least as high as the operator’s own alerting with fewer false alerts. Anything less is reported as a negative result.
Stopping rule: one pass over one window. No tuning of detectors or thresholds after the list is sealed. Any change restarts the study on a new window.
What is borrowed, and what is different.
Almost every idea in Aletheonix comes from somewhere. No novelty or priority claim is made.
| Idea in Aletheonix | Where it comes from |
|---|---|
| Competing hypotheses, discriminating experiments, closed-loop revision | Robot Scientist Adam and Eve (King et al., 2004; 2009) |
| Choosing experiments by expected information gain per cost | Bayesian optimal experimental design (Lindley, 1956) |
| Interestingness, not only surprise, as an objective | OMNI and OMNI-EPIC (Zhang, Lehman, Stanley, Clune); curiosity-driven exploration (Schmidhuber; Oudeyer; Pathak et al., 2017) |
| Preregistration and false-discovery control | Benjamini–Hochberg (1995); online controlled experiments (Kohavi et al.) |
| Separating real change from changes in composition | Stratified analysis; Simpson’s paradox |
| Paired counterfactuals with common random numbers | Simulation methodology |
| Hypothesis-driven perturbation of running systems | Chaos engineering |
| Separating generator from verifier | Limits of self-correction (Huang et al., 2023); the verification gap in AI scientists (Ding et al., 2026) |
| How agent systems fail | MAST (Cemri et al., 2025); Who&When (Zhang et al., 2025); AgenTracer (2026) |
| Hidden ground truth for evaluating discovery agents | DiscoveryWorld (Jansen et al., 2024) |
Autonomous research systems
AI-scientist systems start from a human-chosen area and produce hypotheses, experiments and papers. Their critics converge on one weakness: generation is easy, verification is hard and mostly left to people.
Agent observability
Trace stores, evals and drift detection answer questions a person asks and raise alerts. They do not run experiments, keep a belief model with provenance, or separate “anomalous” from “real, and this lever removes it”.
Automated root-cause analysis
Starts from a given incident. Aletheonix starts from no alert, and must also decide there is something to explain, including measurement problems and waste, not only failures.
References
- King et al. Functional genomic hypothesis generation and experimentation by a robot scientist. Nature, 2004. en.wikipedia.org/wiki/Robot_Scientist
- Lindley. On a measure of the information provided by an experiment. Ann. Math. Stat., 1956. eprints.qut.edu.au/75000/1/75000.pdf
- Zhang, Lehman, Stanley, Clune. OMNI: Open-endedness via Models of human Notions of Interestingness. ICLR 2024. arxiv.org/abs/2306.01711
- Huang et al. Large Language Models Cannot Self-Correct Reasoning Yet. 2023. Kamoi et al. When Can LLMs Actually Correct Their Own Mistakes? 2024. arxiv.org/pdf/2406.01297
- Ding et al. Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap. 2026. arxiv.org/abs/2608.05179
- Cemri et al. Why Do Multi-Agent LLM Systems Fail? (MAST). NeurIPS 2025 D&B. arxiv.org/abs/2503.13657
- Jansen et al. DiscoveryWorld. NeurIPS 2024 D&B. github.com/allenai/discoveryworld
- Zhao et al. Absolute Zero: Reinforced Self-play Reasoning with Zero Data. 2025. arxiv.org/abs/2505.03335
- Towards a Science of Scaling Agent Systems. 2025–2026. arxiv.org/abs/2512.08296
- Zhang et al. Darwin Gödel Machine. 2025. arxiv.org/abs/2505.22954
- ARC-AGI-3. 2026. arcprize.org/blog/arc-agi-3-launch