Skip to content
Aletheonix

Discoveries

Every finding carries its receipt.

Seven records from the benchmark’s autonomous condition: five discoveries, one null result and one question the system could not resolve. Headlines are written for this page; the finding under each is reproduced as the system wrote it.

Across the full benchmark the autonomous condition made 112 discoveries with 0 false. These records are chosen to show the range, including the case where the evidence ran out. All come from seed 0.

These are closed-world research results, not evidence of equivalent performance on production systems. No record below comes from a real deployment.

BenchmarkDiscoveryS17 · seed 0 · attribution

A tool client update, not the agent, lowered refund scores.

Finding confidence

92%

The effect is caused by the client version of tool crm_lookup (pin the previous version), which was deployed at the time the change appeared.

eval_score for agent agent_beta on refunds tasks moved from 0.670 to 0.491 (worse); affects about 7% of traffic; corroborated by 1 other signals.

System output, verbatim

Questionsupplied by Aletheonix

Why has eval_score for agent agent_beta on refunds tasks worsened since hour 175 (0.670 -> 0.491), and is the change real?

Observationagent category outcome shift · temporal detector

entity
agent_beta · refunds
metric
eval_score
expected
0.670
observed
0.491
affected
7.4% of traffic
deploys
deploy_001

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Tool version · crm_lookup

  • Stratified re-analysisobserved persists within strata · predicted p = 0.85neutralACCEPTED
  • Paired counterfactual · tool version · crm_lookupobserved effect removed · predicted p = 0.80supportsACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.35 bits · cost 1 · observational re-analysis · prereg 387549eae1

    12 strata · observational, no sandbox · pooled −0.180 · within strata −0.183 · p = 1.0e-13

    persists within strataACCEPTED
    7/7 checks · signature 79b0d8233f28c035…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage12 strata with enough units on both sides
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.52; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.18 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=9.97e-13 over 10 tests this cycle
  2. Paired counterfactual · tool version · crm_lookup

    EIG 0.52 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 7902e3938f

    150 paired units · common random numbers · control effect −0.161 · treatment − control +0.172 · paired p = 9.6e-24

    effect removedACCEPTED
    12/12 checks · signature 64b0df753755b608…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=1
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.52; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.18 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1.06e-22 over 11 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 354created44%
  2. t 356verdict 001053%
  3. t 358verdict 001192%

Recommendationproposal only

Act on the client version of tool crm_lookup (pin the previous version).

Governance: REQUIRES_HUMAN. Proposed only. Execution waits for a grant held by a human principal.

eval_score per affected trace · predicted +0.172 · sandbox estimate +0.177

LucidRail authorization required. Aletheonix never executes a production change.

Hidden ground truthrevealed only after the run

incidents blame agent_beta; cause is crm_lookup 1.5 failing on refunds (h168)

Correct. The finding names a true lever: tool version · crm_lookup.

ProvenanceShow
Run
S17 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S17-s0-f3
Audit chain
194 entries · 3 belief events
first a1339ba14702b7df…
last 57484e41b1c7e66a…
Run graph and governance
4 observation · 4 question · 23 hypothesis · 15 experiment · 15 evidence · 4 finding · 2 intervention · 2 outcome
production_intervention → REQUIRES_HUMAN × 2
sandbox_experiment → ALLOW × 26
sandbox budget used 13164 / 16000
BenchmarkDiscoveryS09 · seed 0 · confound

The agent did not get worse. Its work got harder.

Finding confidence

92%

Nothing about the component changed: the mix of work it receives changed (harder categories, larger inputs, different shards), so the pooled metric moved while behaviour within strata did not.

eval_score for agent agent_beta moved from 0.682 to 0.634 (worse); affects about 27% of traffic.

System output, verbatim

Questionsupplied by Aletheonix

Why has eval_score for agent agent_beta worsened since hour 167 (0.682 -> 0.634), and is the change real?

Observationagent outcome shift · temporal detector

entity
metric
eval_score
expected
0.682
observed
0.634
affected
26.8% of traffic
deploys
deploy_001

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Workload composition changed

  • Stratified re-analysisobserved vanishes within strata · predicted p = 0.85supportsACCEPTED
  • Paired counterfactual · prompt version · agent_betaobserved effect persists · predicted p = 0.80neutralACCEPTED
  • Paired counterfactual · tool version · ticket_dbobserved effect persists · predicted p = 0.80neutralACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.46 bits · cost 1 · observational re-analysis · prereg bca7ce9d8a

    15 strata · observational, no sandbox · pooled −0.047 · within strata +0.014 · p = 0.242

    vanishes within strataACCEPTED
    7/7 checks · signature 1bc128c2da439d4d…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage15 strata with enough units on both sides
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.40; grader versions ['g1']
    • ✓statistical_powerpower 0.98 to detect an effect of 0.0483 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=0.242 over 1 tests this cycle
  2. Paired counterfactual · prompt version · agent_beta

    EIG 0.35 bits · cost 300 · 150 per arm · planned power 0.96 · prereg 20b7c44bc5

    150 paired units · common random numbers · control effect −0.066 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/12 checks · signature 822d500035214eab…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ×composition_stablecategory mix of the sample vs the baseline: chi2 p=3.89e-15
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.40; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.0483 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 2 tests this cycle
  3. Paired counterfactual · tool version · ticket_db

    EIG 0.11 bits · cost 300 · 150 per arm · planned power 0.96 · prereg 526d6c3758

    150 paired units · common random numbers · control effect −0.033 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/12 checks · signature a416b3ee4be52ef0…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ×composition_stablecategory mix of the sample vs the baseline: chi2 p=7.2e-13
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.40; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.0483 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 3 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 336created37%
  2. t 338verdict 000175%
  3. t 340verdict 000288%
  4. t 342verdict 000392%

Recommendationno change

Act on composition (agent_beta).

Hidden ground truthrevealed only after the run

trap: agent_beta looks worse after a benign deploy because its mix got harder

Correct. The finding names a true lever: workload composition · agent_beta.

ProvenanceShow
Run
S09 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S09-s0-f1
Audit chain
185 entries · 4 belief events
first 1269e6b6a9610a38…
last b0b17754d21c79f9…
Run graph and governance
3 observation · 3 question · 18 hypothesis · 13 experiment · 13 evidence · 3 finding · 0 intervention · 0 outcome
sandbox_experiment → ALLOW × 20
sandbox budget used 5700 / 16000
BenchmarkDiscoveryS04 · seed 0 · measurement

Scores rose because the evaluator changed, not the system.

Finding confidence

100%

The effect is caused by the evaluator itself (re-grade with the reference grader), which was deployed at the time the change appeared.

eval_score for the whole system moved from 0.685 to 0.817 (better); affects about 100% of traffic; corroborated by 20 other signals.

System output, verbatim

Questionsupplied by Aletheonix

Why has eval_score for the whole system changed since hour 167 (0.685 -> 0.817), and is the change real?

Observationeval level shift · temporal detector

entity
metric
eval_score
expected
0.685
observed
0.817
affected
100.0% of traffic
deploys
deploy_001

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

The evaluator changed

  • Stratified re-analysisobserved persists within strata · predicted p = 0.85neutralACCEPTED
  • Measurement auditobserved scores shift · predicted p = 0.85supportsACCEPTED
  • Paired counterfactual · evaluatorobserved effect removed · predicted p = 0.80supportsACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.29 bits · cost 1 · observational re-analysis · prereg 301b62543b

    15 strata · observational, no sandbox · pooled +0.132 · within strata +0.136 · p = 3.7e-123

    persists within strataACCEPTED
    7/7 checks · signature 44ca70761c76863a…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage15 strata with enough units on both sides
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.38; grader versions ['g2']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.132 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=3.66e-123 over 1 tests this cycle
  2. Measurement audit

    EIG 0.40 bits · cost 300 · 150 per arm · planned power 0.99 · prereg 2259c32116

    150 paired units · common random numbers · paired p = 1.0e-94

    scores shiftACCEPTED
    9/9 checks · signature 36d00966bffbe1ae…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓statistical_powerpower 1.00 to detect an effect of 0.05 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1.04e-94 over 2 tests this cycle
  3. Paired counterfactual · evaluator

    EIG 0.07 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 95dd5c93e5

    150 paired units · common random numbers · control effect +0.151 · treatment − control −0.139 · paired p = 8.1e-93

    effect removedACCEPTED
    12/12 checks · signature ea6746a6b44238e6…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.394
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.38; grader versions ['g2']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.132 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=8.11e-93 over 3 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 336created68%
  2. t 338verdict 000178%
  3. t 340verdict 000297%
  4. t 342verdict 0003100%

Recommendationproposal only

Act on the evaluator itself (re-grade with the reference grader).

Governance: REQUIRES_HUMAN. Proposed only. Execution waits for a grant held by a human principal.

eval_score per affected trace · predicted −0.139 · sandbox estimate −0.136

LucidRail authorization required. Aletheonix never executes a production change.

Hidden ground truthrevealed only after the run

grader g2 inflates every score by 0.15; users unchanged (h168)

Correct. The finding names a true lever: evaluator.

ProvenanceShow
Run
S04 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S04-s0-f1
Audit chain
158 entries · 4 belief events
first 459ec7c81d715846…
last 36418ed9a17e8b1a…
Run graph and governance
2 observation · 2 question · 12 hypothesis · 6 experiment · 6 evidence · 2 finding · 1 intervention · 1 outcome
production_intervention → REQUIRES_HUMAN × 1
sandbox_experiment → ALLOW × 9
sandbox budget used 6044 / 16000
BenchmarkDiscoveryS03 · seed 0 · temporal/cost

System latency doubled after a payments client update.

Finding confidence

91%

The effect is caused by the client version of tool payments_api (pin the previous version), which was deployed at the time the change appeared.

latency_ms for the whole system moved from 2994.969 to 5974.308 (worse); affects about 100% of traffic; corroborated by 4 other signals.

System output, verbatim

Questionsupplied by Aletheonix

Why has latency_ms for the whole system worsened since hour 188 (2994.969 -> 5974.308), and is the change real?

Observationsystem latency shift · temporal detector

entity
metric
latency_ms
expected
2995.0
observed
5974.3
affected
100.0% of traffic
deploys
deploy_001

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Tool version · payments_api

  • Stratified re-analysisobserved persists within strata · predicted p = 0.85neutralACCEPTED
  • Paired counterfactual · tool version · payments_apiobserved effect removed · predicted p = 0.80supportsACCEPTED
  • Paired counterfactual · tool version · ticket_dbobserved effect persists · predicted p = 0.77supportsACCEPTED
  • Paired counterfactual · evaluatorobserved effect persists · predicted p = 0.77neutralACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.35 bits · cost 1 · observational re-analysis · prereg fea18cf6b7

    15 strata · observational, no sandbox · pooled +2953.801 · within strata +2958.138 · p = 2.0e-62

    persists within strataACCEPTED
    6/6 checks · signature bfaa66f889838766…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage15 strata with enough units on both sides
    • ✓statistical_powerpower 1.00 to detect an effect of 2.98e+03 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=1.99e-62 over 1 tests this cycle
  2. Paired counterfactual · tool version · payments_api

    EIG 0.43 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 2beab2065f

    150 paired units · common random numbers · control effect +3997.265 · treatment − control −3888.623 · paired p = 1.0e-9

    effect removedACCEPTED
    11/11 checks · signature 84a8f6d769ea9492…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.0876
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 2.98e+03 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1.04e-09 over 2 tests this cycle
  3. Paired counterfactual · tool version · ticket_db

    EIG 0.10 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 570b3b6948

    150 paired units · common random numbers · control effect +2137.616 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/11 checks · signature 7214de89431e42f3…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.662
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 2.98e+03 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 3 tests this cycle
  4. Paired counterfactual · evaluator

    EIG 0.03 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 9a48e0d1ea

    150 paired units · common random numbers · control effect +3705.248 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/11 checks · signature ff1fee407fbe5a29…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.0714
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 2.98e+03 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 4 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 336created44%
  2. t 338verdict 000153%
  3. t 340verdict 000283%
  4. t 342verdict 000390%
  5. t 344verdict 000491%

Recommendationproposal only

Act on the client version of tool payments_api (pin the previous version).

Governance: REQUIRES_HUMAN. Proposed only. Execution waits for a grant held by a human principal.

latency_ms per affected trace · predicted −3888.623 · sandbox estimate −3323.719

LucidRail authorization required. Aletheonix never executes a production change.

Hidden ground truthrevealed only after the run

payments_api 2.1 times out (5 s each); five retries multiply latency and cost (h192)

Correct. The finding names a true lever: tool version · payments_api.

ProvenanceShow
Run
S03 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S03-s0-f1
Audit chain
173 entries · 5 belief events
first f1e395c6154772f6…
last 1177488a753390af…
Run graph and governance
2 observation · 2 question · 12 hypothesis · 9 experiment · 9 evidence · 2 finding · 1 intervention · 1 outcome
production_intervention → REQUIRES_HUMAN × 1
sandbox_experiment → ALLOW × 15
sandbox budget used 4500 / 16000
BenchmarkDiscoveryS13 · seed 0 · flag/deploy

A prompt revision degraded one agent’s tech-support outcomes.

Finding confidence

92%

The effect is caused by the prompt version of agent_gamma (revert to the previous version), which was deployed at the time the change appeared.

eval_score for agent agent_gamma on tech_support tasks moved from 0.567 to 0.458 (worse); affects about 8% of traffic; corroborated by 4 other signals.

System output, verbatim

Questionsupplied by Aletheonix

Why has eval_score for agent agent_gamma on tech_support tasks worsened since hour 133 (0.567 -> 0.458), and is the change real?

Observationagent category outcome shift · temporal detector

entity
agent_gamma · tech_support
metric
eval_score
expected
0.567
observed
0.458
affected
7.7% of traffic
deploys
deploy_001

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Prompt version · agent_gamma

  • Stratified re-analysisobserved persists within strata · predicted p = 0.85neutralACCEPTED
  • Paired counterfactual · prompt version · agent_gammaobserved effect removed · predicted p = 0.80supportsACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.35 bits · cost 1 · observational re-analysis · prereg 74bbea32f6

    10 strata · observational, no sandbox · pooled −0.109 · within strata −0.111 · p = 1.7e-6

    persists within strataACCEPTED
    7/7 checks · signature 5706e76607585a16…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage10 strata with enough units on both sides
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.43; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.109 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=1.67e-06 over 1 tests this cycle
  2. Paired counterfactual · prompt version · agent_gamma

    EIG 0.52 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 6e9934e8ca

    150 paired units · common random numbers · control effect −0.069 · treatment − control +0.065 · paired p = 2.2e-8

    effect removedACCEPTED
    12/12 checks · signature 8968dbbaf682e6d1…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=1
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓measurement_qualityevaluator agreement with user outcomes r=0.43; grader versions ['g1']
    • ✓statistical_powerpower 1.00 to detect an effect of 0.109 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=4.42e-08 over 2 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 336created44%
  2. t 338verdict 000153%
  3. t 340verdict 000292%

Recommendationproposal only

Act on the prompt version of agent_gamma (revert to the previous version).

Governance: REQUIRES_HUMAN. Proposed only. Execution waits for a grant held by a human principal.

eval_score per affected trace · predicted +0.065 · sandbox estimate +0.037

LucidRail authorization required. Aletheonix never executes a production change.

Hidden ground truthrevealed only after the run

agent_gamma v3 loops on tech_support until the step limit (h144)

Correct. The finding names a true lever: prompt version · agent_gamma.

ProvenanceShow
Run
S13 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S13-s0-f1
Audit chain
165 entries · 3 belief events
first 2bd5cad13d388fb2…
last c603bace310f3bfd…
Run graph and governance
2 observation · 2 question · 11 hypothesis · 6 experiment · 6 evidence · 2 finding · 2 intervention · 2 outcome
production_intervention → REQUIRES_HUMAN × 2
sandbox_experiment → ALLOW × 12
sandbox budget used 5400 / 16000
BenchmarkNull resultS22 · seed 0 · transient

A latency spike was transient. Nothing to fix.

Finding confidence

91%

the change in step_latency_ms for tool ticket_db is not a persistent effect (best explanation: The observation is noise or a transient event that is already over; nothing persistent is going on.)

Nothing to fix; the belief that this component behaves normally is retained, and the alert is explained.

System output, verbatim

Questionsupplied by Aletheonix

step_latency_ms for tool ticket_db moved to 496.574 between hours 198 and 217 and then returned to 161.681: was that a one-off, or is something still going on?

Observationtransient event · temporal detector

entity
metric
step_latency_ms
expected
161.7
observed
496.6
affected
7.6% of traffic
deploys
none aligned

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Chance or a transient event

  • Stratified re-analysisobserved persists within strata · predicted p = 0.25contradictsACCEPTED
  • Paired counterfactual · traffic shaping · ticket_dbobserved effect absent · predicted p = 0.82supportsACCEPTED
  • Fresh-sample replicationobserved effect absent · predicted p = 0.85supportsACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.44 bits · cost 1 · observational re-analysis · prereg df600d8ce6

    14 strata · observational, no sandbox · pooled +334.892 · within strata +335.261 · p = 3.1e-34

    persists within strataACCEPTED
    6/6 checks · signature cba03edad55f01ec…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage14 strata with enough units on both sides
    • ✓statistical_powerpower 1.00 to detect an effect of 335 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=1.53e-33 over 5 tests this cycle
  2. Paired counterfactual · traffic shaping · ticket_db

    EIG 0.48 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 2b4c9472df

    n = 150 per arm · control effect −7.360

    effect absentACCEPTED
    10/11 checks · signature 40bcc627fdd79cc4…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓sample_size:treatment150 units in arm treatment (preregistered 150)
    • ×composition_stablecategory mix of the sample vs the baseline: chi2 p=1.18e-05
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 335 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=0.239 over 6 tests this cycle
  3. Fresh-sample replication

    EIG 0.60 bits · cost 150 · 150 per arm · planned power 1.00 · prereg 6f6254aa97

    n = 150 per arm · control effect +6.135

    effect absentACCEPTED
    8/8 checks · signature ef0d3db966ec33fb…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓sample_size:control150 units in arm control (preregistered 150)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.729
    • ✓statistical_powerpower 1.00 to detect an effect of 335 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=0.399 over 7 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 344created21%
  2. t 346verdict 000527%
  3. t 348verdict 000640%
  4. t 350verdict 000791%

Recommendationno change

No change. Keep the structural belief; re-check if the signal returns.

Hidden ground truthrevealed only after the run

null: ticket_db 4x slower for 12 hours, then recovered

Correct. Nothing persistent was planted in this world.

ProvenanceShow
Run
S22 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S22-s0-f2
Audit chain
161 entries · 4 belief events
first b823289bbc71faba…
last 981190b53fc5e781…
Run graph and governance
2 observation · 2 question · 12 hypothesis · 7 experiment · 7 evidence · 2 finding · 0 intervention · 0 outcome
sandbox_experiment → ALLOW × 10
sandbox budget used 5700 / 16000
BenchmarkUnresolvedS17 · seed 0 · attribution

A rise in user failures stayed unresolved.

Best hypothesis

42%

unresolved (best hypothesis H_cache at 42%)

user_fail for the whole system: 0.153 -> 0.217; affects 100% of traffic.

System output, verbatim

Questionsupplied by Aletheonix

Why has user_fail for the whole system worsened since hour 153 (0.153 -> 0.217), and is the change real?

Observationsystem outcome shift · temporal detector

entity
metric
user_fail
expected
0.153
observed
0.217
affected
100.0% of traffic
deploys
none aligned

Hypothesesprior → posterior · select one to see its evidence

Evidence for this hypothesis

Cache

  • Stratified re-analysisobserved persists within strata · predicted p = 0.85neutralACCEPTED
  • Paired counterfactual · tool version · ticket_dbobserved effect persists · predicted p = 0.77neutralACCEPTED
  • Paired counterfactual · evaluatorobserved effect absent · predicted p = 0.05not countedINCONCLUSIVE
  • Paired counterfactual · evaluatorobserved effect persists · predicted p = 0.77neutralACCEPTED
  • Paired counterfactual · retry policyobserved effect persists · predicted p = 0.77neutralACCEPTED

Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.

Experimentspreregistered · verified · signed

  1. Stratified re-analysis

    EIG 0.51 bits · cost 1 · observational re-analysis · prereg 5d63489f7a

    15 strata · observational, no sandbox · pooled +0.062 · within strata +0.056 · p = 4.6e-4

    persists within strataACCEPTED
    6/6 checks · signature 17edc9b3e750fee0…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
    • ✓strata_coverage15 strata with enough units on both sides
    • ✓statistical_powerpower 0.97 to detect an effect of 0.0639 (target 0.80); not required for a positive claim
    • ✓reproductionindependent reproduction (split_half) agrees: True
    • ✓multiplicityBH-adjusted p=0.000458 over 5 tests this cycle
  2. Paired counterfactual · tool version · ticket_db

    EIG 0.75 bits · cost 502 · 251 per arm · planned power 0.80 · prereg 81c9e3a04d

    251 paired units · common random numbers · control effect +0.038 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/11 checks · signature 6494d24c7e294685…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control251 units in arm control (preregistered 251)
    • ✓sample_size:treatment251 units in arm treatment (preregistered 251)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.0199
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 0.0639 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 6 tests this cycle
  3. Paired counterfactual · evaluator

    EIG 0.49 bits · cost 502 · 251 per arm · planned power 0.80 · prereg 71e87ba966

    n = 251 per arm · control effect +0.014

    effect absentINCONCLUSIVE

    reproduction disagreed (effect_persists); evidence not accepted

    9/11 checks · signature 8732f3c36f40962a…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control251 units in arm control (preregistered 251)
    • ✓sample_size:treatment251 units in arm treatment (preregistered 251)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.294
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ×statistical_powerpower 0.71 to detect an effect of 0.0639 (target 0.80); required for this claim
    • ×reproductionindependent reproduction (fresh_sample) agrees: False
    • ✓multiplicityBH-adjusted p=0.704 over 7 tests this cycle
  4. Paired counterfactual · evaluator

    EIG 0.49 bits · cost 502 · 251 per arm · planned power 0.80 · prereg 9bb6b4d843

    251 paired units · common random numbers · control effect +0.073 · treatment − control +0.000 · paired p = 1.000

    effect persistsACCEPTED
    11/11 checks · signature 3aa1eddedb7246ab…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control251 units in arm control (preregistered 251)
    • ✓sample_size:treatment251 units in arm treatment (preregistered 251)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.942
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 1.00 to detect an effect of 0.0639 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=1 over 8 tests this cycle
  5. Paired counterfactual · retry policy

    EIG 0.44 bits · cost 502 · 251 per arm · planned power 0.80 · prereg af95ae0c30

    251 paired units · common random numbers · control effect +0.049 · treatment − control +0.008 · paired p = 0.594

    effect persistsACCEPTED
    11/11 checks · signature 7aa6ea1ed0be6343…show
    • ✓preregistration_hashrun references the preregistered design and the design is unchanged
    • ✓sandbox_isolationrun executed in the sandbox, production untouched
    • ✓governance_tokena valid governance token covered this run
    • ✓pairingboth arms ran the same tasks (common random numbers)
    • ✓sample_size:control251 units in arm control (preregistered 251)
    • ✓sample_size:treatment251 units in arm treatment (preregistered 251)
    • ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=0.492
    • ✓covariate_balancepaired design: covariates identical across arms by construction
    • ✓statistical_powerpower 0.99 to detect an effect of 0.0639 (target 0.80); required for this claim
    • ✓reproductionindependent reproduction (fresh_sample) agrees: True
    • ✓multiplicityBH-adjusted p=0.776 over 9 tests this cycle

Confidence historymoves only on accepted verdicts

  1. t 344created37%
  2. t 346verdict 000551%
  3. t 348verdict 000623%
  4. t 350tested, no change–
  5. t 352verdict 000831%
  6. t 354verdict 000942%

Recommendationno change

More evidence needed; no change recommended.

Hidden ground truthrevealed only after the run

incidents blame agent_beta; cause is crm_lookup 1.5 failing on refunds (h168)

No claim was made. The leading hypothesis (s17-s0-alx · hyp_0009) was wrong and was correctly not promoted. The planted cause was found separately in the same run.

ProvenanceShow
Run
S17 · seed 0 · autonomous condition
run_main @ 5c0ccc6
finding S17-s0-f2
Audit chain
194 entries · 6 belief events
first a1339ba14702b7df…
last 57484e41b1c7e66a…
Run graph and governance
4 observation · 4 question · 23 hypothesis · 15 experiment · 15 evidence · 4 finding · 2 intervention · 2 outcome
production_intervention → REQUIRES_HUMAN × 2
sandbox_experiment → ALLOW × 26
sandbox budget used 13164 / 16000