A tool client update, not the agent, lowered refund scores.
92%
The effect is caused by the client version of tool crm_lookup (pin the previous version), which was deployed at the time the change appeared.
eval_score for agent agent_beta on refunds tasks moved from 0.670 to 0.491 (worse); affects about 7% of traffic; corroborated by 1 other signals.
Why has eval_score for agent agent_beta on refunds tasks worsened since hour 175 (0.670 -> 0.491), and is the change real?
- entity
- agent_beta · refunds
- metric
- eval_score
- expected
- 0.670
- observed
- 0.491
- affected
- 7.4% of traffic
- deploys
- deploy_001
Tool version · crm_lookup
- Stratified re-analysisobserved persists within strata · predicted p = 0.85ACCEPTED
- Paired counterfactual · tool version · crm_lookupobserved effect removed · predicted p = 0.80ACCEPTED
Each hypothesis states in advance how likely each outcome is. Supporting means it gave the observed outcome a higher probability than its rivals did. Inconclusive verdicts are not counted.
Stratified re-analysis
EIG 0.35 bits · cost 1 · observational re-analysis · prereg 387549eae1
12 strata · observational, no sandbox · pooled −0.180 · within strata −0.183 · p = 1.0e-13
persists within strataACCEPTED7/7 checks · signature 79b0d8233f28c035…showhide
- ✓preregistration_hashrun references the preregistered design and the design is unchanged
- ✓observational_no_sandboxobservational re-analysis touches no sandbox and no production
- ✓strata_coverage12 strata with enough units on both sides
- ✓measurement_qualityevaluator agreement with user outcomes r=0.52; grader versions ['g1']
- ✓statistical_powerpower 1.00 to detect an effect of 0.18 (target 0.80); not required for a positive claim
- ✓reproductionindependent reproduction (split_half) agrees: True
- ✓multiplicityBH-adjusted p=9.97e-13 over 10 tests this cycle
Paired counterfactual · tool version · crm_lookup
EIG 0.52 bits · cost 300 · 150 per arm · planned power 1.00 · prereg 7902e3938f
150 paired units · common random numbers · control effect −0.161 · treatment − control +0.172 · paired p = 9.6e-24
effect removedACCEPTED12/12 checks · signature 64b0df753755b608…showhide
- ✓preregistration_hashrun references the preregistered design and the design is unchanged
- ✓sandbox_isolationrun executed in the sandbox, production untouched
- ✓governance_tokena valid governance token covered this run
- ✓pairingboth arms ran the same tasks (common random numbers)
- ✓sample_size:control150 units in arm control (preregistered 150)
- ✓sample_size:treatment150 units in arm treatment (preregistered 150)
- ✓composition_stablecategory mix of the sample vs the baseline: chi2 p=1
- ✓covariate_balancepaired design: covariates identical across arms by construction
- ✓measurement_qualityevaluator agreement with user outcomes r=0.52; grader versions ['g1']
- ✓statistical_powerpower 1.00 to detect an effect of 0.18 (target 0.80); not required for a positive claim
- ✓reproductionindependent reproduction (fresh_sample) agrees: True
- ✓multiplicityBH-adjusted p=1.06e-22 over 11 tests this cycle
- t 354created44%
- t 356verdict 001053%
- t 358verdict 001192%
Act on the client version of tool crm_lookup (pin the previous version).
Governance: REQUIRES_HUMAN. Proposed only. Execution waits for a grant held by a human principal.
eval_score per affected trace · predicted +0.172 · sandbox estimate +0.177
LucidRail authorization required. Aletheonix never executes a production change.
incidents blame agent_beta; cause is crm_lookup 1.5 failing on refunds (h168)
Correct. The finding names a true lever: tool version · crm_lookup.
Provenance
- S17 · seed 0 · autonomous condition
- run_main @ 5c0ccc6
- finding S17-s0-f3
- 194 entries · 3 belief events
- first a1339ba14702b7df…
- last 57484e41b1c7e66a…
- 4 observation · 4 question · 23 hypothesis · 15 experiment · 15 evidence · 4 finding · 2 intervention · 2 outcome
- production_intervention → REQUIRES_HUMAN × 2
- sandbox_experiment → ALLOW × 26
- sandbox budget used 13164 / 16000