AI Analytics Reliability Lab.
Controlled experiments on what makes AI analysts reliable over data. Change one thing at a time, hold the model and questions fixed, and measure what it buys.
Start here
What It Takes to Trust an AI Analyst
Lessons from six experiments on building reliable AI analytics: trusted data, governed definitions, a bounded verification harness, and evaluation that can see silent failure.
Experiment 01 · Grounding
How much does structure buy?
Better grounding materially improves accuracy, but capability improves faster than honesty.
Read: Agentic Analytics: How Much Does Grounding Actually Buy You?Experiment 02 · Reliability
Can the analyst safely say "I don't know"?
Silent error fell from 48.5% to 2.4% across the guardrail study.
Read: Agentic Analytics: Teaching an AI Analyst to Say I Don't KnowExperiment 03 · Protocol
Can an answer show what its claims actually rest on?
Citation repair improved grounded-answer coverage, while the evidence graph exposed missing reasoning structure.
Read: The Evidence Graph: Teaching an AI Analyst to Show Its WorkExperiment 04 · Repair
Which primitive failed, and where should it be fixed?
1,488 graded runs produced a repair matrix mapping failure types to the cheapest correct layer.
Read: The AI-Readiness Repair MatrixExperiment 05 · Ambiguity
Can we predict wrong choices before the agent runs?
In the governed-metric subset, wrong-metric selection dropped from 24% to 0 after ambiguity repair.
Read: Why AI Analysts Pick the Wrong MetricExperiment 06 · Contested
What does an AI analyst do when two governed definitions are both right?
Clarify-by-instruction: 0 of 51. Competing metric attached: 50 of 51. With a disclosure check: 51 of 51.
Read: When Both Numbers Are RightExperiment 06 · Part 2 · Answerability
How does an AI analyst know which questions it can answer?
A closed-world ontology sorts each request into governed, computable, or uninstrumented, so existence and joinability become a graph check rather than a model guess.
Read: How an AI Analyst Knows What It Can Answer
New experiments, as they ship.
I publish each experiment on LinkedIn first, with the method and the numbers. Follow to get the next one.
Field notes, essays and everything else, newest first: all writing →
