Decision Spine
The audit

AI Analytics Reliability Audit

Find the semantic, data-model and guardrail failures that can make your AI analyst confidently wrong.

Approx. 10 working days · Fixed scope · From $5,000

Not ready to book? Score your AI analytics reliability first →

Who this is for

This audit is a fit if:

  • You already have a warehouse or semantic layer.
  • You are testing or deploying an AI analyst.
  • Business users expect reliable natural-language answers.
  • You have real business questions to benchmark.
  • You want to understand failure modes before broader rollout.
What gets tested

Every place a valid query can still be the wrong answer.

  • question interpretation
  • metric selection
  • semantic resolution
  • grain
  • filters
  • joins
  • answerability
  • refusal behavior
  • evidence and provenance
  • consistency across repeated runs
What you provide
  • dbt project or analytics definitions
  • semantic layer or metric definitions, if present
  • existing AI analytics architecture
  • 25 to 50 representative business questions
  • access to a relevant test environment where required
What you get
  • Reliability benchmark

    Correct, incorrect, inconsistent, unsupported, should-refuse and silent-error behavior, measured.

  • Preflight scan

    Structural ambiguity: conflicting definitions, hidden scopes, concept forks, grain mismatches, stale or duplicate objects.

  • Failure map

    Every failure classified by source: selection, construction, coverage, ambiguity or runtime safety.

  • Prioritized repair plan

    Each issue mapped to the cheapest correct fix: model, document, declare, enforce, or remove.

  • Executive readout

    Technical report, executive summary, priority matrix and the highest-risk questions.

  • Focused retest

    Where feasible, repair one to three high-value findings, rerun the benchmark, show the before and after.

The method

Benchmark, diagnose, repair, retest.

Not a generic consulting process. A measurement loop: every claim starts from a number and ends at a re-measured one.

  1. 01

    Benchmark

    Ask realistic business questions and measure what happens.

  2. 02

    Diagnose

    Identify whether the failure comes from selection, construction, coverage, ambiguity or runtime behavior.

  3. 03

    Repair

    Fix the cheapest correct layer: modelling, semantics, documentation, declaration or enforcement.

  4. 04

    Retest

    Re-run the benchmark and measure what actually changed.

Scope and pricing

From $5,000

Fixed scope · Approx. 10 working days

A fixed scope agreed up front, from kickoff to readout. The final figure depends on the size of the stack and the number of business questions in scope. You leave with the benchmark, the failure map, and a prioritized repair plan whichever way the numbers come out.

After the audit

AI Analytics Reliability Sprint

An extension of the Audit: a focused implementation covering semantic modelling, governed metrics, evals, guardrails and agent architecture. Scoped from the highest-risk findings the Audit surfaces.

Book your audit

Can your AI analyst be trusted with your numbers?

Book a short call to work out whether the audit fits your stack, your questions, and where you are in your rollout. No pitch: if it is not a fit, I will tell you.