Open-source tools for reliable AI analytics.
The research runs on tooling anyone can read, run and check. One tool measures behavior, the other catches ambiguity before an agent ever runs.
AI Analytics Harness
A reproducible lab for reliable AI analysts.
A controlled lab for measuring what makes an LLM analyst reliable over data. Change one thing at a time, hold the model and questions fixed, and measure what it buys you.
Five experiments
Grounding, reliability, protocol, repair and ambiguity, each isolating one variable.
Reproducible
A deterministic synthetic environment, so a run is repeatable and a number is checkable.
Model-agnostic
Hold the questions and grounding fixed and swap the model to compare like for like.
Graded behavior
Correct, incorrect, inconsistent, unsupported, should-refuse and silent-error, not just accuracy.
Preflight
Catch ambiguous analytics before AI does.
Preflight finds ambiguous analytics definitions before an AI agent or human chooses the wrong valid object. Static analysis over your metrics, models and semantic definitions, no agent run required.
Finding taxonomy
Scope traps, concept forks, grain mismatch, definition divergence, duplicate or stale models, ambiguous metrics.
Definition layers
Reads metrics, models and semantic-layer definitions where the ambiguity actually lives.
Broad and cheap
The static layer of the reliability stack: it catches structural ambiguity before anything runs.
Built for CI
Designed to run in a pipeline, so ambiguity is caught on the pull request, not in production.
Want to test these failure modes in your own environment?
The tools measure reliability in a controlled lab. The audit runs the same failure modes against your real warehouse, your semantic layer, and your business questions.
