> ## Documentation Index
> Fetch the complete documentation index at: https://docs.platform.nora.my/llms.txt
> Use this file to discover all available pages before exploring further.

# Coverage

> Is your dataset representative of what your Agents actually see?

A dataset that only tests two cases will pass every fix, no matter how brittle. **Coverage** is the measure of how well your dataset matches your production traffic.

## Coverage view

Datasets → **Coverage** tab. Shows two projections:

* **What's in the dataset** — a distribution of examples across features (intents, tags, entities, difficulty).
* **What's in production** — the same distribution, from real traces.

Side-by-side. Gaps between the two are coverage holes.

## Auto-computed dimensions

Nora clusters both the dataset and production traces into topical clusters (using embeddings), then shows the density on both sides. Common patterns:

* **Underrepresented cluster** — production has 30% of traffic in "refunds" but dataset has 5%. Add refund examples.
* **Overrepresented cluster** — dataset has 20% "greetings" but production has 1%. Prune or reweight.
* **Missing cluster** — production has 10% in "billing disputes" but dataset has 0. Full gap.

## Adding a dimension

You can also cover along custom dimensions:

* **Difficulty** — hard cases might be underrepresented if you seeded with easy ones.
* **Language** — 30% of prod is in Korean, but dataset is all English.
* **User tier** — enterprise users test different behavior than free users.

Configure custom dimensions in **Dataset settings → Coverage dimensions**.

## Filling gaps

The Coverage view suggests specific traces to add:

* Identifies production clusters underrepresented in the dataset.
* Picks representative traces from each cluster.
* Offers **Add to dataset** for one-click seeding.

The suggestions are curated — they prefer high-quality traces (successful, well-formed, with citations) over failed ones. Balance: for regression coverage, add failed ones separately.

## Coverage ratio

A single number: **coverage ratio** = 1 - (production entropy accounted for by dataset). 1.0 = perfect coverage; 0.5 = half the production diversity captured; 0.0 = dataset doesn't represent production at all.

Aim for > 0.7 for datasets used in production quality gates. Below 0.5 = the number simulation gives you is barely meaningful.

## When coverage doesn't matter

Some datasets don't need coverage — they're specifically narrow:

* **Regression suite** — should cover *fixed patterns*, not all production. Coverage is intentionally low.
* **Feature-focused test** — testing "the refund flow" doesn't need coverage of "the login flow."

Tag these datasets `narrow` to suppress the coverage warning.

## Coverage over time

Trend chart: coverage ratio over the last N months. Should stay flat or rise. Falling coverage means production drifted and your dataset stayed static — time to refresh.

## Automated coverage refresh

Turn on **Coverage refresh** to have Nora automatically sample new production traces into the dataset weekly, based on coverage gaps. Every added example is marked as `needs review` for annotation — coverage is filled but you decide what the expected outputs are.

Useful for long-lived Flows where production evolves faster than manual dataset curation.
