# AI Evals Results Table

Prompt eval suite: each nightly run is a column and each golden set case a row.

- id: `ai-evals-results-table`
- Page: https://agentic-ui-kit.com/components/ai-evals-results-table
- Live preview: https://agentic-ui-kit.com/live/component/ai-evals-results-table

## Install

```bash
agentic-ui add ai-evals-results-table
npx shadcn@4.17.0 add @agentic-ui/ai-evals-results-table
```

## Metadata

- **Level:** feature
- **Category:** feature
- **Installable in:** Radix UI (registry variant new-york)
- **Capabilities:** runs-by-cases-matrix, score-cell-heatmap, faceted-filtering, full-text-case-search, per-case-trend-sparkline, regression-detection, threshold-driven-pass-rate, column-sorting, cell-drilldown, keyboard-grid-navigation, empty-error-loading-states
- **Intents:** evaluate-ai, compare-over-time, track-status, explore-detail
- **Suits domains:** ai-workspace, devtools, analytics, operations, saas
- **Interaction:** faceted-filtering, case-search, threshold-adjustment, column-sorting, cell-drilldown, keyboard-grid-navigation
- **Screen role:** management
- **Density:** compact
- **Dark mode:** yes
- **Use it for:** prompt eval suite: each nightly run is a column and each golden set case a row, CI console for models that flags which cases regressed relative to the previous run before promoting a version, rubric-based quality audit filtering only the safety cases that fail in the latest run, review of a specific case by opening the cell to read the grader feedback and its latency
- **Avoid when:** there is only one run and therefore no time axis or trend (use a simple results table), the aggregated score distribution is wanted instead of case-by-case detail (use ai-eval-score-chart), the span tree of a single execution is the object of study (use ai-trace-detail)
- **Preserve when adapting:** matrix-orientation-runs-as-columns, threshold-driven-cell-tinting, trend-derivation-from-last-two-runs, keyboard-grid-contract, empty-error-loading-states
- **Replace when adapting:** domain-evaluation-data, criteria-taxonomy, run-labels-and-versions, grader-feedback-copy
- **Adaptation effort:** low
- **Coupling:** low
- **Quality score:** 9.2/10
- **Quality breakdown:** visual 9/10, responsive 9/10, accessibility 9/10, code 9/10
- **Uses primitives:** badge, button, card, table
- **Motion:** profile subtle, library css, reduced-motion fallback yes
- **npm dependencies (registry):** lucide-react@0.577.0
- **Registry dependencies:** lib-primitives
