Guides

Testing

Testing

Platform feature.

Where to Find It

  • Sidebar section: Design My AI
  • Sidebar label: Testing
  • URL path: /model-testing

Sub-pages:

  • /model-testing/eval/:id — Weights (re-rank interactively): Composite scores do not match axis scores × weights. Rankings below may be incorrect. Report generation and publication are blocked until re-aggregation resolves the mismatch.
  • /model-testing/eval/admin — Template Judge Config: Access denied. Super admin only.
  • /model-testing/eval/design-test — Design a Test: Describe your use case. This helps us suggest what to test.
  • /model-testing/eval/history/:id — Test Run in Progress: models • questions • rep • API calls
  • /model-testing/eval/history — Run History: All your past test runs and their results.
  • /model-testing/eval/judges/:id — Edit Judge: Instructions telling the judge how to evaluate model responses.
  • /model-testing/eval/judges/new — New Evaluator: Prompt preview: ...
  • /model-testing/eval/judges — Evaluation: Define how model responses are scored — pick an evaluator model and set your scoring criteria.
  • /model-testing/eval — Model Testing: Create test data, configure evaluation, run tests, and build intelligent routers.
  • /model-testing/eval/routers/:id — Detail Page: Router not found
  • /model-testing/eval/routers/new — Build Your Intelligent Router: Select one or more completed test runs. The router will use these scores to decide which model handles each request.
  • /model-testing/eval/routers — Your AI: Custom AI systems designed from your test results and priorities.
  • /model-testing/eval/run — Measure Models: Test models against your requirements to see which ones deliver.
  • /model-testing/eval/test-data/:id — Questions (): Templates are read-only
  • /model-testing/eval/test-data/new — New Requirements: questions generated:
  • /model-testing/eval/test-data — Requirements: Define what your AI needs to be good at — the questions models will be tested on.
  • /model-testing/judge-calibration — Judge Calibration Registry: Measured judge quality by probe testing. Badges derived from gates and sensitivity, not assumed from model size.

What You Can Do

  • Download JSON
  • ← Back
  • Cancel
  • Next →
    • Add custom capability
  • ← Cancel
  • ↓ Export JSON
  • ← Back to History
  • ← History
  • ← Try again
    • New Run
  • Start your first test run
  • Apply recommendation
  • ← Back to Judges
    • New Evaluator

How It Works

  1. Navigate to /model-testing
  2. Navigate to Weights (re-rank interactively) (/model-testing/eval/:id)
  3. Navigate to Template Judge Config (/model-testing/eval/admin)
  4. Navigate to Design a Test (/model-testing/eval/design-test)
  5. Navigate to Test Run in Progress (/model-testing/eval/history/:id)
  6. Navigate to Run History (/model-testing/eval/history)

Requirements

  • Plan: all
  • Role: admin

Common Issues

Integrity check failed — results quarantined → Check the requirements above or try again. Contact support if the issue persists.

Composite scores do not match axis scores × weights. Rankings below may be incorrect. Report generation and publication are blocked until re-aggregation resolves the mismatch. → Check the requirements above or try again. Contact support if the issue persists.

v • models • mode • Judge: • → Check the requirements above or try again. Contact support if the issue persists.

Report generation blocked → Check the requirements above or try again. Contact support if the issue persists.

⚠️ Partial coverage: not all models scored all tasks. Rankings use only commonly-scored tasks for fairness. → Check the requirements above or try again. Contact support if the issue persists.

Critical Fail > → Check the requirements above or try again. Contact support if the issue persists.

Access denied. Super admin only. → Check the requirements above or try again. Contact support if the issue persists.

Not configured → Check the requirements above or try again. Contact support if the issue persists.

Suggested based on your description — adjust as needed. → Check the requirements above or try again. Contact support if the issue persists.

Important: AI-generated ground truth introduces noise. → Check the requirements above or try again. Contact support if the issue persists.