Testing
Platform feature.
Where to Find It
- Sidebar section: Design My AI
- Sidebar label: Testing
- URL path:
/model-testing
Sub-pages:
/model-testing/eval/:id— Weights (re-rank interactively): Composite scores do not match axis scores × weights. Rankings below may be incorrect. Report generation and publication are blocked until re-aggregation resolves the mismatch./model-testing/eval/admin— Template Judge Config: Access denied. Super admin only./model-testing/eval/design-test— Design a Test: Describe your use case. This helps us suggest what to test./model-testing/eval/history/:id— Test Run in Progress: models • questions • rep • API calls/model-testing/eval/history— Run History: All your past test runs and their results./model-testing/eval/judges/:id— Edit Judge: Instructions telling the judge how to evaluate model responses./model-testing/eval/judges/new— New Evaluator: Prompt preview: .../model-testing/eval/judges— Evaluation: Define how model responses are scored — pick an evaluator model and set your scoring criteria./model-testing/eval— Model Testing: Create test data, configure evaluation, run tests, and build intelligent routers./model-testing/eval/routers/:id— Detail Page: Router not found/model-testing/eval/routers/new— Build Your Intelligent Router: Select one or more completed test runs. The router will use these scores to decide which model handles each request./model-testing/eval/routers— Your AI: Custom AI systems designed from your test results and priorities./model-testing/eval/run— Measure Models: Test models against your requirements to see which ones deliver./model-testing/eval/test-data/:id— Questions (): Templates are read-only/model-testing/eval/test-data/new— New Requirements: questions generated:/model-testing/eval/test-data— Requirements: Define what your AI needs to be good at — the questions models will be tested on./model-testing/judge-calibration— Judge Calibration Registry: Measured judge quality by probe testing. Badges derived from gates and sensitivity, not assumed from model size.
What You Can Do
- Download JSON
- ← Back
- Cancel
- Next →
- Add custom capability
- ← Cancel
- ↓ Export JSON
- ← Back to History
- ← History
- ← Try again
- New Run
- Start your first test run
- Apply recommendation
- ← Back to Judges
- New Evaluator
How It Works
- Navigate to /model-testing
- Navigate to Weights (re-rank interactively) (/model-testing/eval/:id)
- Navigate to Template Judge Config (/model-testing/eval/admin)
- Navigate to Design a Test (/model-testing/eval/design-test)
- Navigate to Test Run in Progress (/model-testing/eval/history/:id)
- Navigate to Run History (/model-testing/eval/history)
Requirements
- Plan: all
- Role: admin
Common Issues
Integrity check failed — results quarantined → Check the requirements above or try again. Contact support if the issue persists.
Composite scores do not match axis scores × weights. Rankings below may be incorrect. Report generation and publication are blocked until re-aggregation resolves the mismatch. → Check the requirements above or try again. Contact support if the issue persists.
v • models • mode • Judge: • → Check the requirements above or try again. Contact support if the issue persists.
Report generation blocked → Check the requirements above or try again. Contact support if the issue persists.
⚠️ Partial coverage: not all models scored all tasks. Rankings use only commonly-scored tasks for fairness. → Check the requirements above or try again. Contact support if the issue persists.
Critical Fail > → Check the requirements above or try again. Contact support if the issue persists.
Access denied. Super admin only. → Check the requirements above or try again. Contact support if the issue persists.
Not configured → Check the requirements above or try again. Contact support if the issue persists.
Suggested based on your description — adjust as needed. → Check the requirements above or try again. Contact support if the issue persists.
Important: AI-generated ground truth introduces noise. → Check the requirements above or try again. Contact support if the issue persists.