start-evals
Start AI evals without overengineering. Create your first 20 test cases in a spreadsheet using PM-Friendly Evals approach.
What this skill does
# Start Evals
Launch your AI evaluation process using the **PM-Friendly Evals approach** (Aman Khan + Hamel Husain).
Start with 20 test cases in a spreadsheet. Scale when ready. Error analysis > automation.
## Entry Point
When this skill is invoked, start with:
```
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
START EVALS
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Start with 20 test cases. Scale when ready.
What AI feature are you evaluating?
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
```
## Usage
```
/start-evals [feature-name]
```
**Examples:**
- `/start-evals "AI product recommendations"` - Generate test cases
- `/start-evals --create-project` - Create Linear project for tracking
- `/start-evals "customer support AI" --count 50` - Generate 50 test cases
## What Happens
1. Invokes the **eval-generator** agent
2. Asks about your AI feature and quality criteria
3. Generates 20 test cases (15 happy path + 5 edge cases)
4. Provides spreadsheet template and workflow
5. Optionally creates Linear project for tracking
## The Philosophy
**Good -> Better -> Best progression:**
| Stage | Test Cases | Process | Tool |
|-------|------------|---------|------|
| **Good** (Week 1) | 20 | Manual review | Spreadsheet |
| **Better** (Month 1-2) | 50-100 | LLM-as-judge | Weekly reviews |
| **Best** (Month 3+) | 200+ | Automated | CI/CD integration |
**Start here.** You're at "Good." Don't jump to automation.
## What You'll Get
```
AI Evals Starter Kit: Product Recommendations
HAPPY PATH (15 cases):
1. Input: "Recommend a laptop under $800 for college"
Expected: Mid-range laptops with student-friendly features, under budget
Pass criteria: All recommendations < $800, suitable for students
2. Input: "Best phone for photography"
Expected: High-end phones with excellent cameras
Pass criteria: Focus on camera quality, not price
...
EDGE CASES (5 cases):
16. Input: "Phone for elderly person"
Expected: Simple, large screen, easy to use
Pass criteria: Prioritizes simplicity over features
Why it's tricky: Must understand implicit needs
...
```
## Week 1 Workflow (2-3 hours)
1. Copy test cases to spreadsheet (10 min)
2. Run your AI against each input (1-2 hours)
3. Record actual outputs
4. Mark pass/fail
5. Look for patterns in failures (30 min)
## After 1-2 Weeks
| Pass Rate | Action |
|-----------|--------|
| 80%+ | Add 10 more test cases |
| <80% | Fix issues, rerun |
| 50-100 cases | Graduate to "Better" approach |
## Common Questions
**Q: 20 seems like too few. Should I start with 100?**
A: No. 20 cases covering your core use case > 100 cases you never run.
**Q: How long does running 20 tests take?**
A: First time: 30-60 min. After that: 15-20 min per run.
**Q: Do I need special tools?**
A: No. Spreadsheet works great. Graduate to tools when manual gets painful.
## Ready to Scale?
| Signal | Next Step |
|--------|-----------|
| You have 50+ test cases or see production failures | `/upgrade-evals` — Systematic error analysis on real traces |
| You need more diverse test inputs | `/generate-test-data` — Dimension-based synthetic data |
| Your AI feature uses retrieval (search, knowledge base) | `/eval-rag` — Separate retrieval from generation evaluation |
## Related Commands
- `/upgrade-evals` - Error analysis on real traces (next step after this)
- `/build-judge` - LLM-as-Judge for subjective failure modes
- `/generate-test-data` - Diverse synthetic test inputs
- `/eval-rag` - RAG-specific retrieval + generation evaluation
- `/calibrate` - Ongoing post-launch calibration
- `/ai-health-check` - Full pre-launch readiness audit
- `/ai-cost-check` - Economic validation
---
**Framework:** PM-Friendly Evals (Aman Khan + Hamel Husain)
**Key insight:** "Error analysis is the most important activity. Start with 20 cases in a spreadsheet."
Related in Data & Analytics
clawarr-suite
IncludedComprehensive management for self-hosted media stacks (Sonarr, Radarr, Lidarr, Readarr, Prowlarr, Bazarr, Overseerr, Plex, Tautulli, SABnzbd, Recyclarr, Unpackerr, Notifiarr, Maintainerr, Kometa, FlareSolverr). Deep library exploration, analytics, dashboard generation, content management, request handling, subtitle management, indexer control, download monitoring, quality profile sync, library cleanup automation, notification routing, collection/overlay management, and media tracker integration (Trakt, Letterboxd, Simkl).
querying-soql
IncludedSOQL query generation, optimization, and analysis with 100-point scoring. Use this skill when the user needs SOQL/SOSL authoring or optimization: natural-language-to-query generation, relationship queries, aggregates, query-plan analysis, and performance or safety improvements for Salesforce queries. TRIGGER when: user writes, optimizes, or debugs SOQL/SOSL queries, touches .soql files, or asks about relationship queries, aggregates, or query performance. DO NOT TRIGGER when: bulk data operations (use handling-sf-data), Apex DML logic (use generating-apex), or report/dashboard queries.
app-store-optimization
IncludedApp Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklists, and tracking ranking changes.
habit-flow
IncludedAI-powered atomic habit tracker with natural language logging, streak tracking, smart reminders, and coaching. Use for creating habits, logging completions naturally ("I meditated today"), viewing progress, and getting personalized coaching.
app-store-optimization
IncludedApp Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklists, and tracking ranking changes.
visualizing-data
IncludedBuilds dashboards, reports, and data-driven interfaces requiring charts, graphs, or visual analytics. Provides systematic framework for selecting appropriate visualizations based on data characteristics and analytical purpose. Includes 24+ visualization types organized by purpose (trends, comparisons, distributions, relationships, flows, hierarchies, geospatial), accessibility patterns (WCAG 2.1 AA compliance), colorblind-safe palettes, and performance optimization strategies. Use when creating visualizations, choosing chart types, displaying data graphically, or designing data interfaces.