Evaluation & Testing¶
Skills for evaluating and testing AI agent skills
SDLC¶
assess-rfe¶
Assess RFEs against quality criteria using a structured rubric.
2 skills - v1.0.0
assess-strat¶
Assess RHAISTRAT strategies against quality criteria using a scored rubric with calibration examples. Scores across four dimensions: feasibility, testability, scope, and architecture.
2 skills - v1.0.0
test-plan¶
End-to-end test planning workflow for RHOAI: generate E2E/UI-focused test plans from Jira strategies, create traceable test cases, implement executable automation code, verify UI tests against live clusters via Playwright, publish to GitHub, resolve review feedback, and score plans with deterministic evidence gates and automated rubrics.
16 skills - v2.0.0
quality-tooling¶
Quality tooling and automation for RHOAI component development. Includes automated repository analysis, build validation, and test pattern extraction.
5 skills - v1.0.0
Generic¶
agent-eval-harness¶
Generic agentic evaluation for skills and agents. Provides end-to-end skills to analyze, test, score, review, and iteratively improve agent skills, plus compare models/configurations and run Design-of-Experiments (ANOVA) sweeps. MLflow support for experiment tracking, tracing, and reporting. Schema-driven evaluation via eval.yaml with support for inline, LLM-based, and external judges.
10 skills - v1.30.0