AI Skills Catalog
Business

Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.

VerifiedFreeMITv1.0.0
SKILL SCORE71Engagement score

Works with

ClClaudeNative
CoCodexPackaged
GPGPTPackaged
GeGeminiPackaged
CuCursorPackaged
OpOpenCodePackaged

Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.

Optimized workflow

This edition turns the source methodology into a repeatable agent workflow with explicit inputs, checkpoints and deliverables.

Quality standard

  • Confirm scope and missing inputs before execution
  • Ground decisions in available evidence and preserve source constraints
  • Return an actionable result with assumptions, risks and next steps

Agent compatibility

The same core method is packaged for Claude, Codex, GPT, Gemini, Cursor and OpenCode.

Permissions & security

Low risk

Source verified · conversion tested · security signals reviewed

Read project files
Write generated skill files