AI Skills Catalog
ai-ml

Advanced Evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines

VerifiedFreeMITv1.0.0
SKILL SCORE74Engagement score

Works with

ClClaudeNative
CoCodexPackaged
GPGPTPackaged
GeGeminiPackaged
CuCursorPackaged
OpOpenCodePackaged

Advanced Evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

Optimized workflow

This edition turns the source methodology into a repeatable agent workflow with explicit inputs, checkpoints and deliverables.

Quality standard

  • Confirm scope and missing inputs before execution
  • Ground decisions in available evidence and preserve source constraints
  • Return an actionable result with assumptions, risks and next steps

Agent compatibility

The same core method is packaged for Claude, Codex, GPT, Gemini, Cursor and OpenCode.

Permissions & security

Low risk

Source verified · conversion tested · security signals reviewed

Read project files
Write generated skill files