Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
Add this skill
npx mdskills install sickn33/agent-evaluation@sickn33? Sign in with GitHub to claim this listing.Identifies critical patterns for LLM agent evaluation but lacks actionable instructions
No comments yet. Sign in to start the discussion.
Threaded comments with markdown support coming soon.