Realistic AI engineering scenarios with reference solutions and problem-specific grading rubrics.
Tool calling · Irreversible actions · Policy enforcement · Escalation
Tool selection · Retrieval · Permissions · Failure handling · Observability
Code Context · Low Latency · Fill-in-the-Middle · Repo Indexing · Acceptance Metrics
Agents · Tool Use · Escalation · Safety · Containment Rate