
Agentic Engineering
FreeOptimize AI-driven development workflows with precision.
Free · Opens the source repo
What Agentic Engineering does
Agentic Engineering is a skill designed for managing development workflows where AI agents take on the bulk of implementation tasks, while human engineers oversee quality and risk management. This approach leverages a structured methodology that emphasizes defining clear completion criteria, breaking down tasks into manageable units, and routing model tiers based on the complexity of the task at hand. The skill is particularly useful for teams looking to enhance their efficiency and effectiveness in AI-assisted software development.
The core operating principles of this skill include the eval-first execution model, which allows teams to define evaluation criteria before implementation begins. By capturing baseline performance and failure signatures, developers can ensure that subsequent implementations meet the desired outcomes. The task decomposition strategy further enhances this process by adhering to the 15-minute unit rule, which promotes the creation of independently verifiable tasks that expose clear risks and completion conditions.
Agentic Engineering also emphasizes the importance of model routing, directing the AI's capabilities according to the complexity of the task. This ensures that simpler tasks are handled efficiently while reserving more advanced models for complex implementations. Additionally, the skill introduces a cost discipline framework, allowing teams to track various metrics associated with each task, such as time taken and token usage. This approach not only aids in resource management but also provides insights into the efficiency of the AI's performance.
This skill is ideal for software development teams that are integrating AI into their workflows and require a structured approach to manage the interplay between human oversight and automated implementation. By utilizing Agentic Engineering, teams can streamline their processes, reduce risks, and improve overall code quality.
When to use it
Use this skill when implementing AI agents in software development to ensure quality control and effective task management.
When not to use it
Avoid this skill in workflows that do not involve AI agents or where human implementation is predominant without AI assistance.
What you can build with it
Managing AI Development Workflows
Utilize this skill to oversee AI-driven development processes, ensuring that human engineers maintain quality and risk controls.
Optimizing Task Decomposition
Apply the 15-minute unit rule to break down complex tasks into manageable units, improving clarity and risk assessment.
Reviewing AI-Generated Code
Use the skill's review checklist to prioritize critical aspects of AI-generated code, ensuring comprehensive quality assurance.
How to install Agentic Engineering
View source1. Install with the skills CLI
npx skills add affaan-m/ecc/agentic-engineering --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by affaan-mAgentic Engineering
Use this skill for engineering workflows where AI agents perform most implementation work and humans enforce quality and risk controls.
Operating Principles
- Define completion criteria before execution.
- Decompose work into agent-sized units.
- Route model tiers by task complexity.
- Measure with evals and regression checks.
Eval-First Loop
- Define capability eval and regression eval.
- Run baseline and capture failure signatures.
- Execute implementation.
- Re-run evals and compare deltas.
Example workflow:
1. Write test that captures desired behavior (eval)
2. Run test → capture baseline failures
3. Implement feature
4. Re-run test → verify improvements
5. Check for regressions in other tests
Task Decomposition
Apply the 15-minute unit rule:
- Each unit should be independently verifiable
- Each unit should have a single dominant risk
- Each unit should expose a clear done condition
Good decomposition:
Task: Add user authentication
├─ Unit 1: Add password hashing (15 min, security risk)
├─ Unit 2: Create login endpoint (15 min, API contract risk)
├─ Unit 3: Add session management (15 min, state risk)
└─ Unit 4: Protect routes with middleware (15 min, auth logic risk)
Bad decomposition:
Task: Add user authentication (2 hours, multiple risks)
Model Routing
Choose model tier based on task complexity:
-
Haiku: Classification, boilerplate transforms, narrow edits
- Example: Rename variable, add type annotation, format code
-
Sonnet: Implementation and refactors
- Example: Implement feature, refactor module, write tests
-
Opus: Architecture, root-cause analysis, multi-file invariants
- Example: Design system, debug complex issue, review architecture
Cost discipline: Escalate model tier only when lower tier fails with a clear reasoning gap.
Session Strategy
-
Continue session for closely-coupled units
- Example: Implementing related functions in same module
-
Start fresh session after major phase transitions
- Example: Moving from implementation to testing
-
Compact after milestone completion, not during active debugging
- Example: After feature complete, before starting next feature
Review Focus for AI-Generated Code
Prioritize:
- Invariants and edge cases
- Error boundaries
- Security and auth assumptions
- Hidden coupling and rollout risk
Do not waste review cycles on style-only disagreements when automated format/lint already enforce style.
Review checklist:
- Edge cases handled (null, empty, boundary values)
- Error handling comprehensive
- Security assumptions validated
- No hidden coupling between modules
- Rollout risk assessed (breaking changes, migrations)
Cost Discipline
Track per task:
- Model tier used
- Token estimate
- Retries needed
- Wall-clock time
- Success/failure outcome
Example tracking:
Task: Implement user login
Model: Sonnet
Tokens: ~5k input, ~2k output
Retries: 1 (initial implementation had auth bug)
Time: 8 minutes
Outcome: Success
When to Use This Skill
- Managing AI-driven development workflows
- Planning agent task decomposition
- Optimizing model tier selection
- Implementing eval-first development
- Reviewing AI-generated code
- Tracking development costs
Integration with Other Skills
- tdd-workflow: Combine with eval-first loop for test-driven development
- verification-loop: Use for continuous validation during implementation
- search-first: Apply before implementation to find existing solutions
- coding-standards: Reference during code review phase
Frequently asked questions about Agentic Engineering
Similar skills
Quality Playbook Generator
Run comprehensive quality audits on any codebase.
PR Draft Summary
Automate PR summary generation for openai-agents-python.
Final Release Review
Streamline your release candidate audits with ease.
Unit Test Vue Pinia
Efficiently write and review unit tests for Vue 3 applications.
Slang Shader Expert
Optimize and integrate Slang shaders with ease.
Telemetry Standards
Ensure consistent event tracking in Supabase Studio.
