
Ablation Planner
FreeDesign ablation studies efficiently for research papers.
Free · Opens the source repo
What Ablation Planner does
The Ablation Planner is a specialized tool designed for researchers and developers who need to conduct ablation studies as part of their machine learning experiments. It is particularly useful when main results pass the result-to-claim check, indicating that claims are supported or partially supported. This skill automates the process of designing ablation studies by leveraging an AI agent to generate a structured plan that addresses the critical questions reviewers typically ask. By doing so, it helps researchers ensure their findings are robust and well-supported by empirical evidence.
The workflow begins with preparing the context, where the tool reads through available project files to gather essential information about the method, experiment results, claims, and compute resources. This comprehensive overview allows the AI agent to design relevant ablations that isolate the contributions of individual components, test sensitivity to hyperparameters, and compare alternative design choices. Each proposed ablation is detailed, specifying what it tests and the expected outcomes, ensuring that every experiment has a clear purpose.
Once the ablation plan is generated, the local executor reviews its feasibility, checking compute budgets and dependencies before implementation. This step is crucial to ensure that the proposed studies can be executed within available resources. The final implementation phase involves creating configurations for each ablation, running them in a prioritized order, and tracking results meticulously. This structured approach not only streamlines the ablation process but also enhances the quality of the research output by ensuring that all relevant questions are addressed systematically.
Overall, the Ablation Planner is ideal for researchers looking to strengthen their experimental designs and provide thorough responses to reviewer inquiries, thereby increasing the likelihood of successful paper submissions.
When to use it
Use this tool when your main results have been validated and you need to design ablation studies for paper submission or when prompted by an auto-review loop.
When not to use it
Avoid using this skill if your results have not been validated or if ablation studies are not relevant to your current research phase.
What you can build with it
Preparing for Paper Submission
When preparing your research for submission, use the Ablation Planner to ensure all necessary ablation studies are designed and documented.
Addressing Reviewer Feedback
If reviewers request additional evidence or clarity on your model's performance, utilize this tool to quickly generate relevant ablation studies.
Optimizing Experimental Design
Use the skill to refine your experimental design by systematically testing the impact of different components and configurations on your model's performance.
How to install Ablation Planner
View source1. Install with the skills CLI
npx skills add wanshuiyin/auto-claude-code-research-in-sleep/ablation-planner --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by wanshuiyinAblation Planner
Systematically design ablation studies that answer the questions reviewers will ask. The reviewer agent leads the design; the local executor reviews feasibility and implements.
Context: $ARGUMENTS
When to Use
- Main results pass
/result-to-claimwithclaim_supported = yesorpartial - The user explicitly requests ablation planning
/auto-review-loopidentifies missing ablations
Workflow
Step 1: Prepare Context
Read available project files to build the full picture:
- Method description and components (from
idea-stage/docs/research_contract.md, legacydocs/research_contract.md, project notes, or method docs) - Current experiment results (from
EXPERIMENT_LOG.md,EXPERIMENT_TRACKER.md, or W&B) - Confirmed and intended claims (from
/result-to-claimoutput or project notes) - Available compute resources (from server notes, run configs, or user-provided budget)
Step 2: Codex Designs Ablations
spawn_agent:
model: gpt-5.6-sol
reasoning_effort: xhigh
message: |
You are a rigorous ML reviewer planning ablation studies.
Given this method and results, design ablations that:
1. Isolate the contribution of each novel component
2. Answer questions reviewers will definitely ask
3. Test sensitivity to key hyperparameters
4. Compare against natural alternative design choices
Method: [description from project files]
Components: [list of removable or replaceable components]
Current results: [key metrics from experiments]
Claims: [what we claim and current evidence]
For each ablation, specify:
- name: what to change (for example, "remove module X", "replace Y with Z")
- what_it_tests: the specific question this answers
- expected_if_component_matters: what we predict if the component is important
- priority: 1 (must-run) to 5 (nice-to-have)
Also provide:
- coverage_assessment: what reviewer questions these ablations answer
- unnecessary_ablations: experiments that seem useful but will not add insight
- suggested_order: run order optimized for maximum early information
- estimated_compute: total GPU-hours estimate
If delegation is unavailable, generate the same plan locally and mark it [pending external review].
Step 3: Parse Ablation Plan
Normalize the response into a structured format:
## Ablation Plan
### Component Ablations (highest priority)
| # | Name | What It Tests | Expected If Matters | Priority |
|---|------|---------------|---------------------|----------|
| 1 | remove module X | contribution of X | performance drops on metric Y | 1 |
| 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |
### Hyperparameter Sensitivity
| # | Parameter | Values to Test | What It Tests | Priority |
|---|-----------|----------------|---------------|----------|
| 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |
### Design Choice Comparisons
| # | Name | What It Tests | Priority |
|---|------|---------------|----------|
| 4 | joint vs separate matching | whether joint adds value | 4 |
### Coverage Assessment
[What reviewer questions these ablations answer]
### Unnecessary Ablations
[Experiments that seem useful but will not add insight - skip these]
### Run Order
[Optimized for maximum early information]
### Estimated Compute
[Total GPU-hours]
Step 4: CC Reviews Feasibility
Before running anything, the local executor checks:
- Compute budget - Can you afford all ablations with available GPUs?
- Code changes - Which ablations need code modifications vs config-only changes?
- Dependencies - Which ablations can run in parallel?
- Cuts - If budget is tight, propose removing lower-priority ablations and ask the reviewer agent to re-prioritize when possible
Step 5: Implement and Run
- Create configs or scripts for each ablation (config-only changes first)
- Smoke test each ablation before the full run
- Run in the suggested order, using descriptive names (for example,
ablation-no-module-X) - Track results in
EXPERIMENT_LOG.md - After all ablations complete, update
findings.mdwith insights
Rules
- The reviewer agent leads the design. Do not pre-filter or bias the ablation list before external review sees it. The reviewer thinks like a reviewer; the local executor thinks like an engineer.
- Every ablation must have a clear
what_it_testsandexpected_if_component_matters. No "just try it" experiments. - Config-only ablations take priority over those needing code changes (faster, less error-prone).
- If total compute exceeds budget, propose cuts and ask for re-prioritization - do not silently drop ablations.
- Component ablations (remove or replace) take priority over hyperparameter sweeps.
- Do not generate ablations for components identical to the baseline (no-op ablations).
- Record all ablation results in
EXPERIMENT_LOG.md, including negative results (for example, component removal had no effect).
Frequently asked questions about Ablation Planner
Similar skills
Scientific Problem Selection
Streamline your research problem selection process.
Nextflow Development
Run nf-core bioinformatics pipelines with ease.
Nature Reviewer Assessment
Simulate peer review for scientific manuscripts.
Research Writing Pipeline
Streamline your scientific writing with structured proposal-first methodologies.
Nature Literature Downloader
Efficiently download academic literature from various sources.
Auto Research
Streamline your NeMo-RL experiments with automated workflows.
