
Monitor Experiment Results
FreeEfficiently track and summarize your experimental outputs.
Free · Opens the source repo
What Monitor Experiment Results does
The Monitor Experiment Results skill is designed for developers and researchers who need to keep track of ongoing experiments and analyze their outcomes. By leveraging SSH commands and integration with platforms like Vast.ai and Weights & Biases, this skill allows users to check the status of running jobs, collect output logs, and retrieve performance metrics in a streamlined manner. This is particularly useful in environments where multiple experiments are conducted simultaneously, and timely insights are crucial.
The workflow begins with checking what experiments are currently running on specified servers or cloud instances. By utilizing SSH commands, users can list active screen sessions or check the status of their Vast.ai instances. Once the running experiments are identified, the skill proceeds to collect the output from each screen session, ensuring that the most recent logs are captured for further analysis. This step is essential for understanding the current state of each experiment and identifying any potential issues early on.
In addition to capturing logs, the skill also checks for JSON result files, which often contain structured data about the experiment's performance. If available, these files are fetched and parsed to extract key metrics. For users utilizing Weights & Biases, the skill can pull detailed training curves and metrics, providing deeper insights into the training dynamics of machine learning models. This comprehensive data collection allows for a thorough comparison of results against known baselines, helping users make informed decisions about their next steps.
Finally, the skill summarizes the collected results in a clear table format, highlighting key metrics and any significant changes compared to baseline experiments. This structured output not only aids in quick assessments but also facilitates communication with team members or stakeholders, especially when integrated with notification systems like Feishu. Overall, this skill is an invaluable tool for anyone involved in experimental research or development, providing a robust framework for monitoring and interpreting results efficiently.
When to use it
Use this skill when you need to monitor the status of experiments, collect output, and analyze results in a structured manner.
When not to use it
This skill is not suitable for real-time monitoring of experiments that require immediate human intervention or for qualitative assessments of results.
What you can build with it
Monitoring Ongoing Experiments
Use this skill to check the status of multiple running experiments on your server, ensuring you have the latest updates.
Collecting Experiment Outputs
Automatically gather logs and results from completed experiments, saving time and reducing manual effort.
Analyzing Performance Metrics
Retrieve and compare key metrics from your experiments, helping to identify trends and inform future research directions.
How to install Monitor Experiment Results
View source1. Install with the skills CLI
npx skills add wanshuiyin/auto-claude-code-research-in-sleep/monitor-experiment --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by wanshuiyinMonitor Experiment Results
⏱ External cadence is appropriate here. This skill waits on an external fact (job completion / progress), so it is a natural
/loop/CronCreatesurface: the wake reads status and self-judges only machine-checkable completion (exit code, file exists, epoch logged) — never quality. This is the additive external-wait shape inshared-references/external-cadence.md. If a scheduled wait here ends in a verdict step (e.g. then audit results), run that verdict once after the wait clears — not re-entered per tick.
Monitor: $ARGUMENTS
Workflow
Step 1: Check What's Running
SSH server:
ssh <server> "screen -ls"
Vast.ai instance (read ssh_host, ssh_port from vast-instances.json):
ssh -p <PORT> root@<HOST> "screen -ls"
Also check vast.ai instance status:
vastai show instances
Modal (when gpu: modal in CLAUDE.md):
modal app list # List running/recent apps
modal app logs <app> # Stream logs from a running app
Modal apps auto-terminate when done — if it's not in the list, it already finished. Check results via modal volume ls <volume> or local output.
Step 2: Collect Output from Each Screen
For each screen session, capture the last N lines:
ssh <server> "screen -S <name> -X hardcopy /tmp/screen_<name>.txt && tail -50 /tmp/screen_<name>.txt"
If hardcopy fails, check for log files or tee output.
Step 3: Check for JSON Result Files
ssh <server> "ls -lt <results_dir>/*.json 2>/dev/null | head -20"
If JSON results exist, fetch and parse them:
ssh <server> "cat <results_dir>/<latest>.json"
Step 3.5: Pull W&B Metrics (when wandb: true in CLAUDE.md)
Skip this step entirely if wandb is not set or is false in CLAUDE.md.
Pull training curves and metrics from Weights & Biases via Python API:
# List recent runs in the project
ssh <server> "python3 -c \"
import wandb
api = wandb.Api()
runs = api.runs('<entity>/<project>', per_page=10)
for r in runs:
print(f'{r.id} {r.state} {r.name} {r.summary.get(\"eval/loss\", \"N/A\")}')
\""
# Pull specific metrics from a run (last 50 steps)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
history = list(run.scan_history(keys=['train/loss', 'eval/loss', 'eval/ppl', 'train/lr'], page_size=50))
print(json.dumps(history[-10:], indent=2))
\""
# Pull run summary (final metrics)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
print(json.dumps(dict(run.summary), indent=2, default=str))
\""
What to extract:
- Training loss curve — is it converging? diverging? plateauing?
- Eval metrics — loss, PPL, accuracy at latest checkpoint
- Learning rate — is the schedule behaving as expected?
- GPU memory — any OOM risk?
- Run status — running / finished / crashed?
W&B dashboard link (include in summary for user):
https://wandb.ai/<entity>/<project>/runs/<run_id>
This gives the auto-review-loop richer signal than just screen output — training dynamics, loss curves, and metric trends over time.
Step 4: Summarize Results
Present results in a comparison table:
| Experiment | Metric | Delta vs Baseline | Status |
|-----------|--------|-------------------|--------|
| Baseline | X.XX | — | done |
| Method A | X.XX | +Y.Y | done |
Step 5: Interpret
- Compare against known baselines
- Flag unexpected results (negative delta, NaN, divergence)
- Suggest next steps based on findings
Step 6: Feishu Notification (if configured)
After results are collected, check ~/.claude/feishu.json:
- Send
experiment_donenotification: results summary table, delta vs baseline - If config absent or mode
"off": skip entirely (no-op)
Key Rules
- Always show raw numbers before interpretation
- Compare against the correct baseline (same config)
- Note if experiments are still running (check progress bars, iteration counts)
- If results look wrong, check training logs for errors before concluding
- Vast.ai cost awareness: When monitoring vast.ai instances, report the running cost (hours * $/hr from
vast-instances.json). If all experiments on an instance are done, remind the user to run/vast-gpu destroy <instance_id>to stop billing - Modal cost awareness: Modal auto-scales to zero — no idle billing. When reporting results from Modal runs, note the actual execution time and estimated cost (time * $/hr from the GPU tier used). No cleanup action needed
Frequently asked questions about Monitor Experiment Results
Similar skills
Spring Boot Testing
Master testing techniques for Spring Boot 4 applications.
GitHub Issues
Manage GitHub issues efficiently with MCP tools.
Geofeed Tuner
Optimize your IP geolocation feeds in CSV format.
Batch Files
Master Windows batch scripting for automation and task management.
Adobe Illustrator Scripting
Automate your Illustrator workflows with ExtendScript.
Plugin Structure
Create and organize Claude Code plugins effectively.
