
Behavioral X-Ray
FreeProbe AI models for hidden behavioral patterns.
Free · Opens the source repo
What Behavioral X-Ray does
Behavioral X-Ray is a powerful tool designed to systematically analyze the behavioral patterns of AI models. By running 30 tailored probe questions across six key dimensions, this skill generates a comprehensive visual report that includes radar charts and actionable insights. It enables users to understand how their AI models truly behave, revealing critical information such as refusal boundaries, hallucination tendencies, and reasoning styles. The skill operates without the need for API keys or external setup, allowing for immediate deployment and testing.
The probing process is straightforward. Users can initiate a full behavioral probe or focus on specific dimensions like refusal or formatting. The results are auto-tagged with behavioral metadata, providing a clear overview of the model's performance in various scenarios. This is especially useful for developers and designers who need to select the right model for specific tasks, debug unexpected behaviors, or conduct compliance audits.
The output is a styled HTML report that breaks down the model's behavior into easily digestible formats, including refusal rates and notable response examples. This allows users to track behavioral drift over time and make informed decisions about model selection and deployment. By utilizing this skill, teams can ensure that they are deploying AI models that meet their operational needs and compliance requirements while minimizing risks associated with unexpected AI behavior.
Behavioral X-Ray is ideal for anyone involved in AI development, testing, or deployment, including developers, compliance officers, and red team assessors. It provides a systematic approach to understanding AI behavior, which is crucial in today's landscape of increasingly complex AI systems.
When to use it
Use this skill when you need to understand an AI model's behavior, compare models for specific tasks, or track behavioral changes over time.
When not to use it
This skill is not suitable for tasks outside its scope, such as real-time model performance monitoring or environment-specific validation.
What you can build with it
Understanding Model Behavior
Use Behavioral X-Ray to gain insights into how your AI model behaves under various conditions, helping you refine your prompts.
Compliance Auditing
Document and assess AI behavior at deployment boundaries to ensure compliance with safety and operational standards.
Red Team Assessments
Conduct systematic evaluations of AI models to identify potential vulnerabilities and safety boundary issues.
How to install Behavioral X-Ray
View source1. Install with the skills CLI
npx skills add sickn33/agentic-awesome-skills/bdistill-behavioral-xray --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by sickn33Behavioral X-Ray
Systematically probe an AI model's behavioral patterns and generate a visual report. The AI agent probes itself — no API key or external setup needed.
Overview
bdistill's Behavioral X-Ray runs 30 carefully designed probe questions across 6 dimensions, auto-tags each response with behavioral metadata, and compiles results into a styled HTML report with radar charts and actionable insights.
Use it to understand your model before building with it, compare models for task selection, or track behavioral drift over time.
When to Use This Skill
- Use when you want to understand how your AI model actually behaves (not how it claims to)
- Use when choosing between models for a specific task
- Use when debugging unexpected refusals, hallucinations, or formatting issues
- Use for compliance auditing — documenting model behavior at deployment boundaries
- Use for red team assessments — systematic boundary mapping across safety dimensions
How It Works
Step 1: Install
pip install bdistill
claude mcp add bdistill -- bdistill-mcp # Claude Code
For other tools, add bdistill-mcp as an MCP server in your project config.
Step 2: Run the probe
In Claude Code:
/xray # Full behavioral probe (30 questions)
/xray --dimensions refusal # Probe just one dimension
/xray-report # Generate report from completed probe
In any tool with MCP:
"X-ray your behavioral patterns"
"Test your refusal boundaries"
"Generate a behavioral report"
Probe Dimensions
| Dimension | What it measures |
|---|---|
| tool_use | When does it call tools vs. answer from knowledge? |
| refusal | Where does it draw safety boundaries? Does it over-refuse? |
| formatting | Lists vs. prose? Code blocks? Length calibration? |
| reasoning | Does it show chain-of-thought? Handle trick questions? |
| persona | Identity, tone matching, composure under hostility |
| grounding | Hallucination resistance, fabrication traps, knowledge limits |
Output
A styled HTML report showing:
- Refusal rate, hedge rate, chain-of-thought usage
- Per-dimension breakdown with bar charts
- Notable response examples with behavioral tags
- Actionable insights (e.g., "you already show CoT 85% of the time, no need to prompt for it")
Best Practices
- Answer probe questions honestly — the value is in authentic behavioral data
- Run probes on the same model periodically to track behavioral drift
- Compare reports across models to make informed selection decisions
- Use adversarial knowledge extraction (
/distill --adversarial) alongside behavioral probes for complete model profiling
Related Skills
@bdistill-knowledge-extraction- Extract structured domain knowledge from any AI model
Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
Frequently asked questions about Behavioral X-Ray
Similar skills
Arize Evaluator
Streamline LLM evaluation workflows on Arize.
Troubleshoot
Analyze logs to understand chat agent behavior.
Agentic Evaluation
Enhance AI outputs through iterative evaluation and refinement.
RAG Evaluation
Evaluate retrieval-augmented generation benchmarks efficiently.
NV-Reason-CXR
Run smoke tests for chest X-ray reasoning models.
Clinical ASR Evaluation
Score and evaluate clinical ASR manifests effectively.
