Mechanistic Interpretability skills
Free agent skills tagged mechanistic interpretability, ready to install into any SKILL.md-compatible agent.
3 skills
Sparse Autoencoder Training
davila7
Train and analyze Sparse Autoencoders for interpretable features.
Developer ToolsintermediatePython · Shell30.2k repo
Nnsight Remote Interpretability
davila7
Access and manipulate neural network internals seamlessly.
Developer ToolsintermediatePython · Shell30.2k repo
TransformerLens Interpretability
davila7
Explore and manipulate transformer model internals.
Research & EducationadvancedPython · Shell30.2k repo
O
Obliteratus
nousresearch
Remove refusal behaviors from LLMs without retraining.
AI & AgentsadvancedPython · Shell228.5k repo
