Performance skills
Free performance skills for AI coding agents, part of our developer tools collection.
12 skills
Half-Quadratic Quantization
davila7
Efficient model quantization without calibration data.
GPTQ Quantization
davila7
Efficiently deploy large LLMs on consumer GPUs with minimal loss.
GGUF Quantization
davila7
Efficient model deployment for consumer hardware.
Flash Attention Optimization
davila7
Enhance transformer performance with optimized attention.
BitsAndBytes Quantization
davila7
Optimize LLMs with efficient quantization techniques.
AWQ Quantization
davila7
Optimize large models with 4-bit quantization.
TensorBoard
davila7
Visualize and debug machine learning models effectively.
TensorRT-LLM
davila7
Optimize LLM inference for NVIDIA GPUs with TensorRT.
Llama.cpp
davila7
Efficient LLM inference on non-NVIDIA hardware.
Speculative Decoding
davila7
Accelerate LLM inference with advanced techniques.
Model Pruning
davila7
Compress LLMs and speed up inference with pruning techniques.
Performance Profiling
davila7
Measure, analyze, and optimize your web performance.
