Quantization skills
Free agent skills tagged quantization, ready to install into any SKILL.md-compatible agent.
13 skills
High-Performance LLM Serving
davila7
Optimize LLM API deployment with high throughput and low latency.
DINO Object Detection
nvidia
Train and deploy 2D object detection models with DINO.
RT-DETR
nvidia
Real-time object detection with competitive accuracy.
vLLM Serving
nousresearch
High-performance LLM serving for optimized inference.
GPTQ Quantization
davila7
Efficiently deploy large LLMs on consumer GPUs with minimal loss.
AWQ Quantization
davila7
Optimize large models with 4-bit quantization.
TensorRT-LLM
davila7
Optimize LLM inference for NVIDIA GPUs with TensorRT.
BitsAndBytes Quantization
davila7
Optimize LLMs with efficient quantization techniques.
GGUF Quantization
davila7
Efficient model deployment for consumer hardware.
Half-Quadratic Quantization
davila7
Efficient model quantization without calibration data.
Vector Search
ruvnet
Efficient vector search for large and small datasets.
TensorRT-LLM
nousresearch
Optimize LLM inference for NVIDIA GPUs.
AgentDB Performance Optimization
ruvnet
Enhance AgentDB efficiency with advanced techniques.
