Cpu Inference skills
Free agent skills tagged cpu inference, ready to install into any SKILL.md-compatible agent.
2 skills
GGUF Quantization
davila7
Efficient model deployment for consumer hardware.
Developer ToolsintermediatePython · Shell30.2k repo
Llama.cpp
Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.
