Inference Serving skills
Free agent skills tagged inference serving, ready to install into any SKILL.md-compatible agent.
3 skills
High-Performance LLM Serving
davila7
Optimize LLM API deployment with high throughput and low latency.
Developer ToolsintermediatePython · Shell30.2k repo
SGLang
davila7
Fast structured generation and serving for LLMs.
AI & AgentsintermediatePython · Shell30.2k repo
TensorRT-LLM
davila7
Optimize LLM inference for NVIDIA GPUs with TensorRT.
Developer ToolsadvancedPython · Shell30.2k repo
