
TensorRT-LLM
FreeOptimize LLM inference for NVIDIA GPUs with TensorRT.
Free · Opens the source repo
What TensorRT-LLM does
TensorRT-LLM is an open-source library designed to enhance the performance of large language model (LLM) inference on NVIDIA GPUs. By leveraging NVIDIA's TensorRT technology, this tool aims to provide significant improvements in throughput and latency, making it suitable for production environments where efficiency is critical. It supports advanced features such as model quantization (FP8, INT4) and in-flight batching, enabling users to achieve up to 100 times faster inference compared to traditional frameworks like PyTorch.
This skill is particularly beneficial for developers and data scientists who deploy LLMs on NVIDIA hardware, such as the A100 or H100 GPUs. With its ability to handle high throughput—up to 24,000 tokens per second for models like Llama 3—and low latency suitable for real-time applications, TensorRT-LLM addresses the growing demand for efficient AI model serving. The library also supports multi-GPU scaling, allowing users to distribute workloads across multiple GPUs or nodes, which is essential for large-scale deployments.
For those looking to implement state-of-the-art performance optimizations, TensorRT-LLM includes features like dynamic in-flight batching, efficient memory management through a paged key-value cache, and optimized attention kernels via Flash Attention. These capabilities not only enhance inference speed but also reduce memory consumption, making it easier to work with larger models without compromising performance.
Overall, TensorRT-LLM is an essential tool for anyone focused on deploying high-performance LLMs in production settings, particularly when using NVIDIA GPUs. It provides the necessary infrastructure to optimize and serve models effectively, ensuring that applications can meet the demands of users in real-time scenarios.
When to use it
Use TensorRT-LLM when deploying LLMs on NVIDIA GPUs and requiring high throughput and low latency for real-time applications.
When not to use it
Avoid TensorRT-LLM if you need a simpler setup or are working with non-NVIDIA hardware, as other solutions may better suit those needs.
What you can build with it
Real-Time Chatbot Deployment
Utilize TensorRT-LLM to serve a chatbot model on NVIDIA GPUs, achieving low latency responses for user queries.
High-Throughput Text Generation
Deploy a large language model for generating text at scale, processing thousands of prompts simultaneously with optimal performance.
Multi-GPU Model Serving
Set up a multi-GPU environment to distribute the workload of a large LLM, ensuring efficient resource use and high availability.
How to install TensorRT-LLM
View source1. Install with the skills CLI
npx skills add davila7/claude-code-templates/inference-serving-tensorrt-llm --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by davila7TensorRT-LLM
NVIDIA's open-source library for optimizing LLM inference with state-of-the-art performance on NVIDIA GPUs.
When to use TensorRT-LLM
Use TensorRT-LLM when:
- Deploying on NVIDIA GPUs (A100, H100, GB200)
- Need maximum throughput (24,000+ tokens/sec on Llama 3)
- Require low latency for real-time applications
- Working with quantized models (FP8, INT4, FP4)
- Scaling across multiple GPUs or nodes
Use vLLM instead when:
- Need simpler setup and Python-first API
- Want PagedAttention without TensorRT compilation
- Working with AMD GPUs or non-NVIDIA hardware
Use llama.cpp instead when:
- Deploying on CPU or Apple Silicon
- Need edge deployment without NVIDIA GPUs
- Want simpler GGUF quantization format
Quick start
Installation
# Docker (recommended)
docker pull nvidia/tensorrt_llm:latest
# pip install
pip install tensorrt_llm==1.2.0rc3
# Requires CUDA 13.0.0, TensorRT 10.13.2, Python 3.10-3.12
Basic inference
from tensorrt_llm import LLM, SamplingParams
# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")
# Configure sampling
sampling_params = SamplingParams(
max_tokens=100,
temperature=0.7,
top_p=0.9
)
# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)
for output in outputs:
print(output.text)
Serving with trtllm-serve
# Start server (automatic model download and compilation)
trtllm-serve meta-llama/Meta-Llama-3-8B \
--tp_size 4 \ # Tensor parallelism (4 GPUs)
--max_batch_size 256 \
--max_num_tokens 4096
# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3-8B",
"messages": [{"role": "user", "content": "Hello!"}],
"temperature": 0.7,
"max_tokens": 100
}'
Key features
Performance optimizations
- In-flight batching: Dynamic batching during generation
- Paged KV cache: Efficient memory management
- Flash Attention: Optimized attention kernels
- Quantization: FP8, INT4, FP4 for 2-4× faster inference
- CUDA graphs: Reduced kernel launch overhead
Parallelism
- Tensor parallelism (TP): Split model across GPUs
- Pipeline parallelism (PP): Layer-wise distribution
- Expert parallelism: For Mixture-of-Experts models
- Multi-node: Scale beyond single machine
Advanced features
- Speculative decoding: Faster generation with draft models
- LoRA serving: Efficient multi-adapter deployment
- Disaggregated serving: Separate prefill and generation
Common patterns
Quantized model (FP8)
from tensorrt_llm import LLM
# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
model="meta-llama/Meta-Llama-3-70B",
dtype="fp8",
max_num_tokens=8192
)
# Inference same as before
outputs = llm.generate(["Summarize this article..."])
Multi-GPU deployment
# Tensor parallelism across 8 GPUs
llm = LLM(
model="meta-llama/Meta-Llama-3-405B",
tensor_parallel_size=8,
dtype="fp8"
)
Batch inference
# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]
outputs = llm.generate(
prompts,
sampling_params=SamplingParams(max_tokens=200)
)
# Automatic in-flight batching for maximum throughput
Performance benchmarks
Meta Llama 3-8B (H100 GPU):
- Throughput: 24,000 tokens/sec
- Latency: ~10ms per token
- vs PyTorch: 100× faster
Llama 3-70B (8× A100 80GB):
- FP8 quantization: 2× faster than FP16
- Memory: 50% reduction with FP8
Supported models
- LLaMA family: Llama 2, Llama 3, CodeLlama
- GPT family: GPT-2, GPT-J, GPT-NeoX
- Qwen: Qwen, Qwen2, QwQ
- DeepSeek: DeepSeek-V2, DeepSeek-V3
- Mixtral: Mixtral-8x7B, Mixtral-8x22B
- Vision: LLaVA, Phi-3-vision
- 100+ models on HuggingFace
References
- Optimization Guide - Quantization, batching, KV cache tuning
- Multi-GPU Setup - Tensor/pipeline parallelism, multi-node
- Serving Guide - Production deployment, monitoring, autoscaling
Resources
Frequently asked questions about TensorRT-LLM
Similar skills
Heap Snapshot Analysis
Investigate V8 heap snapshots for memory issues.
VS Code Performance Workflow
Automate performance investigations in VS Code.
Memory Leak Audit
Prevent memory leaks with effective coding patterns.
CPU Profile Analysis
Analyze V8 and Chrome performance profiles for optimization.
Chat Performance Testing
Benchmark and validate chat UI performance in VS Code.
Vercel React Best Practices
Optimize your React and Next.js applications for performance.
