Developer/NGC Catalog/GPU Instances
Active GPUs
1,248units
▲ 12.4% vs last wk
GPU Utilization
78.3%
▲ 4.1 pts
Inference Req / day
42.7M
▲ 8.9%
Compute Spend (MTD)
$94.2K
▼ 3.2% under budget
GPU Instances 12 RUNNING
All
Running
Queued
Instance GPU Utilization Status Region Uptime
llm-prod-a100
8× A100 80GB · NIM Llama-3.1-70B
8
92% / 640GB
Running us-west-2 14d 6h
training-h200
4× H200 141GB · PyTorch 2.4
4
87% / 564GB
Training us-east-1 2d 11h
inference-b200
2× B200 192GB · TensorRT-LLM
2
64% / 384GB
Serving eu-central-1 7d 22h
rag-embed-l40s
4× L40S · NeMo Retriever
4
31% / 96GB
Idle us-west-2 21d 3h
vision-t4-batch
8× T4 · Triton Inference
8
/ 128GB
Maintenance ap-south-1
Compute Throughput
1H
24H
7D
Training
Inference
Data pipeline
Quarterly Allocation
75%
Used
GPU-hours budget
Enterprise tier · resets in 42 days
Used37,440 h
Remaining12,560 h
Total allocation50,000 h

NGC Catalog — Featured Models

Models
Containers
Helm Charts
Resources
All frameworks ×
LLM
NIM
Computer Vision
Speech
+ Filter
Showing 6 of 4,827 collections
L3
Llama-3.1-70B-Instruct
NVIDIA · ✓ Verified · v1.0.2
NIM LLM Updated 2d
Optimized instruction-tuned 70B parameter model for dialog, reasoning, and code. TRT-LLM accelerated, FP8 quantized for H100/H200.
Pulls2.4M
Context128K
LicenseLlama 3.1
N
NVIDIA NeMo Retriever
NVIDIA · ✓ Verified · v0.4.1
NIM Embeddings RAG
Microservices for enterprise retrieval-augmented generation — embedding, reranking, and vector store connectors with cu optim.
Pulls847K
Latency12ms
LicenseApache 2.0
CV
Segment Anything 2.1
Meta · ✓ Verified · v2.1
Computer Vision Segmentation New
Promptable image and video segmentation. Optimized TensorRT pipeline delivering 8× throughput over reference on L40S.
Pulls1.1M
FPS142
LicenseApache 2.0
AS
Parakeet-RNNT 1.1B
NVIDIA · ✓ Verified · v1.1.0
NIM Speech ASR
State-of-the-art automatic speech recognition with streaming support. RNN-T architecture delivering real-time transcription.
Pulls623K
RTF0.08
LicenseCC-BY 4.0
M
Mistral-Nemo-12B
Mistral AI · ✓ Verified · v0.3
NIM LLM
Multilingual 12B model with 128K context. FP8 quantized and TRT-LLM optimized for efficient inference on a single H100.
Pulls389K
Context128K
LicenseApache 2.0
SD
SDXL TensorRT
Stability AI · ✓ Verified · v1.0
Generative Diffusion
Stable Diffusion XL recompiled for TensorRT. 3.1× speedup over PyTorch on A100 with FP8 execution.
Pulls1.8M
Steps/s8.4
LicenseCommunity
Developer Tools
NGC CLI
Pull containers, deploy models, manage configurations from terminal.
$ pip install ngccli
CUDA Toolkit 12.6
Latest SDK with cuDNN, NCCL, and NVSHMEM for accelerated computing.
v12.6 · 4.2 GB
NVIDIA AI Enterprise
Licensed software stack with support, security, and certification.
License active
DCGM Exporter
GPU metrics for Prometheus — utilization, memory, temperature, power.
v3.3.0
Recent Activity
Live
Deployment llama-3.1-70b-nim scaled to 4 replicas
2 min ago · us-west-2
Training job finetune-vit-l16 checkpoint saved (epoch 12/50)
14 min ago · training-h200
Instance vision-t4-batch entered maintenance window
38 min ago · ap-south-1
NIM endpoint parakeet-asr-prod created and healthy
1 h ago · eu-central-1
API key ngc-prod-key-7f3a rotated successfully
3 h ago · Acme AI Labs