Compute Planner
v1.0.0

AI Model Compute Planner

Estimate compute requirements, VRAM footprint, multi-GPU needs, cost, sharding behavior, and NVIDIA NIM compatibility for any LLM.

AI Model Compute Planner

Use Case Planner
Select an industry and use case to get recommended models and configuration.
?
?
Model Selection?
Define the AI model you plan to analyze.
Estimated VRAM: 13.9 GB
LLaMA8B paramsfp16
B
Architecture Configuration?
Fine-tune the model architecture parameters.
Runtime Configuration?
Adjust inference and training parameters to refine memory estimates.

Vector Database?

Calculated Architecture Results?
Attn / Layer
67.11 M
FFN / Layer
117.44 M
Total Params
6.43 B
KV Cache
0.54 GB
Inf. Memory
13.90 GB
Train Memory
107.04 GB
Inference Memory?
BF16
13.90 GB
Weights
Weights
12.86 GB
KV Cache
0.54 GB
Training Memory?
FP32 + ADAM
107.04 GB
Weights
25.7 GB
Optimizer
51.4 GB
Gradients
25.7 GB
Activations
2.1 GB
CUDA
2.0 GB
Note: Training memory includes FP32 weights, gradients, optimizer states, CUDA overhead, and an approximate activation term based on the KV cache footprint. Actual training usage may vary by implementation.

Compatible GPUs for Inference?

Data Center
NVIDIA DGX B300 – Multi-Tenant Inference Node (8× B300 SXM)High Concurrency
2100 GB
• 10RU• $38/hr• 15.1 kW
High
Headroom
NVIDIA B300 SXM (Blackwell Ultra)High Concurrency
262.5 GB
• SXM• $4.75/hr• 1.89 kW
High
Headroom
AMD Instinct MI350X
288 GB
• OAM• $4/hr
High
Headroom
NVIDIA B200 NVL
192 GB
• SXM• $4.5/hr
High
Headroom
AMD Instinct MI300X
192 GB
• OAM• $3/hr
High
Headroom
NVIDIA H200
141 GB
• SXM• $3.2/hr
High
Headroom
AMD Instinct MI250X
128 GB
• OAM• $2/hr
High
Headroom
NVIDIA H100 80GB
80 GB
• PCIe/SXM• $2.5/hr
High
Headroom
NVIDIA A100 80GB
80 GB
• SXM• $1.8/hr
High
Headroom
Workstation
NVIDIA RTX 6000 Blackwell
64 GB
• PCIe• $1.8/hr
High
Headroom
NVIDIA L40S
48 GB
• PCIe• $1.2/hr
High
Headroom
NVIDIA A40
48 GB
• PCIe• $1/hr
High
Headroom
NVIDIA A10G
24 GB
• PCIe• $0.8/hr
Medium
Headroom
Consumer
NVIDIA L4
24 GB
• PCIe• $0.6/hr
Medium
Headroom
NVIDIA RTX 4090 (Consumer)
24 GB
• PCIe• $0.4/hr
Medium
Headroom
NVIDIA T4
16 GB
• PCIe• $0.3/hr
Low
Headroom

Compatible GPUs for Training?

Data Center
NVIDIA DGX B300 – Multi-Tenant Inference Node (8× B300 SXM)High Concurrency
2100 GB
• 10RU• $38/hr• 15.1 kW
5%
Utilization
AMD Instinct MI350X
288 GB
• OAM• $4/hr
37%
Utilization
NVIDIA B300 SXM (Blackwell Ultra)High Concurrency
262.5 GB
• SXM• $4.75/hr• 1.89 kW
41%
Utilization
NVIDIA B200 NVL
192 GB
• SXM• $4.5/hr
56%
Utilization
AMD Instinct MI300X
192 GB
• OAM• $3/hr
56%
Utilization
NVIDIA H200
141 GB
• SXM• $3.2/hr
76%
Utilization
AMD Instinct MI250X
128 GB
• OAM• $2/hr
84%
Utilization
NVIDIA H100 80GBLow Memory
80 GB
• PCIe/SXM• $2.5/hr
134%
Utilization
NVIDIA A100 80GBLow Memory
80 GB
• SXM• $1.8/hr
134%
Utilization
Workstation
NVIDIA RTX 6000 BlackwellLow Memory
64 GB
• PCIe• $1.8/hr
167%
Utilization
NVIDIA L40SLow Memory
48 GB
• PCIe• $1.2/hr
223%
Utilization
NVIDIA A40Low Memory
48 GB
• PCIe• $1/hr
223%
Utilization
NVIDIA A10GLow Memory
24 GB
• PCIe• $0.8/hr
446%
Utilization
Consumer
NVIDIA L4Low Memory
24 GB
• PCIe• $0.6/hr
446%
Utilization
NVIDIA RTX 4090 (Consumer)Low Memory
24 GB
• PCIe• $0.4/hr
446%
Utilization
NVIDIA T4Low Memory
16 GB
• PCIe• $0.3/hr
669%
Utilization
Workload Characteristics?
?
Standard
?
10 users
?
Calculated RPS0.17 req/s
(10 users × 1 req/min) ÷ 60
?
?
Tokens per Request:768
Requests per Second:0.17
Total Tokens / Sec:128
Est. Capacity Usage0%
B300 is designed for high-concurrency and long-context inference. Lower single-model utilization reflects reserved headroom for concurrent users and KV cache growth.
Hardware Profile?
?

Est. per GPU

?

Per GPU instance

NVIDIA DGX B300 – Multi-Tenant Inference Node (8× B300 SXM)
KV-Cache Optimized
• Data Center• 10RU• NVIDIA• Released 2025• Power: 15.1 kW

Model memory fits on a single GPU.

Estimated VRAM usage per unit: 13.9 GB.

Estimation Results?
Based on your model, workload, and hardware configuration.

Required Hardware

?
8GPUs
Latency-bound
Est. TTFT: 151ms / 3000ms
NVIDIA DGX B300 – Multi-Tenant Inference Node (8× B300 SXM)
1 per replica (for memory)
× 8 replicas (for throughput)

Monthly Compute

5,760hours

Continuous 24/7 operation

Estimated Cost

$218,880/mo

@ $38/hr/GPU

Calculation Logic

Note: GPU count increased to meet TTFT target per NVIDIA NIM guidance.

Primary Model Logic
P * bp + B * S * L * H * 2 * bp_kv
Reference: Sizer Consolidated HTML

Estimates are approximate. Actual performance varies based on specific quantization, batching strategies (e.g., vLLM, TGI), and hardware utilization.

© 2026 AI Model Compute Planner. Built with React & Vite.

base44
Edit with Base44