Hardware & Computational Tools Directory
Showing 97 - 144 of 4,250 verified tools.
Mixtral 8x7B MoE Classic (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Mixtral 8x7B MoE Classic quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Mixtral 8x7B MoE Classic (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Mixtral 8x7B MoE Classic quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Gemma-2 27B Google Research (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Gemma-2 27B Google Research (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Gemma-2 27B Google Research (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Gemma-2 27B Google Research (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Gemma-2 27B Google Research (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Gemma-2 27B Google Research (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Gemma-2 27B Google Research (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 27B Google Research quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Gemma-2 9B High-Precision (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Gemma-2 9B High-Precision (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Gemma-2 9B High-Precision (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Gemma-2 9B High-Precision (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Gemma-2 9B High-Precision (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Gemma-2 9B High-Precision (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Gemma-2 9B High-Precision (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Gemma-2 9B High-Precision quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Command R+ 104B Cohere Enterprise (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Command R+ 104B Cohere Enterprise (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Command R+ 104B Cohere Enterprise (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Command R+ 104B Cohere Enterprise (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Command R+ 104B Cohere Enterprise (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Command R+ 104B Cohere Enterprise (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Command R+ 104B Cohere Enterprise (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R+ 104B Cohere Enterprise quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Command R 35B Enterprise (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Command R 35B Enterprise (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Command R 35B Enterprise (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Command R 35B Enterprise (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Command R 35B Enterprise (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Command R 35B Enterprise (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Command R 35B Enterprise (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Command R 35B Enterprise quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Nemotron-4 340B NVIDIA Synthetic (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Nemotron-4 340B NVIDIA Synthetic (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Nemotron-4 340B NVIDIA Synthetic (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Nemotron-4 340B NVIDIA Synthetic (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Nemotron-4 340B NVIDIA Synthetic (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Nemotron-4 340B NVIDIA Synthetic (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Nemotron-4 340B NVIDIA Synthetic (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Nemotron-4 340B NVIDIA Synthetic quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Phi-3.5 MoE 16x3.8B Microsoft (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Phi-3.5 MoE 16x3.8B Microsoft (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Phi-3.5 MoE 16x3.8B Microsoft (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Phi-3.5 MoE 16x3.8B Microsoft (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Phi-3.5 MoE 16x3.8B Microsoft (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Phi-3.5 MoE 16x3.8B Microsoft (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Phi-3.5 MoE 16x3.8B Microsoft (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3.5 MoE 16x3.8B Microsoft quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Phi-3 Medium 14B High-Density (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Phi-3 Medium 14B High-Density (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Phi-3 Medium 14B High-Density (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Phi-3 Medium 14B High-Density (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.