Hardware & Computational Tools Directory
Showing 145 - 192 of 4,250 verified tools.
Phi-3 Medium 14B High-Density (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Phi-3 Medium 14B High-Density (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Phi-3 Medium 14B High-Density (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Phi-3 Medium 14B High-Density quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Yi-1.5 34B 200K Context (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Yi-1.5 34B 200K Context (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Yi-1.5 34B 200K Context (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Yi-1.5 34B 200K Context (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Yi-1.5 34B 200K Context (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Yi-1.5 34B 200K Context (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Yi-1.5 34B 200K Context (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Yi-1.5 34B 200K Context quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
StarCoder-2 15B Code Synthesis (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
StarCoder-2 15B Code Synthesis (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
StarCoder-2 15B Code Synthesis (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
StarCoder-2 15B Code Synthesis (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
StarCoder-2 15B Code Synthesis (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
StarCoder-2 15B Code Synthesis (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
StarCoder-2 15B Code Synthesis (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for StarCoder-2 15B Code Synthesis quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
CodeLlama 70B Programming Specialist (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
CodeLlama 70B Programming Specialist (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
CodeLlama 70B Programming Specialist (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
CodeLlama 70B Programming Specialist (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
CodeLlama 70B Programming Specialist (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
CodeLlama 70B Programming Specialist (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
CodeLlama 70B Programming Specialist (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CodeLlama 70B Programming Specialist quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Flux.1 Schnell 12B DiT Image Model (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Flux.1 Schnell 12B DiT Image Model (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Flux.1 Schnell 12B DiT Image Model (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Flux.1 Schnell 12B DiT Image Model (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Flux.1 Schnell 12B DiT Image Model (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Flux.1 Schnell 12B DiT Image Model (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Flux.1 Schnell 12B DiT Image Model (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Schnell 12B DiT Image Model quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Flux.1 Dev 12B High-Quality DiT (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Flux.1 Dev 12B High-Quality DiT (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Flux.1 Dev 12B High-Quality DiT (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Flux.1 Dev 12B High-Quality DiT (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Flux.1 Dev 12B High-Quality DiT (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Flux.1 Dev 12B High-Quality DiT (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Flux.1 Dev 12B High-Quality DiT (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Flux.1 Dev 12B High-Quality DiT quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Stable Diffusion 3.5 Large 8B (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Stable Diffusion 3.5 Large 8B (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Stable Diffusion 3.5 Large 8B (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Stable Diffusion 3.5 Large 8B (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Stable Diffusion 3.5 Large 8B (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Stable Diffusion 3.5 Large 8B (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Stable Diffusion 3.5 Large 8B (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion 3.5 Large 8B quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Stable Diffusion XL 6.6B Base (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Stable Diffusion XL 6.6B Base (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Stable Diffusion XL 6.6B Base (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.