Speech Recognition: Workload Profile
Speech Recognition workloads have distinct GPU requirements compared to standard LLM inference. Models range from 0.3B-2B parameters. The primary performance bottleneck is Compute-bound for audio encoding. Key metrics: tokens/second, latency P50/P95/P99, batch throughput, and memory utilization. Production deployments typically use 1-2 GPUs typically enough GPUs.
Recommended GPU Configurations
Recommended GPUs for Speech Recognition: L4, L40S, T4, A10. For models in the 0.3B-2B range, GPU selection depends on: model architecture (encoder-only, decoder-only, encoder-decoder, or diffusion), precision requirements (FP16/FP8/INT4), latency SLAs (real-time vs batch), and throughput targets. Batch sizes of Batch size 16-64 for production are typical for production throughput optimization.
VRAM Requirements
VRAM needs for Speech Recognition vary by model. A 0.3B-2B model at FP16 requires approximately 0 GB for model weights. KV cache for encoder-decoder architectures may require additional 1-4 GB per sequence. Diffusion models additionally need latent space working memory. INT4 quantization reduces weight memory by 75% but increases compute requirements.
Throughput and Latency Expectations
Typical throughput for Speech Recognition on recommended hardware: varies by batch size. For real-time inference, latency targets should be 200-500ms P99. Batch inference can process at 2-10x real-time throughput depending on model complexity and GPU configuration. Production systems should maintain GPU utilization above 70% for cost efficiency.
Production Architecture Patterns
Production Speech Recognition deployment patterns include: dedicated GPU instances per model version; pooled GPU clusters with dynamic model loading; autoscaling based on queue depth or GPU utilization; GPU-backed serverless inference for variable workloads; and multi-model serving on shared GPU instances using model parallelism or MIG partitioning.
Cost Optimization Strategies
Optimize costs for Speech Recognition workloads: choose spot/preemptible GPUs for batch inference with checkpointing; use reserved instances for baseline traffic with on-demand overflow; implement GPU autoscaling to minimize idle capacity; use model quantization to reduce GPU requirements by 2-4x; and batch process non-real-time workloads during off-peak pricing periods.
