All essays
GuideGUIDEFEB 2026

Inference with Audio Models: Whisper, Parakeet Streaming Deployment

Technical deep dive on serving audio models for GPU inference. Whisper and Parakeet architecture comparison, streaming audio processing on GPU, real-time factor benchmarks, batched transcription, and deployment patterns.

01

THE AUDIO MODEL LANDSCAPE: WHISPER VS PARAKEET

Whisper large-v3 (OpenAI, 1.55B parameters) is an encoder-decoder transformer trained on 5 million hours of multilingual audio. It processes 30-second audio chunks through a 2-layer CNN encoder that converts log-Mel spectrograms into a 1,500-token sequence, which is then processed by a 32-layer transformer encoder. The decoder generates text autoregressively. Parakeet-CTC 1.1B (NVIDIA, in collaboration with Suno AI) uses a pure CTC architecture: its 24-layer transformer encoder processes audio directly into text tokens in a single pass, without autoregressive decoding. This architectural difference has profound implications for GPU inference latency, throughput, and streaming capability.

Whisper's encoder-decoder design means inference cost grows with audio duration twice: once for the encoder (proportional to audio frames) and once for the decoder (proportional to generated text length). Whisper large-v3 processes 30 seconds of audio in approximately 450ms on H100 (300ms encoder, 150ms decoder at 100 tokens), achieving a real-time factor of 0.015. Parakeet-CTC achieves an RTF of 0.005 for the same 30-second audio segment (150ms total, single encoder pass, no autoregressive decoding). However, Whisper's decoder provides significantly better language modeling and handles code-switching, punctuation, and formatting natively, while Parakeet-CTC requires an external language model for comparable text quality.

ModelArchitectureRTF (30s audio, H100)
Whisper large-v3Encoder-decoder (1.55B)0.012-0.018
Whisper mediumEncoder-decoder (769M)0.008-0.012
Whisper smallEncoder-decoder (244M)0.005-0.008
Parakeet-CTC 1.1BCTC encoder (1.1B)0.003-0.005
Parakeet-CTC 0.6BCTC encoder (600M)0.002-0.004
Real-time targetStreaming requirementRTF < 1.0
02

AUDIO PREPROCESSING ON GPU: THE HIDDEN COST

Audio preprocessing - loading, decoding, resampling, and spectrogram computation - is the hidden GPU cost in audio inference pipelines and is traditionally done on CPU. A 30-second WAV file at 16kHz takes 2-5ms to decode on CPU, but the log-Mel spectrogram computation (STFT, Mel filter bank, log compression) takes 15-30ms on CPU using librosa or torchaudio. Offloading preprocessing to GPU using custom CUDA kernels reduces this to 2-4ms per 30-second segment on H100. At scale with thousands of concurrent audio streams, CPU preprocessing becomes the bottleneck, limiting the effective throughput of GPU inference pipelines.

The optimal architecture runs preprocessing on the GPU using CUDA-accelerated audio processing libraries (torchaudio's CUDA backend or NVIDIA's cuSignal). A custom preprocessing pipeline on GPU: decode compressed audio (MP3, Opus) using GPU-accelerated codecs, resample to 16kHz, compute log-Mel spectrogram with 80 Mel bands and 25ms window, and normalize - all in approximately 3-7ms for 30 seconds of audio on H100. This GPU preprocessing adds no additional memory pressure (the mel spectrogram is 80x3000 float32 = 0.96 MB per 30-second segment). The key throughput benefit: GPU preprocessing enables end-to-end GPU pipelines where audio enters the GPU, is processed, and the transcription exits, without CPU round-trips that add 20-40ms per inference.

03

STREAMING AUDIO INFERENCE ON GPU

Streaming speech recognition processes audio incrementally as it arrives, returning partial transcriptions with low latency. The challenge for GPU inference is the mismatch between GPU batch-oriented processing and the temporal nature of streaming audio. Whisper is fundamentally non-streaming: its 30-second audio window means partial results are only available after the full window is processed. For lower latency, Whisper can be adapted with a 5-second window and 2-second stride (overlap), where each new 5-second segment re-encodes from scratch, adding 80ms of encoder time per stride on H100. This achieves 500ms end-to-end latency but wastes compute on re-encoding overlapping audio.

Parakeet-CTC is natively streaming. Its CTC architecture processes audio frame-by-frame (each frame is 20ms of audio at 50Hz frame rate). A streaming Parakeet deployment processes 10 frames (200ms) at a time, running a forward pass on the accumulated audio buffer and emitting text tokens when the CTC blank probability drops below a threshold. This achieves 200-300ms end-to-end latency with RTF of 0.005-0.01, meaning GPU time is 1-2ms per 200ms of audio. The GPU cost of streaming Parakeet is 5-10x lower than streaming-adapted Whisper for the same latency target. For applications requiring sub-500ms transcription latency (live captioning, voice assistants, real-time translation), Parakeet-CTC on H100 is the current optimal choice at approximately $0.008 per hour of audio versus $0.035 per hour for Whisper large-v3.

MethodEnd-to-End LatencyCost per Audio Hour
Whisper large-v3 (30s window)2-5 seconds$0.035
Whisper large-v3 (5s streaming)500-800ms$0.065
Parakeet-CTC 1.1B (200ms)200-300ms$0.008
Parakeet-CTC 0.6B (200ms)200-300ms$0.005
Faster-Whisper (CTranslate2)500-1000ms$0.018
04

BATCHED TRANSCRIPTION FOR NON-REAL-TIME WORKLOADS

For offline transcription (batch processing of pre-recorded audio), batching dramatically improves GPU throughput. Whisper large-v3 achieves 2.3x real-time (1 hour of audio processed in 26 minutes) at BS 1 on H100. At BS 32, throughput reaches 18.5x real-time (1 hour of audio in 3.2 minutes), a 7x improvement from batching. The encoder benefits almost linearly from batching up to BS 64, where GPU memory becomes the constraint. The decoder benefits less because it is autoregressive - each audio segment in the batch has a different text length, causing batch divergence where the batch is padded to the longest sequence.

The optimal batch sizes for Whisper on H100: BS 16-32 for long-form transcription (10+ minute audio files), BS 8-16 for shorter segments. Beyond BS 32, the decoder memory overhead from varying text lengths reduces marginal gains. A production batch transcription pipeline on 8x H100 processes approximately 1,000 hours of audio per day with Whisper large-v3, at a cost of approximately $0.012 per audio hour on ClusterBid ($1.15/hr per GPU, 8 GPUs, 220 hours per day effective throughput). For Parakeet-CTC, batch scaling is more efficient because the CTC decoder output length is proportional to audio length (fixed), not generated text length (variable). BS 64 Parakeet-CTC achieves 45x real-time on H100, processing 1 hour of audio in 80 seconds, at $0.003 per audio hour.

05

PRODUCTION DEPLOYMENT PATTERNS

Three deployment patterns dominate audio model serving in 2026. Pattern 1: real-time streaming (Parakeet-CTC on H100 for sub-300ms latency) used by voice assistants, live captioning, and real-time translation. Pattern 2: near-real-time batch (Whisper on H100 or L40S with 5-30 second windows) for meeting transcription, voicemail processing, and asynchronous voice interfaces. Pattern 3: massive offline batch (Whisper or Parakeet on H100 with large batch sizes) for media archives, call center recording analysis, and content moderation at scale.

The GPU selection matters significantly for audio inference. Whisper large-v3 runs comfortably on L40S (16 GB VRAM, 1.2x the latency of H100 but at $0.35/hr, 3.3x cheaper). Parakeet-CTC 1.1B requires 3.5 GB for model weights plus 2-4 GB for activations at moderate batch sizes, fitting on any GPU with 8+ GB VRAM. On ClusterBid, the most cost-effective configuration for mixed audio workloads is L40S for Whisper-based transcription (48 GB VRAM handles BS 32 easily) and H100 for Parakeet-CTC streaming (where the lowest latency matters). A typical production deployment combines both: L40S handles batch transcription of recorded audio at $0.005 per hour, while H100 handles live streams at $0.008 per hour, with a router directing traffic based on the latency requirement of each request.

Deployment PatternRecommended GPUCost per Audio Hour
Streaming (<300ms)H100 SXM$0.008
Near-real-time (<5s)L40S / H100$0.005-0.012
Offline batch (hours)L40S / A100$0.003-0.005
Whisper large-v3 BS 32L40S (48 GB)$0.012
Parakeet-CTC BS 64A100 (40 GB)$0.003
Filed under
Audio Model InferenceWhisper GPU ServingParakeet StreamingSpeech Recognition GPUReal-Time FactorAudio Batch ProcessingStreaming Transcription