WHY CODE MODELS ARE DIFFERENT FROM GENERAL LLMS
Code generation models are distinguished by fill-in-the-middle (FIM), where the model generates code for a cursor position based on both prefix and suffix context. FIM changes inference from left-to-right to bidirectional context encoding followed by middle generation, requiring a special token structure that doubles effective context length.
Code models also operate at higher token efficiency: 3-4 characters per token versus 0.6-0.8 for natural language. However, the vocabulary includes both language and code-specific tokens, making the embedding table 20-30 percent larger. GPU requirements mirror equivalent-size general LLMs, but FIM imposes 50-60 percent more prefill computation.
FILL-IN-THE-MIDDLE: THE LATENCY CONSTRAINT
The critical SLA for IDE code generation is 300-500 ms. For DeepSeek-Coder-V2-Lite (16B), prefill of a 2,048-token FIM context takes 80-100 ms on H100. A typical 20-50 token suggestion at 70-120 tok/s adds 200-400 ms. Total: 280-900 ms for 16B on one GPU, and 450-1550 ms for 236B on 4 GPUs.
16B models on single GPUs are the sweet spot for IDE latency. 70-236B models serve offline code generation (documentation, test generation, refactoring). The recommended architecture: tier 1 pool of 1-2 H100 per node running 16B models for interactive FIM, tier 2 pool of 4-8 H100 nodes running 236B models for comprehensive generation.
| Model | Params | GPUs | FIM Prefill | Gen 50 tok | Total Lat | SLA OK? |
|---|---|---|---|---|---|---|
| CodeLlama-7B | 6.7B | 1 H100 | 45 ms | 250 ms | 295 ms | Yes |
| CodeLlama-34B | 34B | 2 H100 | 120 ms | 450 ms | 570 ms | Marginal |
| StarCoder2-15B | 15B | 1 H100 | 55 ms | 300 ms | 355 ms | Yes |
| DeepSeek-Coder-Lite | 16B | 1 H100 | 85 ms | 280 ms | 365 ms | Yes |
| DeepSeek-Coder-V2 | 236B | 4 H100 | 310 ms | 250 ms | 560 ms | Marginal |
| Codestral-22B | 22B | 2 H100 | 70 ms | 350 ms | 420 ms | Yes |
FIM BATCHING AND INFERENCE OPTIMIZATION
FIM batching treats each request's prefix+suffix as a combined prefill, then decodes the middle. This is supported natively in vLLM since 0.4.0 via `--enable-fim`. A batch of 16 FIM requests with 2K prefix + 1K suffix each prefills in 400-500 ms, then each decodes independently. Effective throughput: 40-60 requests per minute at P50 400-500 ms.
Speculative decoding is particularly effective for code because syntax is predictable. A CodeLlama-7B drafter for DeepSeek-Coder-V2-Lite target achieves 80-90 percent acceptance at speculation length 5, with 1.8-2.2x generation speedup. This brings 50-token generation from 280 ms to 140-160 ms.
| Optimization | Throughput Gain | Lat Impact | Quality | Complexity |
|---|---|---|---|---|
| FP8 Quantization | 1.8-2.0x | -20% | <0.5% pass@1 | Medium |
| Speculative Decode | 1.8-2.2x | -40-50% | None | Med-High |
| FIM Prefix Cache | 1.3-1.5x | -15-25% | None | Low (vLLM) |
| KV Cache INT8 | 1.2-1.3x | -5-10% | <0.1% | Low |
| Continuous Batching | 2-4x | N/A | None | Medium |
MULTI-LANGUAGE SUPPORT AND TOKENIZATION OVERHEAD
Code models must handle multiple programming languages. StarCoder2-15B uses a 49K vocabulary trained on 619 languages. DeepSeek-Coder-V2 uses 128K vocabulary including code and natural language tokens. The larger vocabulary increases the embedding matrix by 2-3 GB compared to general LLMs.
A language detection model (FastText on CPU, <1 ms) identifies the language from the code prefix. The serving infrastructure uses this tag to select model instances and adjust generation parameters. Accuracy above 99.5 percent is required to prevent tokenization errors across languages.
The recommended approach uses a unified model with language tags in the system prompt rather than separate models per language. This adds 50-100 tokens overhead per request (2-3 ms prefill) but reduces infrastructure complexity by 10x.
ENTERPRISE DEPLOYMENT: ON-PREM AND AIR-GAPPED INFRASTRUCTURE
A self-hosted code assistant for 10,000 developers needs 50-200 concurrent requests per second at peak. Recommendation: 8-16 H100 nodes for interactive FIM (16B model) plus 4-8 H100 nodes for advanced generation (236B model). Total: 12-24 H100 GPUs at $30,000-60,000/month rental or $300,000-600,000 CAPEX.
The on-prem deployment uses a model gateway with request caching: an FIM cache achieves 10-25 percent hit rate for common patterns. For air-gapped deployments, model weights (450 GB for DeepSeek-Coder-V2 at FP16) must be physically transferred. Deployment takes 2-4 weeks initially, 1-2 weeks for updates.
B200 AND THE FUTURE OF CODE MODEL SERVING
B200's 192 GB VRAM enables single-GPU serving of 34B-class code models, eliminating TP overhead and reducing FIM prefill latency by 25-35 percent. A single B200 running CodeLlama-34B achieves 35-40 FIM requests per minute at P50 300 ms, equivalent to 2x H100 at 40 percent lower cost.
For 236B MoE models, B200 reduces TP from 4x H100 to 2x B200. At the cluster level, a 10K-developer deployment requires 6-10 B200 GPUs versus 12-24 H100 GPUs. Per-developer GPU cost drops from $3-5/month to $1.50-2.50/month.
