L40S Specs for Fine-Tuning: What Actually Matters
The L40S has 48GB of GDDR6 memory. That single number drives every decision you'll make about L40S LLM fine-tuning - it determines which models fit without quantization, how much gradient accumulation you'll need, and whether QLoRA saves you or not. The L40S is a PCIe card with no NVLink. If you've been reading H100 SXM tutorials that casually mention NVLink-based gradient synchronization, ignore that entire section - it doesn't apply here.
Ada Lovelace brings 733 TFLOPS of FP16 compute and 1457 TOPS of INT8, which is genuinely fast. The memory bandwidth is 864 GB/s - lower than the H100 SXM5's 3.35 TB/s, but the comparison isn't as lopsided as those numbers suggest. For LoRA fine-tuning, you're not hammering memory bandwidth the way inference does. You're loading a frozen base model and updating a small set of adapter weights. The bandwidth constraint matters much less than raw VRAM capacity and compute density. On fine-tuning specifically, the L40S trades well.
What the L40S lacks for fine-tuning is NVLink. Across multiple L40S GPUs, you're communicating over PCIe, which caps gradient synchronization bandwidth at roughly 64 GB/s bidirectional per card. For multi-GPU LoRA runs on 7B-13B models this is acceptable - the gradient tensors are small. For full SFT on 70B models requiring all-reduce across 4-8 GPUs, PCIe becomes a real bottleneck and you'll want H100 SXM or at minimum H100 PCIe with NVLink bridges.
| Specification | L40S | H100 SXM5 |
|---|---|---|
| VRAM | 48 GB GDDR6 | 80 GB HBM3 |
| Memory Bandwidth | 864 GB/s | 3.35 TB/s |
| FP16 TFLOPS | 733 | 1,979 |
| FP8 TFLOPS | 1,457 | 3,958 |
| GPU Interconnect | PCIe Gen4 x16 | NVLink 4 (900 GB/s) |
| TDP | 350W | 700W |
| Spot Rate (2026) | $1.20-1.50/hr | $2.00-2.50/hr |
VRAM Math for Fine-Tuning: 7B to 70B With QLoRA, LoRA, and Full SFT
Fine-tuning memory requirements are not the same as inference. During training you're holding the model weights, optimizer states, gradients, and activations simultaneously. For full SFT in BF16, the rule of thumb is 4x the inference VRAM - an 8-billion parameter model needs roughly 32 GB just for model + optimizer states (AdamW stores two moments per parameter, adding another 2 bytes per parameter at mixed precision). Activations during forward pass add another 4-8 GB depending on sequence length and batch size. That pushes a clean full SFT on Llama 3 8B to around 36-42 GB on a single GPU.
This is exactly where the L40S 48GB fits and the A100 40GB doesn't. L40S fine-tuning in full BF16 SFT works comfortably for 7B-8B models with batch size 4 and sequence length 2048. You have headroom. The A100 40GB version, still widely available and cheaper, forces you into gradient checkpointing or reduced batch sizes at this model size. The L40S 48GB is the right card for 7B full SFT without tricks.
LoRA changes the math dramatically. With LoRA (rank 8-16, targeting attention + MLP layers), the adapter parameters for a 7B model are roughly 10-40 million parameters - a fraction of the 7 billion base. Your optimizer states are tiny. The base model loads in BF16 at about 16 GB, activations add 4-6 GB, LoRA adapters and their gradients add maybe 2-3 GB. A 7B LoRA run fits in 24-26 GB comfortably. QLoRA loads the base in 4-bit NF4, dropping it to around 4-5 GB, with LoRA adapters in BF16. The entire 7B QLoRA run sits at 10-14 GB - a single 48GB L40S can run 3-4 simultaneous fine-tuning jobs.
For 13B models, full SFT needs 52-60 GB of VRAM - outside the L40S. But LoRA on 13B fits in 32-36 GB, leaving room for reasonable batch sizes. QLoRA on 13B squeezes to 18-22 GB. The 70B model is where L40S hits its hard wall: full SFT requires 280-320 GB across 8 H100s. LoRA on 70B needs 80-100 GB. QLoRA on 70B at 4-bit gets you to 40-48 GB, which technically fits one L40S - but the memory pressure leaves almost no room for activations at any meaningful batch size. In practice, QLoRA 70B on L40S requires batch size 1 with aggressive gradient checkpointing. It runs, but slowly.
| Model + Method | VRAM Required | L40S Fit? | GPU Count |
|---|---|---|---|
| 7B Full SFT (BF16) | ~40 GB | Single GPU | 1x L40S |
| 7B LoRA (BF16) | ~26 GB | Single GPU | 1x L40S |
| 7B QLoRA (NF4) | ~12 GB | 3x per GPU | 1x L40S |
| 13B Full SFT (BF16) | ~56 GB | Too large | 2x L40S |
| 13B LoRA (BF16) | ~34 GB | Single GPU | 1x L40S |
| 13B QLoRA (NF4) | ~20 GB | Fits well | 1x L40S |
| 70B LoRA (BF16) | ~90 GB | Too large | 4x H100 |
| 70B QLoRA (NF4) | ~46 GB | Tight fit | 1x L40S |
L40S vs H100 SXM Throughput on Real Fine-Tuning Runs
The H100 SXM5 is faster than the L40S for fine-tuning. That's not in dispute. On full SFT with BF16, H100 SXM5 runs roughly 2.3-2.8x more tokens per second than L40S on equivalent model sizes. The question is whether that speed difference justifies the price difference - and for many fine-tuning scenarios it absolutely does not.
On a Llama 3 8B full SFT run with batch size 4, sequence length 2048, an L40S achieves approximately 8,500-10,000 training tokens per second. An H100 SXM5 under the same config hits 22,000-26,000. H100 is 2.4x faster - but costs roughly 2x more per hour in 2026 spot markets ($1.20-1.50 vs $2.00-2.50). Your cost per token trained ends up comparable, sometimes favoring L40S when spot rates dip to the low end. For LoRA specifically, where compute intensity drops because you're only updating adapter weights, the L40S performance gap narrows - H100 advantage drops to roughly 1.7-2.0x while maintaining the same hourly cost advantage.
Wall-clock time is where H100 earns its premium on tight schedules. A fine-tuning run on a 1-billion-token dataset at the numbers above takes roughly 28 hours on an L40S versus 11 hours on an H100 SXM5. If you're iterating quickly on research experiments and clock time matters more than cost, H100 SXM makes sense. If you're running a production fine-tuning job and cost efficiency is the primary constraint - which is most teams doing recurring domain adaptation or instruction tuning - L40S is the more rational choice.
| Workload | L40S (tokens/s) | H100 SXM5 (tokens/s) | Cost/M tokens |
|---|---|---|---|
| 8B Full SFT, bs=4 | ~9,000 | ~24,000 | $0.037 vs $0.029 |
| 8B LoRA, bs=8 | ~14,000 | ~26,000 | $0.024 vs $0.026 |
| 13B LoRA, bs=4 | ~8,000 | ~18,000 | $0.042 vs $0.038 |
| 7B QLoRA, bs=16 | ~19,000 | ~38,000 | $0.018 vs $0.018 |
When L40S Wins - and When H100 Is Worth the Premium
L40S wins decisively for budget-constrained teams running 7B-13B LoRA or QLoRA. If you're fine-tuning Llama 3 8B, Mistral 7B, or Phi-3 Medium with LoRA adapters - doing domain adaptation, instruction tuning, or chat template alignment - you will get roughly equivalent cost per token trained as H100 and sometimes better. The L40S handles these workloads without breaking a sweat. Your GPU hours cost 40-60% less than equivalent H100 SXM capacity, and the absolute wall-clock time for typical fine-tuning datasets (10-50M tokens) is 6-18 hours on a single L40S. Acceptable for most production fine-tuning cycles.
L40S is also the right choice when you need to run many parallel small fine-tuning experiments. Research teams that run hyperparameter sweeps - testing different LoRA ranks, learning rates, or dataset mixes - can run 3-4 simultaneous QLoRA experiments on a single 48GB L40S. An equivalent H100 gives you one 7B experiment at a time unless you use MIG (which adds operational complexity). The dollar-per-experiment math strongly favors L40S for exploration-heavy workflows.
H100 SXM is the right choice when you're doing full SFT on 30B-70B models, when you need the NVLink fabric for fast multi-GPU all-reduce, or when you're on a hard deadline and need minimum wall-clock time. It's also the better choice if your workflow involves alternating between fine-tuning and inference serving on the same GPU - H100 SXM's HBM3 bandwidth advantage shows up heavily in inference, so the GPU earns its cost on both sides of the workflow. For pure fine-tuning workloads at 7B-13B scale, the L40S cost advantage is hard to ignore.
Practical Setup: Getting LoRA and SFT Running on L40S
The thing nobody tells you about running fine-tuning on L40S versus H100 SXM is the memory clock difference. GDDR6 on L40S runs at 18 Gbps effective, while HBM3 on H100 operates very differently in terms of access patterns. In practice this means your gradient accumulation steps need to be higher on L40S to maintain effective batch size without OOM errors. Start with gradient_accumulation_steps=4 and scale from there. With HuggingFace Transformers + PEFT, a typical L40S training script for 8B LoRA looks like batch_size=2, gradient_accumulation=4, giving an effective batch size of 8.
Flash Attention 2 is mandatory on L40S for fine-tuning - not optional. Without it, the attention computation during training keeps much larger activation tensors in memory, and you'll hit OOM on sequences longer than 1024 tokens with any reasonable batch size. With Flash Attention 2 enabled, you can push to 4096 sequence length on 7B LoRA with comfortable headroom. Install via pip install flash-attn and pass attn_implementation='flash_attention_2' to your model loading call.
For multi-GPU L40S setups doing LoRA across 2-4 cards, use DeepSpeed ZeRO-2 rather than ZeRO-3. ZeRO-3 partitions optimizer states across GPUs, which is designed for the NVLink bandwidth of H100 SXM. On PCIe-connected L40S cards, ZeRO-3's aggressive communication pattern actually hurts throughput - you're bottlenecked on the PCIe bandwidth during the parameter gathering passes. ZeRO-2 partitions only optimizer states (not parameters), reducing communication overhead and better matching the L40S interconnect profile. Throughput improvement over naive DDP is typically 25-40% on 4x L40S with ZeRO-2.
Where to Source L40S in 2026: Provider Landscape and Pricing Reality
L40S sits in an awkward position in the cloud market. Hyperscalers don't surface it prominently - AWS, Azure, and GCP have iterated toward H100 and H200 for enterprise AI workloads, and the L40S gets buried in niche instance families (AWS g6 for inference-adjacent workloads, Azure NVads for visualization-adjacent use cases). You won't stumble across L40S capacity doing a normal cloud comparison. This is where the neocloud and bare-metal provider market has a real advantage - many operators sitting on Ada Lovelace inventory from the 2023-2024 purchase cycle have L40S available at rates hyperscalers don't match.
Current market rates for L40S in H1 2026 run $1.20-1.50/hr per GPU on hourly spot, dropping to $0.90-1.10/hr on 1-3 month reserved contracts. Compare this to H100 PCIe at $1.50-2.00/hr spot and H100 SXM5 at $2.00-2.50/hr spot. The L40S discount versus H100 PCIe is real but modest; versus H100 SXM it's substantial. For LoRA fine-tuning on 7B-13B models where throughput per dollar matters more than absolute speed, that 40-60% hourly discount translates directly to lower training costs.
ClusterBid's network of 340+ data centers includes a meaningful slice of L40S capacity that isn't listed on hyperscaler portals. Operators who purchased L40S nodes for visualization and rendering workloads find themselves with GPUs well-suited for LoRA fine-tuning as that market has shifted. If you need L40S capacity for a fine-tuning run - whether it's a one-off domain adaptation or a recurring weekly instruction tuning job - a direct quote through ClusterBid typically surfaces 20-40% lower rates than hyperscaler equivalents because you're accessing inventory that isn't competing in the same auction markets.
| GPU | Spot $/hr | Reserved $/hr | Best For |
|---|---|---|---|
| L40S 48GB | $1.20-1.50 | $0.90-1.10 | 7B-13B LoRA/SFT |
| A100 40GB PCIe | $1.00-1.30 | $0.75-0.95 | 7B LoRA (tighter) |
| H100 PCIe 80GB | $1.50-2.00 | $1.20-1.60 | 13B-30B SFT |
| H100 SXM5 80GB | $2.00-2.50 | $1.60-2.00 | 30B+ full SFT |
| H200 SXM 141GB | $2.50-3.50 | $2.00-2.80 | 70B+ SFT, speed priority |
The Decision: A Simple Framework for L40S vs H100 Fine-Tuning
The L40S fine-tuning decision comes down to three questions. First: what is your target model size? If it's 7B-13B, L40S handles full SFT and LoRA without restriction. If it's 30B-70B full SFT, you need H100 or H200 regardless. Second: is this LoRA or full SFT? For LoRA at any size up to 13B, L40S is our first recommendation - the cost savings are significant and throughput is sufficient. For full SFT on 13B+, move to H100 PCIe 80GB minimum. Third: how important is wall-clock time versus total cost? If you're shipping a product update that needs the fine-tuned model by tomorrow morning, H100 SXM wins on wall-clock. If you're running weekly fine-tuning jobs and a 3-4x longer run time is acceptable for 50% lower cost, L40S is the rational choice.
There's a use case the community undervalues: rapid iteration with QLoRA on L40S. A research team running 10-15 experiments per week testing different fine-tuning approaches will spend roughly $180-270/week on a single L40S doing QLoRA on 7B models (3-hour jobs at $1.35/hr). The equivalent H100 SXM job completes in 1.2 hours but costs $2.25/hr - $45 per experiment versus $29. At 15 experiments per week, that's $675 versus $435. The H100 is cheaper per experiment because of its speed advantage. But for teams where a 3-hour run is acceptable overnight, L40S amortizes very well. The full financial picture depends entirely on whether you value researcher-hours or GPU-hours.
We'd pick L40S without hesitation for: any team doing repeated LoRA or QLoRA on 7B-13B models with a cost-per-run budget under $50, any organization running fine-tuning on a reserved schedule where overnight jobs are acceptable, and any research team doing hyperparameter exploration on smaller models. We'd pick H100 SXM for: full SFT on any model 30B+, teams needing NVLink for multi-GPU gradient synchronization, and inference workloads that need to share GPU capacity with fine-tuning. The price gap has narrowed in 2026 but hasn't closed - L40S remains the correct answer for the majority of production LoRA workflows.
Pricing Note
All GPU pricing referenced in this article reflects H1 2026 spot market rates sourced from ClusterBid's active order flow. Prices fluctuate based on GPU availability, contract term, region, and provider. Current live pricing for L40S, H100, and H200 capacity is available at clusterbid.com/inventory. Pricing is accurate as of the article's publication date and can change with GPU market conditions.
