All essays
TechnicalDEEP DIVEFEB 2026

GPU Infrastructure for Legal AI and Document Intelligence: Compliance, Throughput, and Deployment Reality in 2026

GPU infrastructure guide for legal AI in 2026: data residency compliance, long-context sizing, batch throughput math, and SOC 2 provider checklist for legaltech CTOs.

01

Legal AI Workloads Don't Look Like General AI Workloads

Most GPU infrastructure guides are written for the standard AI inference pattern: short prompts, short outputs, thousands of concurrent requests. Legal AI is the structural opposite. You have long inputs - sometimes very long - moderate outputs, medium concurrency, and batch jobs that run overnight against deadlines measured in court dates, not SLAs.

Contract review is the highest-volume workload for most legal AI deployments. You're feeding 50-200 pages of dense legal text into a model to extract specific clauses, flag risk provisions, compare against standard templates. The GPU isn't being hammered with 10,000 simultaneous short requests. It's processing one giant document at a time, with the bottleneck being memory bandwidth rather than raw compute. Legal language is dense enough that a 150-page commercial agreement easily hits 60,000-80,000 tokens. A cross-border M&A agreement with exhibits crosses 200,000 tokens.

Case research and deposition analysis push context requirements even further. A full deposition transcript - 300-400 pages - runs 120,000-200,000 tokens. A litigation file across a complex matter can be millions of tokens total, though you rarely push that through a single context window at once. eDiscovery is the opposite problem: millions of short documents (1-5 pages each) that need classification, privilege review, and relevance scoring at high throughput. The GPU requirement here is about volume processing speed, not context length. These four workload types - contract review, case research, deposition analysis, and eDiscovery - drive most legal AI GPU infrastructure decisions in 2026, and they have different hardware profiles.

02

How Context Length Determines Your GPU Budget: The Memory Math for Legal Documents

The legal AI GPU infrastructure question that trips people up most often is simple: how much VRAM do I actually need? The answer depends on three variables - the model you're running, the context length of your documents, and your batch size. Get the math wrong and you either buy too little and hit out-of-memory errors mid-batch, or overspend on VRAM you're not using.

Take a 70B parameter model at FP8 precision - a reasonable choice for production legal AI in 2026. The model weights occupy approximately 70GB of VRAM. The KV cache for a single sequence is roughly 160KB per token for a model like Llama 3 70B (80 transformer layers, 8 grouped-query attention heads, 128 head dimension, FP8). At 64,000 tokens - about 150 pages - that's roughly 10GB of KV cache. At 128,000 tokens - a full agreement with exhibits - about 20GB. At 512,000 tokens, 80GB. This is why the H100's 80GB VRAM is a hard constraint for legal AI: model weights plus KV cache leave you no room for anything beyond ~64K tokens. You hit the ceiling at precisely the context length where the interesting legal AI work starts.

The H200 changes this significantly. With 141GB of HBM3e, a 70B FP8 model has ~71GB free for KV cache - enough for single-sequence contexts up to roughly 440,000 tokens, or practical batch processing at 128K tokens with batch sizes of 3-4. For most legal contract review workloads, the H200 is the right minimum configuration. The B200 at 192GB gives meaningful headroom above that - covering 256K context at batch size 4, or accommodating 1M-token depositions with tensor parallelism across two GPUs instead of four.

Context LengthKV Cache (70B FP8)Min VRAM RequiredGPU Config
32K tokens (~75 pages)5 GB75 GBH200 SXM (141 GB)
128K tokens (~300 pages)20 GB90 GBH200 SXM (141 GB)
256K tokens (~600 pages)40 GB110 GBH200 or B200
512K tokens (~1,200 pages)80 GB150 GBB200 (192 GB)
1M tokens (~2,400 pages)160 GB230 GB2x H200 or 1x B300
03

Data Residency Is Non-Negotiable: US and EU Requirements for Attorney-Client Privileged AI

The legal sector has a confidentiality obligation that does not bend to infrastructure convenience. In the US, ABA Model Rule 1.6 requires lawyers to make reasonable efforts to prevent unauthorized disclosure of client information. In practice, bar associations in New York, California, and most other jurisdictions have issued ethics opinions that extend this obligation explicitly to cloud-hosted AI tools. The question is no longer 'is using cloud AI permissible' - most bars have approved it with conditions. The question is: what specific conditions does your general counsel need to document in writing?

Data residency is the first condition that matters. Most major law firms have adopted policies requiring that client data - including data fed to AI models for analysis - stays within the jurisdiction of the engagement. US litigation data stays in the US. EU client files stay in EU/EEA data centers. This sounds straightforward until you try to operationalize it. 'Runs on AWS us-east-1' is not a guarantee that your data never transits EU infrastructure for backup, replication, or network routing. Bare metal single-tenant capacity in a named, auditable data center facility - Equinix DC6 in Ashburn, QTS Atlanta, whatever your firm's legal ops team can verify - is the only approach that gives your partners a defensible paper trail.

EU-based law firms and multinational practices face added complexity from GDPR Article 44, which prohibits transferring personal data to third countries without adequate protections. Post-Brexit UK GDPR creates a parallel set of requirements for firms with UK client work. In practice, this means you need separate GPU deployments for EU-touching data and US-touching data, often with no shared infrastructure. Many legal AI teams running on single-region cloud GPUs are unknowingly violating their firms' own data governance policies because the cloud provider's shared infrastructure doesn't map cleanly to the legal definition of 'data residency.'

04

Batch Processing Throughput - GPU Sizing for 10K to 1M Pages Per Day

Let's do the throughput math that legal AI CTOs actually need. The core question is: given a document processing workload, how many GPUs do I need to hit a specific daily volume? Start with a realistic assumption for legal document processing - a 70B FP8 model processing 8,000-token documents (roughly 20 pages of legal text) at FP8 precision, with a single H200 SXM running batch_size=8.

A single H200 SXM can sustain approximately 8,000-12,000 tokens per second for standard legal document inference at these settings. Call it 10,000 tokens/second as a working estimate. At 8,000 tokens per document, that's 1.25 documents per second, or 108,000 documents per day. One H200. That sounds like you can handle everything with one GPU - but that's optimal throughput with perfectly batched uniform-length documents. Real legal workloads have variable document lengths, pipeline overhead, storage I/O latency, and job scheduling overhead. Apply a 40% efficiency discount and you're at ~65,000 documents per day per H200. For eDiscovery processing at 1M pages - roughly 200,000 documents - you need 3-4 H200s running continuously. For 100K documents with a 4-hour deadline window, you need 8-10 H200s running in parallel.

These numbers shift meaningfully when you move to longer-context workloads. Contract review at 64,000 tokens per document (150-page agreements) drops throughput to roughly 150 documents per hour per H200 - a batch of 10,000 agreements takes 67 hours on a single GPU. For time-sensitive M&A due diligence where you have 3 days to review 500 agreements, you're looking at a 4-GPU deployment at minimum. The B200 roughly matches the H200 on throughput-per-dollar for long-context batch workloads, with the advantage of its larger VRAM eliminating multi-GPU KV cache coordination overhead for documents in the 128K-256K range.

Daily VolumeAvg Doc SizeGPUs Required (H200)Throughput Mode
10K pages/day5K tokens1x H200Standard inference
100K pages/day5K tokens2-3x H200Batch processing
100K pages/day64K tokens (long-form)8-10x H200Long-context inference
1M pages/day5K tokens20-25x H200eDiscovery scale
1M pages/day32K tokens80-100x H200Contract review at scale
05

Which GPU for Legal AI in 2026: H100 vs H200 vs B200 vs B300

The H100 SXM is not the right choice for production legal AI in 2026, and we'd say that even at its current spot price of around $1.80-2.50/hr. The 80GB VRAM cap forces you into tensor parallelism for any document over 64K tokens, which adds networking overhead to what should be a simple single-GPU inference job. You also lose the ability to batch multiple mid-length documents together efficiently. The H100 PCIe is even more constrained at the same 80GB. Both are fine for development and testing, or for eDiscovery workloads where you're processing short documents at very high volume and never need a large context window.

The H200 SXM is the workhorse for most legal AI deployments in 2026. At $3.07-3.50/hr spot, it handles the full range of legal document workloads - contract review up to 150 pages, deposition analysis at 128K tokens, RAG systems with generous context windows. The 141GB HBM3e eliminates the multi-GPU requirement for most single-document workflows, which simplifies your inference serving setup considerably and reduces infrastructure cost. For a law firm processing 50,000-200,000 documents per month, a cluster of 4-8 H200s covers most workloads comfortably with room for peak demand.

The B200 at 192GB becomes compelling when you're regularly dealing with 256K+ token documents - a full litigation file, combined deposition transcripts, or multi-contract review. At current spot pricing of $3-6/hr, you're paying a premium over H200 but eliminating tensor parallelism complexity for long-context jobs. The B300 at 288GB HBM3e and roughly $5-6.50/hr spot is for the edge of what legal AI asks for today - 1M-token context windows for full case files, or serving multiple concurrent sessions of a 405B parameter model that your firm has fine-tuned on its own historical documents. Most legal AI teams don't need B300 today, but the teams planning 2027 infrastructure for very large models or ultra-long-context workflows should price it now.

GPUVRAMSpot PriceMax Context (70B FP8)Best Legal AI Use Case
H100 SXM80 GB$1.80-2.50/hr~60K tokenseDiscovery, short-doc batch
H200 SXM141 GB$3.07-3.50/hr~440K tokensContract review, depositions
B200192 GB$3.00-6.00/hr~750K tokensLong-form contracts, large RAG
B300 NVL288 GB$5.00-6.50/hr~1.3M tokensFull case files, 405B models
06

8 Things Every Legaltech Team Must Verify Before Signing a GPU Contract

The standard GPU provider qualification checklist AI teams use is wrong for legal AI. Most of it focuses on network latency, provisioning speed, and spot instance interruption rates - none of which matter for attorney-client privileged workloads compared to the compliance questions. Here is the actual checklist your legal ops and security teams should be running.

First: SOC 2 Type II certification, not Type I. Type I is a point-in-time assessment; Type II covers a period of at least six months and proves the provider's controls actually function continuously. Get the report, read the 'control exceptions' section, and confirm it was issued within the last 12 months. Second: a signed Data Processing Agreement that explicitly covers your AI inference workloads and names the specific data centers where processing will occur. Generic DPAs that say 'US-based infrastructure' are insufficient - you need facility-level specificity. Third: single-tenant bare metal confirmation. Ask directly: 'On the physical server where my GPU jobs run, do other customers' workloads also run?' Anything other than an unambiguous no means you're on shared infrastructure.

Fourth: VRAM data remanence policy. Ask how the provider ensures that GPU memory is cleared between customer workloads. This is rarely documented publicly and the answer varies widely - some providers do cryptographic wiping, some rely on OS-level process isolation that is not equivalent, some have no formal policy at all. Fifth: audit rights - can your security team physically visit the facility and review access logs? Sixth: employee background checks covering all personnel with physical or remote access to hardware. Seventh: if any of your client work involves healthcare litigation, confirm HIPAA BAA availability. Medical records in discovery are increasingly common and require a signed BAA. Eighth: data egress rights - can you retrieve your fine-tuned model weights and cached embeddings without penalty if you switch providers?

07

Why Single-Tenancy Is the Only Defensible Option for Legal AI Workloads

Here is the thing most GPU infrastructure guides do not tell you about shared GPU infrastructure: VRAM is not guaranteed to be zeroed between jobs on many multi-tenant platforms. When your batch job finishes and the next customer's job starts on the same physical GPU, the HBM3e memory may still contain residual data from your workload. Some platforms address this with explicit memory wipe routines; many do not document their policy at all. For most AI workloads this is a theoretical risk. For legal AI processing attorney-client privileged documents, it is a compliance issue that your general counsel cannot sign off on.

Beyond memory remanence, shared GPU platforms introduce a second risk that is harder to quantify: side-channel attacks. HBM3e memory bandwidth contention and power consumption patterns can theoretically be observed by a co-tenant on the same physical package, enabling inference about the workload running alongside them. This is not a common attack vector in practice, but it is the type of risk that bar ethics opinions and law firm security policies are beginning to address. The ABA's 2023 Formal Opinion 498 on virtual practice and data security points toward this category of risk without naming it explicitly.

Single-tenant bare metal eliminates both concerns. Your GPU jobs are the only jobs running on that physical server. VRAM contains only your data. No co-tenant exists to observe your memory access patterns. This is not available from AWS, Azure, or GCP in their standard GPU instance types - their 'dedicated host' options provide physical server isolation for CPU instances but GPU instances remain on shared pools in most regions. The providers that offer true single-tenant GPU bare metal are primarily neoclouds and colocation-backed bare metal providers. Identifying compliant single-tenant GPU capacity in US-jurisdiction data centers - which is genuinely scarce - is exactly the problem that ClusterBid's sourcing desk is structured to solve.

08

Where to Source Compliant Single-Tenant GPU Capacity for Legal AI in 2026

Single-tenant H200 and B200 GPU capacity with written data residency guarantees, SOC 2 Type II coverage, and DPA availability is not something you can just self-serve on the major cloud platforms. It requires sourcing from providers who specifically build this offering - and because the compliance overhead increases their operating costs, the supply is thinner than general GPU capacity. In mid-2026, the waiting time for compliant single-tenant H200 bare metal with all required compliance documentation is typically 2-4 weeks, compared to immediate availability for standard shared GPU instances.

The key variables to specify when sourcing for legal AI: jurisdiction (US-only or EU-only), minimum VRAM per node (141GB H200 for most contract review, 192GB B200 for long-form or batch work at scale), number of nodes, and contract term. Longer-term commitments unlock 20-35% price discounts over spot rates - meaningful when you're running GPU capacity 24/7 for a firm-wide document intelligence platform. Short-term spot capacity is fine for eDiscovery bursts and pilot projects; reserved capacity makes sense for production systems processing client documents continuously.

The legal AI infrastructure buildout cycle typically starts with a pilot on 2-4 GPUs, runs 3-6 months of production validation, and then scales to 8-20+ nodes for full deployment. The compliance requirements need to be established in the pilot - retrofitting data residency and single-tenancy onto an existing deployment is significantly harder than building it in from day one. If your team is at the pilot or early-production stage, the conversations to have now are about which data center facilities your compliance team will accept, whether you need HIPAA BAA coverage in addition to SOC 2, and what your peak batch throughput requirements look like - not just what your daily average is.

Filed under
Legal AI InfrastructureDocument IntelligenceData ResidencySOC 2 ComplianceLong-Context LLMGPU Compute Legal TechSingle-Tenancy GPU