TOGETHER AI: FULL-STACK AI INFRASTRUCTURE
Together AI emerged from the open-source AI community. Founded by former Perplexity AI engineers and researchers, the company positions itself as a full-stack AI infrastructure platform, meaning they provide the entire stack: raw GPU compute (via their cloud), a managed inference API (competitive with OpenAI and Anthropic), a training platform (competitive with Anyscale and Modal), and open-source tools (the RedPajama dataset, inference engine contributions to vLLM, and the Together C++ inference runtime). This vertical integration is their core strategic bet: developers can start with the API and graduate to raw GPU access as their needs evolve, all within Together's ecosystem.
As of Q2 2026, Together AI operates approximately 16,000 GPUs across five data centers (US West, US East, EU West, EU Central, and a new facility in APAC Southeast). The fleet is primarily H100 (12,000) and H200 (3,000), with approximately 1,000 A100s for cost-sensitive workloads. Their infrastructure is built on a highly customized stack that extends well beyond standard cloud offerings.
| Product | Description | GPU Backend | Price Range | Target Audience | Key Differentiator |
|---|---|---|---|---|---|
| Inference API | Hosted LLM API (100+ models) | H100, H200 | $0.10-3.50/M tokens | Developers, SaaS | Custom inference engine |
| Training Platform | Managed training with RedPajama | H100 clusters | $2.50-4.00/GPU-hr | AI teams, researchers | Custom distributed library |
| Together Cloud | Raw GPU compute (bare metal) | H100, H200, A100 | $2.00-3.50/GPU-hr | Infrastructure teams | Full K8s + Slurm support |
| RedPajama Dataset | Open-source training data | N/A (download) | Free | Researchers | 3 trillion tokens |
THE CUSTOM INFERENCE ENGINE: TOGETHER'S MOAT
Together AI's most significant technical achievement is their custom inference engine. While most GPU clouds run standard vLLM or TGI for LLM inference, Together built a highly optimized C++ inference runtime that incorporates advanced techniques: FlashAttention-3 with FP8 block quantization, continuous batching with dynamic KV cache allocation, speculative decoding with n-gram draft models, and tensor parallelism with pipeline parallelism overlapped to hide communication latency. The result is 1.5-3x higher throughput per GPU compared to standard vLLM deployments for equivalent model sizes and quality targets.
The impact of this engine on GPU economics is substantial. For Llama 3 70B inference, Together achieves approximately 1,400 tokens per second on a single 8x H100 node versus 700-900 tokens/second on standard vLLM. This 1.6-2.0x throughput improvement means Together can offer inference API pricing that is competitive with or below OpenAI's while operating on fewer GPUs and maintaining healthy margins. Together's inference API pricing for Llama 3 70B is $0.88 per million tokens input and $1.08 per million tokens output, compared to OpenAI's GPT-4o at $2.50/$10.00.
| Model | Together Price (Input) | Together Price (Output) | OpenAI Equivalent | OpenAI Price (In/Out) | Speed (tok/s) |
|---|---|---|---|---|---|
| Llama 3 8B | $0.14/M | $0.14/M | GPT-4o-mini | $0.15/$0.60 | 2,800 |
| Llama 3 70B | $0.88/M | $1.08/M | GPT-4o | $2.50/$10.00 | 1,400 |
| Mixtral 8x22B | $0.90/M | $0.90/M | Claude 3.5 Sonnet | $3.00/$15.00 | 900 |
| DeepSeek-V3 | $0.80/M | $0.80/M | GPT-4o | $2.50/$10.00 | 1,100 |
| Qwen 2.5 72B | $0.90/M | $0.90/M | Claude 3 Opus | $15.00/$75.00 | 1,200 |
TRAINING PLATFORM AND DISTRIBUTED SYSTEMS
Together AI's training platform provides a managed training experience on top of their GPU clusters. The platform uses a custom distributed training library called Together Train, built on top of PyTorch FSDP with optimizations for their specific InfiniBand fabric. Key features include automatic sharding of large models (up to 405B parameters), integrated experiment tracking with W&B and MLflow, and automated checkpointing to S3-compatible blob storage with 15-minute checkpoint intervals for crash recovery.
The training platform is priced at a premium over raw GPU rental ($2.50-4.00/GPU-hr versus $2.00-3.00 for bare metal) but includes the training library, a managed Jupyter environment, pre-configured Docker images, and infrastructure support. For teams that are building on open-source models and need a production training pipeline, Together's platform competes with Anyscale, Modal, and RunPod's training offering. The differentiator is Together's deep integration with their own inference API: models trained on Together can be deployed to the inference API with a single click, creating a seamless train-to-serve pipeline.
GPU FLEET STRATEGY AND DATA CENTER FOOTPRINT
Together AI operates a fleet that is 75 percent H100, 19 percent H200, and 6 percent A100. The fleet is distributed across five data centers: two in the US (Ashburn, VA and Dallas, TX), two in Europe (Frankfurt and Amsterdam), and one in APAC (Tokyo). All nodes are connected via InfiniBand NDR fabric with a spine-leaf topology that provides full bisection bandwidth across the cluster. Together's maximum single-job size is 512 GPUs, limited by their current InfiniBand spine capacity, with plans to scale to 2,048 GPUs in 2026.
Together differentiates their GPU cloud in two ways. First, all GPUs run their custom inference engine and training stack, which means users get the performance benefits without needing to deploy or configure the software themselves. Second, Together provides a unified API for both inference and training-a developer can call Together's inference API for production serving and submit training jobs through the same API, simplifying the infrastructure stack to a single provider relationship.
OPEN-SOURCE STRATEGY AND COMMUNITY POSITIONING
Together AI's open-source contributions serve as a strategic moat. The RedPajama dataset (3 trillion tokens of open-source training data) has become a standard pre-training dataset for open LLMs, and Together's engineers are core maintainers of components in vLLM, PyTorch, and FlashAttention. These contributions create a virtuous cycle: the open-source community builds on Together's tools, which drives model compatibility with Together's inference engine, which makes the Together API the preferred hosted option for models trained on their stack.
The business impact is measurable. Models that are trained on RedPajama data or use Together's training libraries are overwhelmingly deployed on Together's inference API-Together estimates 40-50 percent of the models on their inference API have some lineage connection to their open-source work. This community-led growth strategy has enabled Together to grow inference API revenue without a direct sales team, relying instead on developer word-of-mouth and GitHub visibility.
COMPETITIVE POSITION IN THE AI INFRASTRUCTURE MARKET
Together AI occupies a unique position as the only GPU provider that is simultaneously a raw compute provider, a managed inference API, and a training platform. This breadth creates both advantages and tensions. The advantage is a unified developer experience: a model trained on Together's platform deploys to their inference API with one click. The tension is competition with customers who use raw compute but also offer their own inference APIs (like open-source model companies). Together navigates this by treating the inference API as a separate business that happens to share GPU infrastructure with the raw compute offering.
The financial trajectory supports the strategy. Together AI's inference API revenue is growing at 200-250 percent year-over-year and now accounts for 40 percent of total revenue, up from 15 percent in 2024. Raw compute rental accounts for 50 percent, and training platform for 10 percent. The inference API has higher margins (estimated 55-65 percent gross margin versus 20-30 percent for raw compute) and creates a more defensible business because switching costs are higher once an application integrates with Together's API.
