Together AI
Run open-source AI models on fast, affordable cloud infrastructure
Overview
Together AI is infrastructure for teams building on open-source models rather than locking into a single closed API. You get serverless inference across a large catalog of chat, vision, image, video, and audio models, plus dedicated GPU instances and full clusters when you need to fine-tune or train something yourself. Pricing is transparent and metered: pennies per million tokens for the cheaper models, up to a few dollars per million for the largest ones, with separate rates for image generation, video generation, transcription, and embeddings.
It's a strong fit if your team already knows what open-source models it wants to run and just needs somewhere fast and cheap to run them, especially if you want to avoid being tied to one closed-model vendor. It's a weaker fit if you want a simple flat-rate plan or a polished no-code interface, because everything here assumes you're comfortable with APIs, tokens, and GPU-hour math. The sheer number of models and pricing tiers can also be a lot to parse the first time you land on the pricing page.
Key Features
Serverless Inference: Call dozens of open-source chat, vision, image, video, and audio models through one API, billed per token or per unit generated.
Dedicated Endpoints: Single-tenant GPU instances for predictable latency and guaranteed capacity.
GPU Clusters: On-demand or reserved H100, H200, and B200 clusters for training and large-scale workloads.
Fine-Tuning: Supervised fine-tuning, DPO, and full fine-tuning on standard and specialized open models.
Code Sandbox: Isolated compute for running AI-generated code and code interpreter sessions.
Pricing
Starting price
Usage-based, from $0.05/1M input tokens on serverless inference (no flat monthly plan)
Serverless Inference: Pay-per-use, from $0.05/1M input tokens up to roughly $4.50/1M tokens depending on the model; images from $0.002 each, video from $0.14 each.
Dedicated Inference: Single-tenant GPU instances, e.g. H100 at $5.49/hour, B200 at $8.99/hour.
GPU Clusters: On-demand from $3.99/GPU-hour (H100); reserved capacity from $3.09/GPU-hour with longer commitments.
Fine-Tuning: Per-token pricing from $0.48/1M tokens for standard models, minimum $4 per job.
Disclaimer: pricing may change, confirm on Together AI's own pricing page before buying.
Pros
Large Model Catalog: Access to dozens of open-source chat, image, video, and audio models through one API.
Competitive Pricing: Per-token rates are generally cheaper than closed-model providers for comparable open models.
Full Infrastructure Stack: Covers inference, fine-tuning, and raw GPU clusters in one place, no need to stitch together separate vendors.
No Vendor Lock-In: Because it's built on open models, you can move your workloads elsewhere if needed.
Cons
No Flat Subscription: Everything is metered, so costs are harder to predict than a fixed monthly plan.
Developer-Only: No consumer-facing app, you need to write code to use any of it.
Overwhelming Pricing Page: Dozens of models and instance types with different per-unit rates makes cost estimation a real exercise.
Model Quality Varies: You're dependent on the underlying open-source model's capabilities, which can lag behind top closed models like GPT or Claude on hard reasoning tasks.
What Makes It Unique
Full-Stack Open Model Infrastructure: One of the few platforms offering serverless inference, fine-tuning, and raw GPU clusters together, so you can go from prototyping to training your own model without switching providers.
Kay Score
7.7
/ 10
Tool Information
Pricing
Usage-based, from $0.05/1M input tokens on serverless inference (no flat monthly plan)
Category
Coding & Development
Platform
Web / iOS / Android
Last Updated
