
Full-stack AI cloud for inference, fine-tuning, and GPU clusters
Gallery
4 items




About Together AI
Together AI, styled as the "AI Native Cloud," is a full-stack platform for running, fine-tuning, and training open-source AI models at production scale. Rather than building its own closed frontier models, Together AI focuses on making the best open-source models, such as DeepSeek, Llama, Qwen, GLM, and Kimi, fast, affordable, and easy to deploy through serverless inference APIs, dedicated infrastructure, and its own vertically integrated GPU cloud. The platform's OpenAI-compatible API lets teams switch from closed-model providers with minimal code changes while accessing over 200 models across text, vision, image, audio, and video.
Founded in June 2022 by Vipul Ved Prakash, Ce Zhang, Chris Ré, and Percy Liang, Together AI is grounded in systems research from its team, including foundational work like FlashAttention, and channels that research directly into product performance gains such as 2x faster inference and up to 90% faster pre-training. Beyond serverless inference, the platform offers batch processing for large asynchronous workloads, provisioned throughput with SLA-backed capacity, dedicated single-tenant GPU deployments, and on-demand or reserved GPU clusters spanning NVIDIA H100 through GB300 hardware. For teams shaping their own models, Together AI provides LoRA and full fine-tuning, direct preference optimization, and evaluation tooling, all priced per token based on model size.
Together AI serves close to 500,000 developers and enterprises, with customers like Cursor, Decagon, Zoom, Quora, ElevenLabs, and Salesforce building production AI products on its infrastructure; Cursor uses the platform for real-time, low-latency inference at scale, while Decagon achieved 6x cost reduction per conversational turn versus GPT-5 mini. The company also contributes actively to the open-source ecosystem through datasets like RedPajama-V2 and ongoing research published at venues including NeurIPS, ICML, and ICLR, and offers a startup accelerator program to help early-stage companies build on open models.
Key Features
- Serverless inference API access to 200+ open-source text, vision, image, audio, and video models
- Batch inference for cost-effective asynchronous processing at scale
- Provisioned throughput with token-based capacity and uptime SLAs
- Dedicated model and container inference on single-tenant GPU hardware
- On-demand and reserved GPU clusters spanning H100, H200, B200, and GB200/GB300
- Fine-tuning via LoRA or full fine-tuning with supervised and DPO methods
- Sandbox environments and code interpreter for AI app and agent development
- Managed high-performance storage with zero egress fees
- Evaluations tooling to measure and compare model quality
- OpenAI-compatible API for easy migration from closed-model providers
Pros
- OpenAI-compatible API makes migrating from closed-model providers straightforward
- Transparent per-model, pay-as-you-go pricing across 200+ open-source models
- Vertically integrated GPU cloud offers competitive on-demand and reserved rates
- Backed by deep systems research, including FlashAttention and other efficiency breakthroughs
- Full-stack coverage from inference to fine-tuning to raw GPU compute in one platform
- Proven at scale with customers like Cursor, Zoom, Quora, and ElevenLabs
Cons
- Pricing spans many separate model and product pages, making total cost estimation more complex than flat-rate competitors
- Dedicated GPU and reserved cluster pricing largely requires contacting sales rather than transparent self-serve rates
- Focus on open-source models means access to closed frontier models like GPT or Claude isn't the platform's core strength
- Fine-tuning costs vary significantly by model size and technique, requiring careful comparison before committing
- Provisioned throughput and PTU-based pricing has a learning curve for teams new to capacity-based billing
Pricing
Together AI uses pay-as-you-go pricing across its products. Serverless inference is billed per model, priced per 1M tokens for text (e.g., MiniMax M3 at $0.30 input/$1.20 output, GLM-5.2 at $1.40 input/$4.40 output, gpt-oss-120B at $0.15 input/$0.60 output), per image for image generation (e.g., FLUX.1 [schnell] at $0.0027/image), per video for video models (e.g., ByteDance Seedance 2.5 at $0.115/video, Google Veo 3.0 at $1.60/video), and per audio minute or character for speech models. Dedicated Inference runs on single-tenant GPUs starting at $5.49/GPU/hour on-demand for NVIDIA HGX H100 and $8.99/hour for HGX B200, with reserved options available via sales. GPU Clusters offer on-demand rates from $3.99/hour (H100) to $8.19/hour (B200), with reserved pricing dropping as low as $3.19/hour for 181+ day H100 commitments. Sandbox compute costs $0.0446/vCPU/hour and $0.0149/GiB RAM/hour, with Code Interpreter sessions at $0.03 per 60-minute session. Fine-tuning is priced per 1M tokens processed, ranging from $0.48 (LoRA, up to 16B parameters) to $8.00 (full fine-tuning, 70-100B parameters) for standard models, with specialized model pricing (e.g., DeepSeek-R1, GLM-5) ranging $5-$40 per 1M tokens plus a minimum job charge. Managed Storage costs $0.16/GiB/month.
Claim Verified Creator Badge
Are you the founder of Together AI? Display this listing's verified badge on your website to show your customers that your product has been vetted and listed on AI Central Resources.
<a href="https://www.aicentralresources.com/tool/together-ai" target="_blank" rel="noopener"> <img src="https://www.aicentralresources.com/badges/featured-badge-dark.svg" alt="Featured on AICentralResources" width="200" height="54" style="border: none;" /> </a>
* Place this HTML snippet in your website's footer, landing page, or press section. This creates a search-friendly backlink directly to your verification page.
Connect with Together AI
Frequently Asked Questions
Together AI is a full-stack AI cloud platform, often called the 'AI Native Cloud,' that provides serverless and dedicated inference, GPU clusters, fine-tuning, and systems research to help developers build and scale AI applications on open-source models.
Together AI's serverless inference is pay-as-you-go and priced per model, for example MiniMax M3 costs $0.30 per 1M input tokens and $1.20 per 1M output tokens, while GPU clusters start around $3.99/hour on-demand for NVIDIA HGX H100.
Yes, Together AI's serverless inference API is OpenAI-compatible, making it straightforward to switch existing applications from closed models to open-source alternatives with minimal code changes.
Yes, Together AI supports both LoRA and full fine-tuning using supervised fine-tuning or direct preference optimization, with pricing based on model size and the volume of tokens processed during training.
Together AI was founded in June 2022 by Vipul Ved Prakash, Ce Zhang, Chris Ré, and Percy Liang, and has grown to serve nearly 500,000 developers and enterprises including Cursor, Zoom, and Quora.
Yes, Together AI operates its own vertically integrated GPU cloud with clusters of NVIDIA H100, H200, B200, GB200, and GB300 hardware, available on-demand or as reserved capacity for 7 days to 180+ days.





