What you need to know
- The round — $800M Series C, closed 1 July 2026, at an $8.3B post-money valuation. That is a 2.5x jump from the $3.3B valuation Together AI carried after its $305M Series B in February 2025.
- Who wrote the cheque — led by Aramco Ventures (via its US venture arm, Prosperity7 Ventures), with Vista Equity Partners, General Catalyst, Emergence Capital, Nvidia, March Capital, Pegatron and S Ventures also participating.
- The number that matters more than the headline — Together AI says its annualised bookings crossed $1.15B last quarter, and the round arrived with more than 500MW of compute capacity separately committed by investors.
- The customer story — named accounts including Cursor and Cognition are reportedly running production traffic on open-weight models (DeepSeek, Nemotron, MiniMax, Kimi) through Together instead of closed-model APIs, citing cost reductions the company frames as 6x to 60x.
Don't read this round as "open-weight inference is now definitively cheaper" — read it as "there is now $8.3B of investor conviction that enough enterprises believe that to make it a fundable thesis." Those are different claims. The first requires you to benchmark your own workload; the second just tells you the market is moving and you should have an opinion.
The round, in detail
Together AI is a San Francisco-based company, founded in 2022 by CEO Vipul Ved Prakash alongside Stanford's Percy Liang and ETH Zürich / University of Chicago's Ce Zhang, that rents GPU capacity and layers a managed inference, fine-tuning and training API on top of it — the pitch is "run any open-weight model, from Llama to DeepSeek to Qwen, without operating your own GPU fleet." The Series C is its fourth institutional round, following a 2023 Series A led by Kleiner Perkins and last year's Series B.
| Round | Date | Amount | Valuation | Lead investor(s) |
|---|---|---|---|---|
| Series A | 2023 | $102.5M | Undisclosed | Kleiner Perkins |
| Series B | Feb 2025 | $305M | $3.3B | General Catalyst, Prosperity7 Ventures |
| Series C | 1 Jul 2026 | $800M | $8.3B | Aramco Ventures |
The valuation trajectory — 2.5x in under 18 months — is the loudest single data point in the round, but the bookings figure is the one that should carry more weight for builders deciding whether this is a durable platform or a hype cycle. Multiple outlets independently reported that Together AI's annualised bookings surpassed $1.15B in the quarter before the raise, up from a fraction of that a year earlier, driven by what the company and investors describe as a tripling of open-weight model usage industry-wide over twelve months. Bookings are forward commitments, not trailing revenue, so treat the figure as a demand signal rather than an audited financial result — but a $1.15B bookings run rate is large enough that it is not plausibly fabricated wholesale; it is corroborated across independent trade press covering the same press materials.
The 500MW-plus compute commitment is the part most likely to be under-covered relative to its importance. It is capacity pledged by investors — expected to be capitalised independently of the $800M equity cheque — earmarked to support a roughly 50-fold expansion of Together AI's infrastructure footprint over the next five years. In practical terms, that is Together AI trying to make the same supply guarantee that has historically kept large enterprises glued to hyperscalers: "we will have the GPUs when you need them, at the scale you need them." Whether a four-year-old infrastructure company can actually deliver that at hyperscaler reliability is the open question the next few quarters will answer.
Why customers say they're switching
The commercial argument behind the round rests on customer cost claims. Together AI's announcement and the investors backing it point to accounts including Cursor, Cognition and Decagon running production inference on open-weight models — DeepSeek, Nemotron, MiniMax and Kimi were named — instead of closed frontier APIs, with cost reductions the company characterises as ranging from 6x to 60x for comparable or better output quality. Emergence Capital, a participating investor, cited Decagon specifically as having cut inference costs roughly sixfold after the switch.
| Claim | Source | Status |
|---|---|---|
| $800M raised at $8.3B valuation | Company release + independent trade press | Verified — 2+ independent sources |
| Annualised bookings past $1.15B | Company release + independent trade press | Verified — 2+ independent sources |
| 500MW+ investor-committed compute | Company release + secondary coverage | Verified — 2+ independent sources |
| 6x–60x customer cost savings (aggregate range) | Company release + investor commentary | Verified as a stated claim — self-reported, not independently audited |
| Decagon's ~6x reduction specifically | Emergence Capital blog (investor in the round) | Single source with qualifier — investor-adjacent, treat as directional |
Every source for the cost-savings numbers traces back to Together AI's own announcement or a firm that just invested in Together AI. That doesn't make the numbers false — but it means nobody independent has audited the methodology, the baseline closed-model provider, or whether the comparison controls for output quality and latency. Treat "6x to 60x" as an upper-bound marketing range that reflects the best cases in the customer base, not a typical result you should pencil into your own budget.
The underlying mechanism is straightforward and doesn't need Together AI's press release to be true: open-weight models such as DeepSeek, Kimi and GLM have closed much of the quality gap with closed frontier models on coding and agentic benchmarks over the past two quarters, while token pricing for open-weight inference has kept falling as more providers compete on the same underlying weights. When two providers can serve the same open model, the only real differentiation left is price, latency and reliability — which pushes margins down for everyone in that layer, Together AI included. That's a structurally different economics story than the closed-model API business, where the model itself is the moat. We've covered the open-weight quality catch-up and the wider inference-platform landscape in more depth elsewhere — this piece is specifically about what an $800M round says about the economics, not a rehash of that ground.
The competitive layer Together AI is fighting to lead
Together AI is not alone in this space, and the round should be read against a crowded field. Fireworks AI competes almost head-on — managed inference and fine-tuning across the same open-weight catalogue. Modal Labs, which closed its own $355M Series C in 2026, sells more general-purpose serverless GPU compute rather than a model-specific API. Groq and Cerebras compete on custom silicon built for low-latency token generation rather than GPU rental economics. Baseten and DeepInfra round out the field with narrower, often cheaper, inference-as-a-service offerings. We've compared these platforms head-to-head before — see the DeepInfra vs Together vs Fireworks vs Groq breakdown for the mechanics — but the short version is that Together AI's $800M round is, by a wide margin, the largest single funding event this specific layer of the stack has seen in 2026, and it puts real distance between Together and the rest of the field on balance-sheet strength alone.
That distance matters because this is an infrastructure business with real capital intensity — GPUs, data-centre leases, power contracts — not a software business with near-zero marginal cost. DeepInfra's own $107M Series B and Modal's $355M raise show the whole category is attracting capital, but Together AI's round is roughly double the next-largest of those and comes with the 500MW compute commitment layered on top. If the inference layer consolidates the way cloud infrastructure did a decade ago, capital advantage compounds — the well-funded providers can absorb price wars that starve smaller rivals of margin.
What this means for Indian builders
For teams building in India, the round is a useful external data point rather than a reason to switch providers overnight. India already has domestic momentum in this exact layer — Neysa's $1.2B Series B is effectively India's version of the same GPU-cloud thesis, and the IndiaAI Mission's subsidised GPU-hour programme gives domestic builders a lower-cost route to compute that doesn't depend on any single US-based provider's pricing. The practical read for a Bengaluru or Pune team evaluating inference vendors: Together AI's scale and the broader wave of capital validate that open-weight inference is a durable category worth building on, but your actual vendor decision should weigh Together AI, Fireworks, DeepInfra and India's own emerging GPU-cloud options on your specific latency, data-residency and cost requirements — not on whose funding round made the bigger headline. Data residency in particular deserves a second look given DPDP Phase 2's cross-border transfer rules, which can tilt the calculation toward a domestic or India-region deployment even when a US provider is nominally cheaper per token.
What this means for UK builders
The UK angle is less about a single competing vendor and more about capital flow. London's AI infrastructure scene — the DeepMind-alumni wave documented in 112 DeepMind alumni startups and reflected in the record month covered in the UK's first £1bn funding month — has produced plenty of model and applications-layer companies but comparatively few infrastructure plays at Together AI's scale. That's partly a market-size issue: US hyperscalers and inference specialists have a much larger addressable customer base to underwrite $800M rounds. For a UK team, the practical takeaway is the same due-diligence exercise as for Indian builders — benchmark cost per token and latency against your workload, and weigh EU/UK data-residency requirements under the EU AI Act's transparency obligations, which is exactly the kind of migration-cost UK teams should factor in before assuming an open-weight switch is a pure win.
Run a two-week shadow test before committing production traffic: mirror 5-10% of real requests to an open-weight model via Together AI (or a comparable provider) alongside your existing closed-model API, compare cost-per-request, p95 latency and your own evaluation harness's quality score, then decide. This costs a few hundred dollars in API spend and tells you far more than any vendor's funding announcement.
Migrating a production workload wholesale to an open-weight model because a competitor's funding round made headlines. Cost-per-token comparisons that don't control for prompt length, output length, retries and fallback-to-closed-model rates routinely overstate savings — the "60x" end of Together AI's range is almost certainly an outlier case, not a typical outcome.
Every article here is written by a Verified Builder. Want your name on the next one?
AI Tech Connect lists AI engineers, founders and researchers across India and the UK — and the people hiring browse it to find them. Adding your profile is free.
Become a Verified Builder →So — what should you actually do?
Three moves, depending on where you are today:
- If you're already spending five figures a month or more on a closed-model API for a workload with predictable, high-volume traffic — coding assistants, customer-support agents, batch summarisation — run the shadow test above. This is the exact profile of workload where Cursor- and Cognition-style savings are most plausible, because volume amortises the integration cost.
- If you're pre-product-market-fit or low-volume, the migration overhead (prompt re-tuning, eval re-baselining, fallback handling) usually isn't worth it yet. Stay on a closed API for now and revisit once volume justifies the engineering time.
- If you're choosing an inference vendor from scratch, Together AI's balance sheet strength is now a genuine advantage — a well-capitalised vendor is less likely to disappear or spike prices mid-contract — but weigh it against Fireworks, DeepInfra, Modal and, for India/UK teams specifically, data-residency and compliance requirements that a US-based vendor may not solve as cleanly as a regional one.
Primary coverage: Together AI's Series C announcement via Businesswire and TechCrunch's independent report.