Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale
The 2026 inference market is no longer GPU time rental. It bifurcates into custom silicon (Groq, Cerebras, SambaNova), GPU platforms (Baseten, Together, Fireworks, Modal), and API-first marketplaces (Replicate, DeepInfra). Fireworks raised price $1/hr per GPU on May 1, 2026, and $4B valuation on 10T+ tokens/day tells you the volume-driven model works. Baseten closed $300M Series E at $5B in January 2026. The competitive positioning rule is simple: Fireworks optimizes latency, Together optimizes catalog breadth, Baseten optimizes enterprise polish, Modal optimizes Python-native DX, Replicate optimizes multimodal reach, Anyscale optimizes distributed Python. This lesson gives you a matrix you can hand a founder. Name the three market segments (custom silicon, GPU platforms, API-first) and map each vendor to a segment. Explain why the "per-token" API pricing model compresses toward the serving engine's cost curve, not the hardware's. Compute effective cost per request across at least three vendors and explain when per-minute (Baseten, Modal) beats per-token. Identify which platform is the right default for a given workload (serverless bursty, steady high-throughput, fine-tuned variants, multimodal). You evaluated managed hyperscaler platforms. You decided you need a narrower, faster provider — Fireworks for latency, Together for breadth, Baseten for a fine-tuned custom model. Now you have six real choices and the pricing pages do not line up. Fireworks shows $/M tokens; Baseten shows $/minute; Modal shows…
Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale: The 2026 inference market is no longer GPU time rental. It bifurcates…
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.