The first phase of the AI boom focused on building large models. Companies spent billions of dollars on chips, data centers, research teams and training runs. The main question was simple: who could build the most capable model?

That question still matters, but another market has grown beside it. Once a model exists, it must answer questions, write code, create images, process documents and support software products every day. That work is called inference. It means a trained model takes new information and produces an answer or prediction.

Inference now looks like a major business on its own. Fireworks AI, Baseten, Together AI, Modal, Groq and OpenRouter have all built large businesses around different parts of this market. Their recent funding rounds and growth figures show how quickly investors and customers have started to value the infrastructure that sits between AI models and end users.

The scale of demand also keeps rising. One recent market tracker measured 82.7 trillion tracked AI inference tokens per week as of September 13, 2026, up 18% in 30 days and 252% since May 2026. Open-source models accounted for 82.7% of that tracked volume, with an effective price near $0.92 per million tokens.

This creates a simple but important shift. The value of AI no longer sits only inside the model. A large part of the value now sits in the system that serves that model quickly, cheaply and reliably.

Why inference needs its own infrastructure

Training and inference may use the same chips, but the business problems differ.

Training focuses on building model intelligence. A company can spend months on a huge training run and then release the model. Inference starts after that point. It must handle real customer requests at any hour, with changing traffic and different response times.

A customer may send a short question at one moment and a very long document at the next. Another customer may need an answer in milliseconds. An AI coding tool may send a large number of requests in a short period. An agent may call several models during one task.

That creates a difficult infrastructure problem. The provider must select the right hardware, keep enough capacity available, control costs, reduce delays and maintain reliable service.

The market therefore has a new economic unit: the token.

A cloud provider once sold compute hours. An AI inference provider can sell access to model output at a price linked to tokens. This creates a market where speed, model quality, hardware efficiency and price all matter at the same time.

OpenRouter has pushed this idea even further. The platform lets developers access hundreds of models through one interface and route requests across different providers. Stripe agreed to acquire OpenRouter in August 2026. Reuters reported a deal value slightly above $8 billion, while OpenRouter had more than 10 million developers and companies, more than 400 models and over 10 trillion tokens of daily volume.

That deal offers one of the clearest signs that inference has become a separate technology market.

Fireworks AI shows how large the opportunity can become

Fireworks AI now sits near the center of this market. In July 2026, the company announced a $1.505 billion Series D at a $17.5 billion valuation. At the same time, it reported more than $1 billion in annualized revenue and more than 40 trillion tokens served each day.

The more unusual figure sits inside that token number. More than 95% of Fireworks’ daily tokens come from models specialized on customer data and specific jobs.

That detail changes the picture. Fireworks is not only selling access to general AI models. It helps companies create and serve models that fit their own products, data and workflows.

A legal company may need a model that works well with legal documents. A coding company may need a model tuned for software development. A business may want a smaller model that handles one task at a much lower cost than a large frontier model.

Fireworks sits in the middle of that process. It gives companies tools for customization, evaluation, fine-tuning and inference. Its customers can then serve those models at large scale.

The company therefore treats inference as part of a larger system for what it calls specialized intelligence. That idea could become one of the strongest parts of the AI infrastructure market.

Baseten is turning inference into a major platform

Baseten has followed a similar path, with a stronger focus on model deployment and production infrastructure.

In June 2026, Baseten raised $1.5 billion in Series F at a $13 billion valuation. The company said its revenue had grown 20 times over the previous year, while inference volume had grown 40 times. Its customer list includes Cursor, Notion, Lovable, Harvey, HubSpot, OpenEvidence, Abridge and Decagon.

Those numbers show how enterprise AI demand can translate into infrastructure demand.

A company does not need to build a new foundation model to create a large inference bill. It can use an existing model inside a coding product, customer support system, research tool or internal workflow. Each customer interaction then creates another inference request.

As AI software reaches more users, the number of model calls rises. Baseten benefits from that growth without needing to own the application itself.

This creates an attractive position. The company can serve many AI products at once, while each customer adds more traffic as its own product grows.

Together AI is building a broader AI cloud

Together AI takes a wider approach. The company combines inference with model training, fine-tuning, GPU infrastructure and other parts of the AI stack.

That gives Together AI a different advantage. Customers can use the same platform for more than model serving.

The company raised $800 million at an $8.3 billion valuation in July 2026. Reports placed its annual bookings above $1.15 billion, although that figure includes products beyond inference. This makes direct comparison with Fireworks and Baseten difficult, but it also shows the size of the wider AI infrastructure opportunity.

Together AI also benefits from the growth of open models. Open models can give companies more control over cost, customization and deployment. As their quality improves, demand can move away from expensive closed systems for some workloads.

That does not mean closed models will disappear. It means the inference market can support many model types at the same time.

Groq shows why hardware still matters

Software alone does not define inference. Hardware remains critical.

Groq has built a specialized inference cloud around its own technology. The company raised $650 million in June 2026 and another $350 million in August. The August round valued Groq at $3.5 billion. Groq said its network had reached 13 data centers and more than six million developers.

The company also has a notable relationship with Nvidia. In December 2025, Nvidia entered a non-exclusive licensing agreement with Groq. Nvidia later announced an LPX platform that incorporates Groq inference technology.

That relationship shows a wider change in the chip market. The best hardware for training a huge model may not always offer the best economics for every inference workload.

Inference needs fast response times, strong memory performance and high efficiency. Different model types may favor different hardware designs.

That opens space for companies that build chips specifically for model serving.

A new hardware race is forming around inference

The hardware shift has become even clearer with Positron.

In September 2026, Positron raised $875 million at a $5 billion valuation. Its Asimov processor uses a memory-first design with up to 2.3 terabytes of memory per chip. The company targets AI inference and plans to place the processor inside its Titan server systems.

This shows why memory has become such a major issue.

Large AI models need fast access to huge amounts of data while they produce responses. A chip can have enormous computing power and still struggle with a workload if memory access becomes the limiting factor.

That creates room for new hardware designs. Nvidia remains the dominant force, but Groq, Positron, Cerebras, Etched and other companies now compete for specific inference workloads.

The result looks less like one chip winning everything and more like a market with several types of specialized hardware.

The router may become as important as the server

OpenRouter points to another layer above infrastructure.

A developer does not always need to choose one model and one provider forever. Different requests can go to different systems.

A simple request may use a cheap small model. A difficult reasoning task may use a stronger model. A coding request may go to a model with strong software performance. A private enterprise task may require a particular region or deployment type.

Software can make those decisions automatically.

That gives model routers a powerful role. They can compare price, speed, availability and model quality, then send each request to the most suitable provider.

Stripe’s acquisition of OpenRouter shows that this layer has serious commercial value. Stripe also has an interest in AI billing and token usage, which makes the combination especially logical. AI companies need both model access and a way to measure and charge for that consumption.

Inference is also becoming a global infrastructure problem

The location of inference now matters.

An enterprise may want a model close to its customers to reduce delay. It may need data to stay within a certain country or region. It may want access to several hardware providers rather than depend on one cloud.

Equinix, Nvidia and Together AI announced Equinix Inference Exchange in September 2026. The service combines Equinix infrastructure, Nvidia reference architectures and Together AI’s inference platform. The goal centers on faster enterprise deployment, lower latency, flexibility and better cost control across models, providers and regions.

That development points to a larger market structure. Inference now touches chips, data centers, networking, model platforms, routing and enterprise software.

The infrastructure chain keeps getting deeper.

The biggest challenge is commoditization

The market still carries a major risk.

If many companies offer similar models, similar hardware and similar deployment tools, price can become the main difference. Customers may switch providers whenever another company offers lower token prices or faster service.

Inference companies also face high infrastructure costs. A provider needs capacity available even when demand changes. GPU leases, data center costs, power and networking can reduce margins.

This concern has already appeared among investors. Newcomer reported in May that some investors questioned whether inference startups could keep strong margins while they lease GPU capacity and compete with hyperscalers that have deeper infrastructure resources.

That makes differentiation essential.

A company needs more than a cheap API. Stronger positions may come from proprietary hardware, better scheduling, custom model tools, enterprise integration, routing technology or deep customer relationships.

Why open models could make the market bigger

Falling token prices do not automatically mean a smaller inference market.

Cheap inference can encourage more software use. A company may avoid a costly AI feature at a high token price, then add that feature once the cost falls enough.

The same pattern can appear with agents. A single task may require many model calls. If each call becomes cheaper, developers can afford more complex software.

Tracked market data supports this view. Open-source models accounted for 82.7% of tracked inference tokens in September 2026, while their demand-weighted effective price sat near $0.92 per million tokens.

Lower prices can therefore create more consumption.

That creates an unusual market dynamic. Inference companies can lose revenue per token while gaining far more tokens.

The next stage could center on specialized intelligence

The strongest long-term opportunity may sit between general models and end-user software.

Companies want AI that understands their own data, workflows and customers. They may not need the largest model for every task. They may need a smaller, specialized system that performs one job extremely well at a lower cost.

Fireworks already reports that more than 95% of its daily token volume comes from models specialized for customer data and specific tasks.

That model could spread across industries.

Coding, legal work, customer service, healthcare, research and finance all have different needs. A single general model can serve them, but specialized models can offer better cost, control and performance for specific jobs.

Inference startups can sit directly inside that transition.

The market now has several layers

The emerging AI infrastructure market has a clear structure.

Chip companies provide the physical compute. Cloud and neocloud companies provide access to that compute. Inference platforms handle deployment and model performance. Routing platforms choose where each request should go. Application companies turn those capabilities into products.

No single layer needs to own the entire stack.

That creates several possible winners.

Fireworks has strong scale and specialized-model demand. Baseten has strong enterprise growth and production deployment. Together AI has a broad infrastructure strategy. Groq has specialized hardware and inference cloud capacity. OpenRouter has built a major routing marketplace. Positron and other chip startups now target the hardware layer with designs made for inference.

Serving AI may become bigger than building AI

The central idea behind the new market is simple. A model only creates value when software can use it.

Training creates the intelligence. Inference delivers that intelligence to a customer.

As AI products gain users, every interaction creates another request. As agents become more capable, each task can create many requests. As model prices fall, more companies can afford to place AI inside everyday software.

That creates a large recurring infrastructure market.

The latest funding numbers make the shift hard to ignore. Fireworks has reached a $17.5 billion valuation, Baseten has reached $13 billion, Together AI has reached $8.3 billion, Groq has reached $3.5 billion, Positron has reached $5 billion, and OpenRouter has attracted a reported acquisition price above $8 billion.

The AI industry therefore has a new question alongside model quality: who can serve intelligence at the right price, speed and scale?

That question has already created a new market. The next stage of AI may depend less on who builds the biggest model and more on who builds the best system for serving millions of models, billions of requests and trillions of tokens.

Also Read – Open-Source Startups: How Free Software Becomes a Business

By Arti

Leave a Reply

Your email address will not be published. Required fields are marked *