Liquid Inference is an LLM router that runs an instant auction for every AI request. You send an ordinary chat request, and providers compete to answer it at the lowest price that meets the rules you set. Instead of committing to one vendor's rate card, you describe what the request needs — model, region, speed and minimum quality — and the cheapest provider that satisfies those requirements answers. The product is built for developers and teams already working with large language models through coding agents, SDKs or direct API calls, and who want access to hundreds of open- and closed-weight models without managing a separate account, key and billing relationship for each one.
Choosing a model provider has usually meant accepting a published, fixed rate and paying it whether or not cheaper capacity exists somewhere else at that exact moment. Inference capacity fluctuates with supply and demand, so providers often hold spare capacity they would be willing to sell at a lower price, but buyers have had no practical way to discover and capture that price request by request. Liquid Inference turns every prompt into a small auction to close that gap. Providers publish a price per token for each model and can change it at any time, lowering it when they have spare capacity. The result is that a buyer pays the lowest marginal cost available for that request while still getting the model, region, speed and quality constraints they specified.
The auction runs in three steps. First, providers set their prices: each provider sets a price per token for each model and may change it whenever they like, so the offers on the book move with supply and demand. Second, the cheapest match answers: you set the rules — model, region, speed and minimum quality — and the cheapest provider that meets them answers your request. Third, you pay for what you use: before the model starts, you know the most a request can cost, and you are billed only for the tokens it produces. Because the cap is fixed before the first token is generated, the price of a request is known up front rather than discovered later on the invoice.
Control over routing is explicit. A price limit is placed on every request before the first token, so you know the maximum a request can cost and pay only for the tokens produced. Routing rules let you choose the model, region, speed and minimum quality, and only providers that meet those requirements compete for the work. To support trust in the marketplace, Liquid Inference tests every provider regularly with standard benchmarks and publishes the results, so the speed and measured quality of providers are visible and the same rules apply to everyone. Billing is itemized, showing what every request cost line by line, and public price history lets you download live and past prices for every model so you can send work when prices are low.
Liquid Inference covers hundreds of open- and closed-weight models, spanning families such as GPT, Qwen, Gemini, GLM, Deepseek, Gemma, Kimi, Grok, Claude, Mistral, O, Nemotron, Ernie, Mimo, Llama, Phi, Nova, Command, Hermes, Granite, Sonar and many more. Model counts per family are listed on the site alongside the lowest current offer in USD per million output tokens, which ranged from $0.01 for Nemotron to $4.62 for O at the time of capture. Compatibility is deliberately simple: one key covers both OpenAI Chat Completions and Anthropic Messages on a single base URL, so existing code works unchanged, and getting started requires changing one URL. The product also supports full multi-modal requests.
The overall approach is to treat inference as a live market rather than a static catalogue. Providers compete dynamically on price for every single prompt, offering the lowest marginal cost at the moment the request is made. You can create routing rule presets or use the Auto routing algorithms to decide how requests are matched, and the rules you set determine which offers are allowed to compete. Because providers can raise or lower prices at will and stop quoting when they are full, the offers on the book reflect current capacity. The buyer's side of the market is equally concrete: a single base URL, one API key, and code that already speaks the OpenAI or Anthropic API continues to work.
Practically, this means a lower cost per request for equivalent work, because you are matched to the cheapest provider that meets your stated requirements rather than paying a fixed published rate. It means more predictable spend, since a price limit is fixed before the first token and you are billed only for produced tokens, with itemized billing showing what each request cost. It gives you leverage over supplier choice — model, region, speed and quality are yours to specify — while independent benchmark results keep providers comparable. And because prices move with supply and demand, public price history gives you the information needed to time workloads when prices are low.
Concrete workflows follow from those capabilities. A developer using an agentic coding tool such as Claude Code, Codex, Cursor, OpenCode, Cline or Pi can point the tool at Liquid Inference by changing the base URL, keeping their existing client while gaining access to every listed model. A team planning spend can pick a model and a monthly output volume in the cost estimator and see the lowest offer now and the estimated cost per month. A workload that is price-sensitive rather than latency-sensitive can be scheduled using public price history when prices are low. A user who needs multi-modal support can send those requests through the same single key. And anyone already running inference on their own models can join as a provider, post a price and be paid for the tokens they serve.
For buyers, sign-up is free with email and the first 500 users receive $20 of free inference; referring others earns 20% of referred fees as free inference, with 10% for second-level referrals. Access is through an API key and the base URL https://router.inference.ai.exchange/v1. The documented list of compatible clients and SDKs includes Claude Code, Codex, Pi, Oh My Pi, Cursor, OpenCode, Cline, Roo Code, Kilo Code, Goose, Open WebUI, Cherry Studio, n8n, the OpenAI SDK, the Anthropic SDK, the Claude Agent SDK, the OpenAI Agents SDK, the Vercel AI SDK, LiteLLM, LangChain and curl. On the supply side, anyone running models can create an account, register the deployment they want to serve, have an operator review and list it, then connect their agent and start quoting. Registration as a provider is free, and Liquid Inference charges its fee to the customer, not to the provider.
In short, Liquid Inference replaces fixed inference pricing with a per-prompt auction in which providers compete to serve your request at the lowest price that still satisfies your rules. You keep the models and tools you already use, gain a price cap before the first token, itemized and predictable billing, benchmark-backed provider trust and public price history, while providers gain a live market for their spare capacity. The value proposition is straightforward: complete your AI requests and pay the lowest offer that meets your requirements.