Hopscotch AI is a developer platform that provides access to more than 500 AI models from providers including Anthropic, OpenAI, Google, DeepSeek, Moonshot AI, Qwen, and Meta through a single API. Rather than maintaining a separate integration, account, and bill for every provider, developers send requests to one base URL with one key and one balance. The platform is aimed at developers and engineering teams building applications on top of large language models who also need to understand and control what those applications spend. Its stated purpose is simple: control AI costs and access top models through one API.
The problem Hopscotch addresses is operational sprawl. A team that wants to use several model providers usually ends up managing a different SDK, a different set of credentials, a different dashboard, and a different invoice for each one. Model quality, speed, and price change constantly, so teams want to move between Claude, GPT, Gemini, and open-weight models as their needs shift, but every switch means more integration work. Hopscotch describes itself as the intelligence layer for AI and has raised $7.5m to build it. It consolidates provider access into one connection so that choosing a model becomes a configuration change instead of a project.
The core of the product is a single API that aggregates provider access. You set a new base URL and key in the OpenAI SDK you already use, and after that switching models means changing the model name, with no new SDK to install. Model names are provider-prefixed, such as anthropic/claude-sonnet-5, and Hopscotch guarantees that the model you name is the model that runs; the Activity log shows which provider served it. Supported endpoints include chat completions and the Responses API, plus a models endpoint that lists every model you can call. A curl example and a Python quickstart are provided in the documentation.
Routing profiles let you define backups. In the example shown on the site, Sonnet runs first, then GPT, then Gemini if the earlier attempts fail. Resilience works in layers: if a provider returns a 429 rate-limit response, Hopscotch moves the request to another of its keys for that provider, then to the next model in your routing profile. During an outage, Hopscotch retries your request first, then moves to the next model on your list if you use a routing profile, and for chat requests a final attempt runs your model through a backup provider. If every attempt fails, you receive an error. The Upstreams view shows how requests are distributed across providers, such as 40% Anthropic, 19% OpenAI, 15% Google, 13% DeepSeek, and 8% Moonshot AI in the example given.
Spending controls are the other half of the product. You can give each key a credit limit that resets daily, weekly, or monthly, set monthly limits for individual teammates and for the whole workspace, and cap how fast the account can spend, $50 per minute by default. A request that would cross a limit is refused before it reaches the provider, which the Activity log records with a rejected outcome, zero attempts, and no upstream fetch. An owner can also pause all spending at once. This matters most for agents: if an automated loop runs unattended, the limits stop it from draining the balance.
Visibility comes from three connected views. Keys show who may call, Requests show what happened, and Usage shows what it cost. The Activity log lists every request with the model, the provider that served it, the outcome, the number of attempts, total tokens, cost, and latency, and it exports to CSV. Usage breaks spend down by model, provider, key, and teammate, so costs can be attributed. Your code can also look up any request's tokens and cost through the API. Logged outcomes include successful calls, client aborts, truncated responses, and requests rejected before fetch.
The playground lets you run one prompt on up to three models side by side, billed through your key like any other request, so you can compare answers and the cost of each before changing your code. The model catalog is the reference behind that comparison: each of the 500+ models lists its context window and its price per million tokens, and where several providers serve the same open-weight model, the catalog shows each provider and its price. Example entries include anthropic/claude-sonnet-5 at 2.00 in and 10.00 out per million tokens, openai/gpt-5.6 at 4.00 in and 20.00 out, and deepseek/deepseek-v4-flash at 0.44 in and 1.32 out with four upstreams.
Hopscotch does not alter your requests. Prompts reach the model as written and answers come back unchanged, and by default the platform never stores your prompts or the model's responses. Some features that providers run on their own servers, such as web search, audio, and hosted tools, are not supported through Hopscotch's provider accounts, and the docs list them. You can also bring your own provider keys: add an OpenAI key, and Hopscotch sends that provider's requests on your key, with the provider billing you directly and Hopscotch charging nothing for those requests. If your key fails, Hopscotch does not silently switch to its own key.
Taken together, the design keeps every part of multi-provider work, including access, routing, budgeting, and reporting, inside one connection. Your application speaks the OpenAI-compatible chat completions API it already uses; behind that endpoint Hopscotch holds the provider relationships, applies your routing profile, enforces your limits before a request is forwarded, and records the result. Because billing is pass-through, you pay each provider's list price per token with no markup and no added fees, and providers that bill you directly through your own keys cost nothing on the Hopscotch side. The result is that the model can change without the code, the budget, or the reporting changing with it.
The practical benefit is less integration work and more predictability. Moving between any models in the catalog, from any provider, keeps your key, balance, and limits the same, so experimenting with a new model costs one line of code rather than a new integration. Spend limits refuse over-budget requests before they reach a provider, which protects prepaid balances from runaway agents and unexpected usage. Because the Activity log and Usage views record the model, provider, tokens, cost, and speed of every request, cost questions can be answered with data rather than estimates, and CSV export makes that data usable in spreadsheets and reports.
Concrete workflows follow from those capabilities. A team evaluating a cheaper model can run a prompt on three candidates in the playground, compare answers and per-request cost, then change one line to move production traffic. A team worried about provider outages can define a routing profile, with Sonnet first and GPT then Gemini if it fails, and let Hopscotch handle retries, rate limits, and backup providers automatically. A platform owner can issue a key per environment or teammate with a monthly credit limit, set a workspace ceiling, and watch spend broken down by model and provider. A team with existing provider contracts can attach its own keys and keep paying those providers directly while still using one endpoint and one log.
Hopscotch is built for developers and engineering teams, and its Product Hunt listing appears under Developer Tools and Artificial Intelligence. Integration is deliberately minimal: any client that can point at https://api.hopscotchlabs.ai/v1 and send a bearer token works, including the OpenAI SDK for Python and curl. There are no plans or subscriptions; you add prepaid credit by card starting at $5, and auto top-up can refill your balance when it drops below an amount you choose. As a Product Hunt launch offer, the first 250 Product Hunt users to sign up receive $50 in free model credits by redeeming code HOPSCOTCH50OFF in the Billing tab.
The takeaway is a single connection that replaces many. Hopscotch AI turns 500+ models from Anthropic, OpenAI, Google, DeepSeek, Moonshot AI, and others into one OpenAI-compatible API, adds routing profiles and retries so a bad provider does not become a bad user experience, and enforces per-key, per-teammate, and workspace spend limits before requests are forwarded. You pay provider rates with no token markup, you can bring your own keys, and every request is logged with its model, provider, tokens, cost, and speed. For teams that want model choice without provider sprawl or surprise bills, that combination is the whole point.