So the question “what can replace the ChatGPT API?” is now too broad. One project wants cheaper text generation, another needs a million-token context, a third needs image and voice processing, a fourth needs web search and tools for an AI agent, and another may care most about running a model on its own infrastructure.
Technically, OpenAI API is the more accurate term than ChatGPT API: ChatGPT is a user-facing product, while the API is a separate developer platform. But “ChatGPT API” is still widely used, so in this article I use alternatives to mean services through which modern language and multimodal models can be connected to an application.
What counts as an OpenAI API alternative?
In 2026, it is useful to divide these services into three groups.
The first is direct APIs from model developers. OpenAI provides GPT, Anthropic provides Claude, Google provides Gemini, xAI provides Grok, DeepSeek provides its own models, and so on. You work directly with the company developing the model.
The second is multi-model platforms and API gateways. OpenRouter, Together AI, Hugging Face Inference Providers, Amazon Bedrock, and others offer unified access to several model families. These are not direct alternatives to Claude or Gemini: they primarily solve an infrastructure problem involving routing, billing, resilience, and provider choice.
The third group is specialised AI APIs. Perplexity, for example, is primarily useful as a web-search and research API; Replicate provides one access point to a large number of models for images, video, audio, and other tasks; NVIDIA NIM offers a way to run models on your own infrastructure behind a familiar API.
Putting all three groups in one table and trying to name the “best model” produces a fairly unhelpful comparison. So let’s start with direct OpenAI alternatives.
Main direct alternatives to the OpenAI API
The table below is a quick guide to some of the most notable APIs worth considering when developing a new AI product. Prices are per 1 million input and output tokens for the selected model, where the provider publishes a comparable pay-as-you-go rate.
| Provider | Example current model | Context | Input / Output | What distinguishes the platform |
|---|---|---|---|---|
| OpenAI | GPT-6 Sol | 1.05 million | $2 / $10 | Responses API, reasoning, web/file search, computer use, broad tool set |
| Anthropic | Claude Sonnet 5.5 | 1 million | $2 / $10 | Coding, agents, long tasks, tool use, vision |
| Gemini 3.8 Flash | ~1.05 million | $0.75 / $3.75 through the end of 2026 | Text, images, video, audio, PDF, Google Search/Maps, code execution | |
| xAI | Grok 4.7 | 500,000 | $2 / $6 | Reasoning, web search, X search, code execution |
| DeepSeek | V4.1 Flash | 1 million | From $0.15 / $0.60 | Low cost, OpenAI/Anthropic compatibility, vision, tools |
| Mistral | Medium 3.5 | 256,000 | $1.50 / $7.50 | Open-weight ecosystem, multimodal, agents, self-hosting scenarios |
| Alibaba | Qwen3.8 Max | Up to 1 million | $2 / $6 | Broad Qwen lineup, including very inexpensive Flash models |
| Moonshot | Kimi K3 | 1 million | Depends on plan | Long context, vision, reasoning, OpenAI-compatible API |
| Z.ai | GLM-5.3 | Long context | Depends on plan | Coding, reasoning, long-horizon agentic tasks |
| Meta | Muse Spark 1.3 | Long context | Public preview | New first-party Model API, multimodality, search, and tools |
| Cohere | Command A | 256,000 | $2.50 / $10 | RAG, citations, reranking, enterprise scenarios |
| Yandex | YandexGPT 5.1 Pro | Depends on model | Regional pricing | Russian infrastructure, AI Studio, agents, search, MCP |
| Sber | GigaChat | Depends on model | Via Cloud.ru for new customers | Russian infrastructure, multimodal, functions, documents |
This table should not be read as a quality ranking. GPT-6 Sol, Claude Sonnet 5.5, Gemini 3.8 Flash, and DeepSeek V4.1 Flash occupy different price and architecture segments, and the same token count does not imply the same result. The table is mainly a snapshot of the market’s scale and the differences between providers’ approaches.
Anthropic Claude API
Claude remains one of OpenAI’s most obvious direct competitors, especially for software development, large context windows, and agent scenarios. At the time this article was prepared, the current lineup includes Claude Sonnet 5.5, Opus 5.5, and more specialised models.
Claude Sonnet 5.5 has a context window of up to 1 million tokens and a maximum response of up to 128,000 tokens. Its standard price is $2 per million input tokens and $10 per million output tokens. The more capable Opus 5.5 costs $4/$20, while Fable 5.1, designed for especially demanding reasoning and long-running agent tasks, costs $10/$50.
The API supports images as input, tool use, extended/adaptive thinking, and discounted batch processing. Anthropic is also steadily developing Claude Code and agent scenarios, so Claude is worth considering as a platform for tool-using AI systems, not just as a text generator.
One important detail for developers is that Anthropic’s Messages API differs from the OpenAI Responses API. Compatibility can be arranged through third-party gateways or infrastructure platforms, but using Claude’s native capabilities usually means accounting for Anthropic’s own model interface.
Google Gemini API
Gemini API perhaps best illustrates how outdated the term “language model” has become. Gemini 3.8 Flash accepts text, images, video, audio, and PDFs, has a context window of about 1.05 million tokens, and supports function calling, structured outputs, file search, code execution, Google Search grounding, Google Maps, and thinking.
Through December 31, 2026, Gemini 3.8 Flash has an introductory price of $0.75 per million input tokens and $3.75 per million output tokens. Rates of $1.50 and $7.50, respectively, are announced from January 1, 2027.
Its compatibility with the OpenAI SDK is also notable. Google provides a beta interface that can make it possible to switch from OpenAI in simple scenarios by changing the API key, base URL, and model name. This does not mean that every tool and parameter is fully compatible, but it lowers the cost of trying another provider.
Gemini is worth evaluating separately for products that need text, images, audio, video, search, and document processing through one API, even if their main backend already uses OpenAI.
xAI Grok API
xAI was initially seen mainly as the Grok ecosystem inside X, but its API is now a standalone platform. The current Grok 4.7 has a 500,000-token context window and costs $2 per million input tokens and $6 per million output tokens at a normal context length.
The model supports reasoning, function calling, and structured outputs. xAI tools provide web search, X search, and code execution. For requests with context longer than 200,000 tokens, Grok 4.7 pricing rises to $4/$12, which matters when comparing it with models that have a million-token context window.
xAI also has models with context windows of up to one million tokens, and the company is developing voice and image/video capabilities alongside its text API. Grok can therefore no longer be viewed only as a specialised model for working with X data, although X Search remains one of its notable features.
DeepSeek API
DeepSeek stands out above all for its price-to-capability ratio. The current deepseek-flash corresponds to V4.1 Flash, has a context window of up to 1 million tokens and a maximum output of up to 384,000 tokens. It supports JSON output, tool calls, the Responses API, and vision.
Its pricing is unusual: the rate depends on the time of day. For V4.1 Flash during off-peak hours, one million input tokens without a cache hit cost $0.15, and output tokens cost $0.60. During peak hours, the price doubles to $0.30/$1.20. Cache hits cost even less.
DeepSeek’s support for OpenAI and Anthropic formats is especially useful for migration. It does not guarantee identical model behaviour or 100% support for every field, but it lets you test DeepSeek without completely rewriting the client layer.
The low price makes the API interesting for high-volume classification, document processing, generation, and background tasks. Still, comparing DeepSeek with more expensive models only by token price is misleading: the meaningful measure is the cost of a successfully completed task at the required quality.
Mistral AI
Mistral is interesting because it combines a commercial API with a strong open-weight strategy. That matters to projects that find a hosted API convenient today but may need their own deployment or tighter infrastructure control in the future.
Mistral Medium 3.5 has a 256,000-token context window and costs $1.50/$7.50 per million input and output tokens. Mistral Large 3 is less expensive at $0.50/$1.50, and Small 4 costs $0.15/$0.60. The models offer structured outputs, function calling, Document Q&A, batch processing, and agent tools.
Mistral has another practical advantage: some models are available with open weights. So the choice is broader than “which API should I call?” The same ecosystem can be accessed through a cloud service, a third-party inference provider, or your own infrastructure.
Alibaba Qwen and Model Studio
Qwen has long since moved beyond being “just another Chinese model”. Alibaba Cloud Model Studio offers many models of different sizes and specialisations, including Qwen3.8 Max and Qwen3.8 Flash.
Qwen3.8 Max supports up to 1 million context tokens and costs $2/$6 per million input and output tokens through the international API. Qwen3.8 Flash is much less expensive at $0.15/$0.47. Some models use tiered pricing based on the length of the input context.
Qwen is especially interesting when a developer needs a family of models rather than one all-purpose model: from inexpensive high-volume processing to large reasoning and multimodal tasks. Open Qwen models are also widely available through third-party inference providers and can be self-hosted.
Moonshot Kimi
Moonshot AI develops the Kimi family, and its current flagship Kimi K3 has a context window of up to 1 million tokens, native image processing, reasoning, and structured output using JSON Schema. Its API is compatible with OpenAI clients, so Kimi is relatively easy to add as an alternative backend.
One interesting feature of K3 is its high max_completion_tokens limit, which makes the model suitable for genuinely long generations and multi-step tasks. It supports custom tools and automatic caching.
Individual platform features, however, develop independently of the model itself. For example, at the time this article was prepared, web search for K3 was being updated, so a production system should not assume that all functionality from older Kimi APIs is equally available in the new stack.
Z.ai and GLM
Z.ai develops the GLM family, and GLM-5.3 is positioned primarily for coding, reasoning, and long agentic workflows. The model supports several reasoning-effort levels, with thinking built into its normal operation.
GLM is also appearing at third-party inference providers, making the ecosystem interesting even without a direct link to Z.ai. For example, GLM-5.3 and Flash variants are available through several multi-model platforms.
Direct price comparisons are difficult because Z.ai’s public pricing changes and is often presented less transparently than OpenAI or Anthropic pricing. So it would be misleading to list a third-party host’s price as “the GLM API price”. For a specific project, compare each access method separately.
Meta Model API
Meta was long known mainly as the provider of open Llama models, which people ran through third-party providers. That changed in 2026, when Meta introduced its own Meta Model API, now in public preview.
The current Muse Spark 1.3 is aimed at long-horizon agentic coding and multimodal tasks. The platform offers an OpenAI-compatible interface, structured output, parallel tool calling, built-in search, and multimodal inputs. Alongside its main model, Meta is developing separate Muse Voice and Muse Image APIs.
For now, this is a fast-developing new platform rather than an established commercial API like OpenAI, Google, or Anthropic. But Meta’s first-party API matters: choosing Meta models no longer necessarily starts with finding a third-party host.
Cohere: a special focus on RAG
Cohere is more useful to consider as an enterprise and RAG-focused platform than as a universal OpenAI replacement “for everything”. Command A has a 256,000-token context window, supports tool use, structured outputs, reasoning, images, and citations, and costs $2.50/$10.
The ecosystem’s strength is not limited to its generative model. Cohere has separate Embed and Rerank models, making the service a good fit for architectures where much of the value comes from searching corporate data, ranking documents, and producing answers with sources.
That advantage may not matter for a project that only needs a general content generator. If a product is built around complex searches through an internal knowledge base, comparing Command A’s price with GPT or Claude alone would be too superficial.
Yandex AI Studio and GigaChat for Russian projects
Projects focused on Russian infrastructure, local payment, and local data processing should also consider Yandex AI Studio and GigaChat. They are not merely “Russian models”, but their own platforms with their own tools and ecosystems.
Yandex AI Studio brings together Model Gallery, YandexGPT, Agent Atelier, AI Search, and MCP Hub. In 2026, Yandex expanded the platform’s pricing: different token types, caching, and agent tools are billed separately, so the old comparison of a single price per thousand tokens no longer fully describes the cost of a complex workflow.
GigaChat supports multimodal work, function calling, batch processing, and document handling. From September 1, 2026, commercial access for new customers moved to Cloud.ru infrastructure, so it is better to check current prices and model availability there rather than rely on old GigaChat API pricing tables.
These platforms may not be the first choice for an international project. For a product that needs Russian infrastructure and local billing, they address organisational requirements that ordinary model benchmarks do not capture at all.
What does the same volume of tokens cost?
It is hard to interpret prices per million tokens without a real-world scale. Consider a hypothetical request with 100,000 input tokens and 10,000 output tokens. It could be an analysis of a large document set followed by a detailed report.
| API and model | Cost for this volume |
|---|---|
| OpenAI GPT-6 Sol | $0.30 |
| Claude Sonnet 5.5 | $0.30 |
| Gemini 3.8 Flash | $0.1125 |
| Grok 4.7 | $0.26 |
| DeepSeek V4.1 Flash, off-peak | $0.021 |
| DeepSeek V4.1 Flash, peak | $0.042 |
| Mistral Medium 3.5 | $0.225 |
| Qwen3.8 Max | $0.26 |
| Qwen3.8 Flash | $0.0197 |
| Cohere Command A | $0.35 |
This is not a table of the cost for the same result. The models differ in quality, reasoning, speed, performance on specific tasks, and the number of tokens they actually need. Gemini 3.8 Flash is also still on introductory pricing, while Grok becomes more expensive for contexts over 200,000 tokens.
Even so, the figures make one point clear: API prices can differ by an order of magnitude, not just by a few percent. Choosing a model for a production product based only on a general impression from a public chatbot is increasingly hard to justify.
Why token price is a poor metric on its own
An inexpensive model may need more attempts, a longer prompt, an extra result check, or a follow-up call to a stronger model. A more expensive model may complete the task in a single request. In that case, the “expensive token” is cheaper at the business-process level.
Beyond the base rate, account for caching, reasoning tokens, higher prices for long contexts, built-in web search and other tools, retries, rate limits, and response latency. In an agentic scenario, a single user request may generate dozens of internal model and tool calls.
For production, measure the cost of a successful workflow instead. For example, calculate the average cost to handle one customer enquiry, review one contract, or prepare one product for publication at the required quality.
OpenAI-compatible APIs are becoming a standard of their own
One notable trend in 2026 is the spread of interfaces compatible with the OpenAI SDK. DeepSeek, Kimi, and Meta Model API offer this mode; Google provides its own compatibility layer, and some infrastructure products support both the OpenAI Responses API and Anthropic Messages API.
In practice, this means an application can be made much less dependent on one provider. For an initial experiment, it may indeed be enough to change base_url, the API key, and the model name.
But “OpenAI-compatible” does not mean “fully interchangeable”. Providers differ in reasoning parameters, caching schemes, tool use, structured output, context limits, image handling, and built-in tools. The more complex the application, the less likely it is that migration will come down to one line of configuration.
A sound approach is to keep business logic separate from the LLM layer, maintain your own set of evaluation tasks, and test a new provider on real scenarios before switching production traffic.
Multi-model platforms: when the gateway matters more than the model
A direct API has one simple advantage: fewer intermediate layers and access to all the model developer’s capabilities. It also creates dependence on one provider. If a product uses several models or needs to switch quickly between them, aggregators and cloud AI platforms become worth considering.
OpenRouter
OpenRouter is one of the more straightforward options: a single API provides access to more than 500 models and over 80 providers. You can choose a specific inference provider, use automatic routing, and set up a fallback.
OpenRouter charges a 5.5% platform fee on its standard plan. The convenience of unified billing and routing therefore comes on top of the models’ own prices.
This approach is useful when an application needs to compare models easily or withstand an individual provider outage. If a product depends heavily on unique OpenAI Responses API or Gemini features, however, an extra gateway may limit access to specialised capabilities instead.
Together AI and Fireworks AI
Together AI and Fireworks AI sit between simple API aggregators and full inference infrastructure. They provide access to many open-weight models while also offering dedicated deployments, fine-tuning, and more controlled hosting.
Together provides models such as Qwen, Kimi, GLM, DeepSeek, and others. Fireworks also combines serverless inference with fine-tuning and dedicated deployments.
These platforms are especially useful when a project wants to work with open models but is not ready to operate GPU infrastructure itself. As demand grows, it can move from serverless to a more predictable dedicated deployment.
GroqCloud and Cerebras
Groq and Cerebras compete on inference speed rather than on the number of built-in business tools. That is a separate criterion and becomes critical for real-time interfaces, coding assistants, and agentic loops with many sequential calls.
For example, Groq reports about 500 tokens per second for GPT-OSS 120B and about 1,000 tokens per second for the 20B variant. Cerebras reports around 3,000 tokens per second for GPT-OSS 120B.
Advertised speed cannot automatically be applied to every prompt and workload: actual latency depends on input length, queues, and the model. Still, these platforms are worth testing separately when response latency is part of the user experience rather than just a technical metric.
Hugging Face Inference Providers
Hugging Face Inference Providers turns the familiar Hugging Face ecosystem into a unified layer over several inference providers. You can use hundreds of models through one interface and choose a specific backend or a strategy such as fastest or cheapest.
Hugging Face says it adds no markup to the provider’s price. This makes the platform convenient for experimenting with open models and for projects that already use Hugging Face Hub as their main model catalogue.
Cloudflare AI Gateway and Workers AI
Cloudflare offers two related approaches. Workers AI runs models directly on Cloudflare infrastructure, while AI Gateway can sit in front of OpenAI, Anthropic, Google, and other APIs.
In this setup, Gateway handles logging, caching, rate limiting, routing, and unified billing. With Unified Billing, Cloudflare charges a 5% fee on credits purchased through Cloudflare, while the provider’s base token price is not increased.
For an application already running on Workers and Cloudflare, this layer may be simpler than separate AI infrastructure. But Gateway cannot make a weak model stronger: it is an infrastructure alternative, not a model alternative.
Microsoft Foundry
Microsoft Foundry is primarily an option for companies already using Azure. Its catalogue includes Microsoft/OpenAI, Anthropic, DeepSeek, Meta, Mistral, Cohere, Hugging Face, and other providers. Deployments can be serverless or run on managed compute infrastructure.
The main advantage is usually not the lowest token price. The point is to make models part of the same enterprise environment for access management, networking, monitoring, policies, and corporate billing.
Amazon Bedrock
Amazon Bedrock serves a similar role within AWS. The platform now supports more than 100 foundation models, including models from Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, OpenAI, xAI, Meta, Mistral, Qwen, Z.ai, and other providers.
For new applications, AWS recommends the Bedrock runtime endpoint. In addition to its own Invoke and Converse APIs, the platform supports Chat Completions, Responses API, and Messages API for compatible models. This makes migration much easier for applications already built around OpenAI or Anthropic.
Bedrock is especially logical when AI is one component of a larger AWS architecture and a unified IAM model, deployment regions, and enterprise infrastructure matter more than a direct contract with each model developer.
NVIDIA NIM
NVIDIA NIM differs from the platforms above: it is primarily a ready-made deployment layer for models on your own or otherwise controlled GPU infrastructure. NIM for LLMs provides OpenAI-compatible endpoints, including /v1/responses and /v1/chat/completions, as well as the Anthropic-compatible /v1/messages endpoint.
This option is not for everyone. A serverless API is almost always simpler for a small project. But if an organisation wants to control hardware, networking, data, and the model lifecycle, NIM makes it possible to keep a familiar programming interface while removing an external inference service from the critical path.
Replicate
Replicate is difficult to compare with OpenAI if you look only at LLMs. Its strength is a large catalogue of different kinds of models: image generation and editing, video, audio, 3D, and language models.
Replicate has more than 100 official models with always-available endpoints, stable API schemas, and predictable pricing, plus thousands of public models. For a project that needs different media models, this can sometimes be easier than integrating several specialised providers.
Specialised alternatives
Not every AI API has to replace OpenAI in its entirety. Sometimes a specialised service handles one part of a product better.
Perplexity API
Perplexity builds its API around search and working with current information. The platform offers Search API for ranked search results, Agent API for complex scenarios using tools and third-party models, and Embeddings API. Sonar API remains a separate option for web-grounded answers.
If a product’s main task is to find current sources, filter the web, and produce an answer with citations, Perplexity makes more sense to compare not only with GPT or Claude but also with a “LLM plus a separate search provider” combination.
Separate image, video, and voice APIs
OpenAI, Google, and xAI are gradually bringing multimodal features together within their platforms, but the specialised market is still here. Dozens of image and video model families are available through Replicate, fal, and other inference services; voice has separate platforms such as ElevenLabs.
The choice of an “OpenAI API alternative” therefore also depends on product architecture. Sometimes one general-purpose provider reduces the number of integrations. In other cases, a specialised model provides such an advantage in quality, speed, or price that a separate API is fully justified.
How to choose an API for a specific project
There is no universal winner, but the scenarios are fairly clear.
If a project already uses the OpenAI SDK and you want to try alternatives without a major refactor, start with providers and gateways that have an OpenAI-compatible interface: DeepSeek, Kimi, Meta Model API, the Gemini compatibility layer, OpenRouter, and some infrastructure platforms.
If the top priority is the cost of high-volume text processing, DeepSeek V4.1 Flash, Qwen3.8 Flash, and the cheaper OpenAI tiers stand out among published rates. But test the model on your own classification, extraction, or generation tasks instead of choosing from a price list.
If the product depends on long context, do not compare only the advertised million-token window. Quality near the context limit, higher long-context rates, caching, and how much context is useful to send to the model all matter.
For AI agents, other factors become critical: reliable tool use, structured output, reasoning, state management, web/file search, and the cost of a multi-step workflow. Differences between platforms can matter more here than a few dollars per million tokens.
If self-hosting is required, look at open-weight ecosystems such as Mistral, Qwen, DeepSeek, Kimi, and other models, then choose how to run them: your own vLLM, NVIDIA NIM, Together, Fireworks, Hugging Face, or another inference provider.
For Russian projects with local billing and infrastructure, Yandex AI Studio and GigaChat remain separate candidates. For applications deeply integrated with AWS or Azure, it is often more sensible to check Bedrock or Microsoft Foundry first, because IAM, networking, and compliance may matter more than the lowest token price.
What I would build into the architecture of a new AI application
The main change in the 2026 market is not the arrival of a model that is “better than ChatGPT”. What matters more is that you no longer have to choose one provider forever.
For a new application, it makes sense to separate model access from the main business logic. An internal application layer might receive a task such as “extract data from a document” or “prepare an answer using search”, then select a backend and translate the request into the format used by OpenAI, Anthropic, Gemini, or another API.
That does not mean you need to build a complex universal router from day one. It is enough to avoid scattering one SDK’s provider-specific fields throughout the project, keep model configuration separate, and maintain a small set of real test tasks.
The next important element is evals. A dozen or a hundred representative examples from your own product are more useful than an abstract leaderboard. They let you test a new API for answer quality as well as latency, average output length, the number of retries, and total cost.
After that, choosing a model becomes much less emotional. One model can handle inexpensive high-volume operations, another can handle complex requests, and a third can be used only as a fallback. Switching providers no longer means rewriting half the application.
What has changed since the era of “one ChatGPT API”
The 2026 AI API market can no longer be reduced to a choice between GPT and one or two alternatives. Claude, Gemini, Grok, DeepSeek, Mistral, Qwen, Kimi, GLM, and Meta are developing as independent platforms; Yandex and GigaChat serve local scenarios; and services such as OpenRouter, Bedrock, and Foundry let you separate model choice from infrastructure choice entirely.
That makes the question “which API is best?” less useful. For a real product, a more important combination is how well a model handles a specific task, how much the whole workflow costs, its latency, the tools it needs, where data is processed, and how difficult it will be to change providers six months from now.
The most practical strategy is no longer to guess the market’s single winner. Build a small set of your own evaluation tasks, test several suitable models, and leave the architecture room to switch between them. The cost of that flexibility is much lower than it was a few years ago, and the choice of APIs is considerably wider.
Main official sources
- OpenAI API pricing: https://developers.openai.com/api/docs/pricing
- OpenAI models: https://developers.openai.com/api/docs/models
- Anthropic models: https://platform.claude.com/docs/en/about-claude/models/overview
- Google Gemini models: https://ai.google.dev/gemini-api/docs/models
- Google OpenAI compatibility: https://ai.google.dev/gemini-api/docs/openai
- xAI models: https://docs.x.ai/developers/models
- DeepSeek pricing: https://api-docs.deepseek.com/quick_start/pricing/
- Mistral pricing: https://mistral.ai/pricing/api/
- Alibaba Model Studio pricing: https://www.alibabacloud.com/help/en/model-studio/model-pricing
- Kimi API documentation: https://platform.moonshot.ai/docs
- Z.ai / GLM: https://z.ai/
- Meta Model API: https://ai.meta.com/
- Cohere Command A: https://docs.cohere.com/docs/command-a
- Yandex AI Studio: https://yandex.cloud/ru/docs/ai-studio/
- GigaChat API: https://developers.sber.ru/docs/ru/gigachat/
- OpenRouter pricing: https://openrouter.ai/pricing
- Together AI: https://www.together.ai/
- Fireworks AI: https://fireworks.ai/
- Groq models: https://console.groq.com/docs/models
- Cerebras Inference: https://inference-docs.cerebras.ai/
- Hugging Face Inference Providers: https://huggingface.co/docs/inference-providers/
- Cloudflare AI Gateway: https://developers.cloudflare.com/ai-gateway/
- Microsoft Foundry models: https://learn.microsoft.com/azure/foundry/
- Amazon Bedrock models: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards.html
- NVIDIA NIM API: https://docs.nvidia.com/nim/large-language-models/latest/api-reference.html
- Replicate official models: https://replicate.com/docs/topics/models/official-models
- Perplexity API: https://docs.perplexity.ai/
Need to add AI to a digital product?
We can help select an API for the task, design a reliable integration, and keep your team in control of data, spending, and releases.
Discuss AI-driven development