✨ Strapi MCP is now Generally Available - let your agents manage your Strapi content ✨

Ecosystem●16 min read

Top 7 AI APIs for Developers: Technical Comparison for Full-Stack Teams

●January 22, 2026●Updated on September 14, 2026
Top 7 AI APIs for Full-Stack Developers

If you are picking an AI API for a full-stack app, here are the current per-token prices, context windows, rate limits, service-level agreements (SLAs), and SDK retry behavior for seven providers, plus how to wire whichever you choose into Strapi 5.

Treat the AI API you pick for a chatbot, summarizer, or image-tagging feature as a planning decision for your monthly bill, latency floor, and migration schedule for as long as the model you pinned stays supported. Plan for these constraints before the integration ships: model retirements and rate-limit tiers can require later changes. Model churn is now the norm: API access was the most common release type for notable 2025 models, 47 of 102.

This guide compares seven AI APIs for developers building full-stack applications: OpenAI, Anthropic, Google Gemini, Amazon Bedrock, Azure OpenAI (Microsoft Foundry), Cohere, and Mistral. Each section covers current models, per-token pricing, context windows, and the operational details that differ most: rate limits, SDK retry behavior, deprecation cadence, and terms, then sections on shared patterns and on wiring your pick into a headless CMS like Strapi 5.

In brief:

1. OpenAI API: GPT-5.6 Family and GPT-6 Astra

OpenAI's flagship models are now GPT-6 Astra and the GPT-5.6 trio (Sol, Terra, Luna). All four share a 1,050,000-token context window and 128,000-token max output. GPT-4o and GPT-4o-mini still answer API calls at 128,000 tokens but are no longer flagship, and the gpt-4o-2024-05-13 snapshot shuts down October 23, 2026 (deprecations page) with gpt-5.6-sol as the recommended replacement.

Per million tokens (pricing page): gpt-5.6-luna costs $0.20 input and $1.20 output. gpt-5.6-terra costs $2.00 and $12.00, and gpt-6-astra $10.00 and $50.00. Sol's $4.00/$20.00 rate is promotional through November 21, 2026. Large-context variants of GPT-5.6 run about 2× standard rates, and regional data-residency endpoints add a 10% uplift for models released on or after March 5, 2026.

Computed from those rates, a chatbot moving 10M tokens a day split evenly between input and output costs $7.00 per day on Luna and $70.00 on Terra. Prompt caching is enabled by default, discounts cached input "up to 90%" depending on model, and kicks in at 1,024 tokens on GPT-5.6. Cache writes cost 1.25× the uncached input rate. The Batch API halves both input and output prices for jobs that can wait 24 hours.

Rate limits run across six usage tiers: Tier 1 unlocks at $5 paid spend with a $100 monthly cap; Tier 5 requires $1,000 paid and allows $200,000 per month (rate-limits page), and the x-ratelimit-remaining-requests and x-ratelimit-reset-requests headers report your headroom. The official rate-limit page lists dollar thresholds only; the 30-day waiting period that circulates in older guides has no source. The 99.9% uptime SLA is a Scale Tier feature, not standard.

For populating CMS fields, OpenAI's Structured Outputs "ensures the model will always generate responses that adhere to your supplied JSON Schema, so you don't need to worry about the model omitting a required key."

Tool definitions and results count as billable input tokens, so long schemas cost money on every call. The openai Node SDK retries 408, 409, 429, and 5xx responses twice by default. If you want embeddings plus generation against Strapi content, the walkthrough on RAG-based search with OpenAI covers the full pipeline.

Compliance covers SOC 2 Type 2, ISO 27001:2022, ISO 27701:2019, and a HIPAA BAA for qualifying customers, and API data is not used for training by default, though abuse-monitoring logs persist up to 30 days even with Zero Data Retention.

2. Anthropic Claude API: Claude 5 Models and Deprecation Cadence

Anthropic's current lineup splits by context size. Claude Fable 5.1, Opus 5, and Sonnet 5 offer 1M-token windows out of beta with no special header. Claude Haiku 4.5, Opus 4.5, and Sonnet 4.5 stay at 200K tokens. The 1M models can emit up to 128K output tokens in a single synchronous request.

Standard pricing effective June 29, 2026: Haiku 4.5 at $1.00 input / $5.00 output per million tokens, Sonnet 4.5 at $3.00 / $15.00, Opus 4.5 at $5.00 / $25.00. Cache reads cost 0.1× the input rate ($0.10 per million on Haiku). The Batch API discount is 50%, which is where the often-quoted $0.50/$2.50 Haiku figure comes from. That is batch pricing, not the standard rate. Caching and batch discounts stack.

Prompt caching charges 1.25× input for a five-minute cache write or 2× for a one-hour write. Minimum cacheable length varies: 512 tokens on Fable 5.1 and Opus 5, 1,024 on Sonnet 5 and Sonnet 4.5, 4,096 on Opus 4.5 and Haiku 4.5. That last number matters if you planned to cache a short system prompt on Haiku; it won't qualify.

Rate limits use named tiers: Start ($500 monthly cap), Build ($1,000), Scale ($200,000), and Custom. On Start, Haiku 4.5 allows 1,000 requests, 2,000,000 input tokens, and 400,000 output tokens per minute. Headers anthropic-ratelimit-requests-remaining and retry-after report state, and the API distinguishes a 429 from a 529 overloaded_error, which deserve different retry behavior.

Plan for Claude deprecations as an operational cost. Sonnet 4 and Opus 4 retired June 15, 2026; Claude 3 Haiku retired April 20, 2026 (deprecation schedule). Opus 4.5 is safe until at least November 24, 2026 and Sonnet 4.5 until at least September 29, 2026, with a minimum 60 days' notice. Pin model IDs in config, not code. There is no published SLA for standard access; the Priority Tier's 99.5% target exists on paper, but capacity commitments "are no longer available for purchase."

The @anthropic-ai/sdk TypeScript package exposes anthropic.messages.stream(), which returns a MessageStream helper for SSE events. For schema-driven work, the tutorial on auto-generating content types with Claude shows one practical application.

3. Google Gemini API: Uniform 1M Context and Free-Tier Data Terms

Every current Gemini model, from 2.5 Flash-Lite through the 3.8 Flash series, accepts 1,048,576 input tokens and returns up to 65,536 output tokens. Inputs span text, images, video, and audio. A document-processing example that pairs Gemini with a Strapi backend is the PDF summarizer with Gemini. The gemini-3-pro-preview model shut down March 9, 2026, so any code pinned to it needs a swap to 3.1 Pro Preview or a Flash variant.

Paid-tier pricing per million tokens: Gemini 2.5 Flash-Lite at $0.10 input / $0.40 output; 2.5 Flash and 3.5 Flash-Lite both at $0.30 / $2.50; 3.8 Flash at $0.75 / $3.75 through December 31, 2026, doubling to $1.50 / $7.50 from January 1, 2027. Gemini 2.5 Pro charges $1.25 for prompts up to 200k tokens and $2.50 above that, with $10.00 / $15.00 output. That 200k breakpoint is a pricing tier, not a free allowance.

The free tier has a data-use condition. Per the Terms of Service, "When you use Unpaid Services, Google uses the content you submit and any generated responses to provide, improve, and develop Google products and services… Human reviewers may read, annotate, and process your API input and output." Paid Services carry the opposite guarantee, and users in the EEA, Switzerland, and UK get paid-tier protections on the free tier as well. Prototype on free, but move proprietary or customer content to a paid project before any real data flows through.

Implicit caching is on by default for Gemini 2.5 and newer. Explicit caching needs 2,048 tokens on 2.5 models or 4,096 on 3.5 Flash, and storage costs $0.50 per million tokens per hour on most Flash models through December 31, 2026.

The Batch API is paid-tier only, 50% off, 24-hour target. Rate limits are per project: Tier 1 needs active billing ($250 cap), Tier 2 needs $100 paid plus three days ($2,000), Tier 3 needs $1,000 paid plus 30 days; the Gemini Developer API rate-limit documentation publishes no uptime commitment.

The Live API runs over a stateful WebSocket, and barge-in is the default: user speech interrupts the model mid-response. Sessions cap at 15 minutes for audio-only and two minutes for audio plus video; audio output on the 2.5 Flash Native Audio model bills at $12.00 per million tokens.

4. Amazon Bedrock: Multi-Vendor Access Behind One SLA

Bedrock's value is breadth under one IAM boundary. Pick it when procurement wants one contract across model vendors. The model catalog includes Amazon Nova plus models from Anthropic, Meta, Mistral, OpenAI, Cohere, DeepSeek, Google, and xAI. Nova context windows range from 128k (Micro) through 300k (Lite, Pro, Sonic) to 1M (Premier).

For the listed Claude models, Bedrock pricing matches the corresponding published Anthropic rates. Claude Haiku 4.5 on Bedrock costs $1.00 / $5.00 per million tokens and Opus 5 $5.00 / $25.00 (Bedrock pricing) via global cross-region inference. Batch inference is "50% lower price compared to on-demand inference pricing" for supported models, but prompt caching does not work with batch, so pick one discount per workload. Global cross-region inference (the global. model prefix) saves about 10% over geographic routing.

The Bedrock SLA commits to 99.9% monthly uptime with 10%, 25%, or 100% service credits as availability falls below 99.9%, 99.0%, and 95.0%. In-scope programs include SOC 1/2/3, ISO 27001, PCI DSS, FedRAMP Moderate and High, HIPAA eligibility, and GDPR. AWS PrivateLink keeps inference traffic off the public internet.

Customization now has four methods: supervised fine-tuning, reinforcement fine-tuning, distillation, and custom model import. Where the CMS itself lives is a separate call: compare self-hosted Strapi versus Strapi Cloud. Throttling and retry behavior arrives as a ThrottlingException in the SDK rather than an HTTP header, and the AWS SDK's standard retry mode applies full jitter from a 1,000 ms base. @aws-sdk/client-bedrock-runtime streams via InvokeModelWithResponseStreamCommand, which returns an AsyncIterable body.

5. Azure OpenAI in Microsoft Foundry: Privacy Guarantees and Opaque Pricing

Azure AI Foundry is now Microsoft Foundry, with the portal at ai.azure.com. The GPT-5.6 Luna, Sol, and Terra variants landed July 9, 2026 with retirement set for January 11, 2028. The retirement schedule marks the 2024-05-13 gpt-4o deprecated with an October 1, 2026 shutdown (replacement gpt-5.1).

Pricing is the weak spot for planning. The official pricing page shows placeholder values rather than numeric per-token prices and states that "Prices are estimates only and are not intended as actual price quotes," so budget from the live pricing page for your region.

Microsoft publishes explicit verbatim guarantees. Its data-privacy page states: "Prompts and completions are not used to train, retrain, or improve the base models" and "Prompts and completions are NOT available to OpenAI or other providers of Models sold by Azure." Regional deployments process within the chosen geography. Choose Azure when a written no-training-on-prompts guarantee is the blocker.

Azure guarantees availability "at least 99.9% of the time," though Developer deployments carry no SLA. Microsoft recommends Entra ID tokens over API keys for production, scoped to https://ai.azure.com/.default. Since May 7, 2026, quota is tracked at the subscription level across seven tiers, and every resource and region in a subscription draws from the same pool.

Fine-tuning covers supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement fine-tuning (RFT) across the gpt-4o and gpt-4.1 families, with RFT limited to o4-mini and invitation-only on gpt-5. Llama 2 is not on the list. By default streaming buffers until content filtering completes; Asynchronous Filter delivers token by token.

Hosting Strapi next to it is covered in Strapi on Azure. Compliance adds ISO 42001 to SOC 2 Type 2, PCI DSS 4.0, HIPAA BAA, FedRAMP High, and EU Data Boundary.

6. Cohere API: Embed, Rerank, and Command for RAG Pipelines

Cohere's Command family now includes command-a-plus-05-2026 (128k context, 64k output), command-a-03-2025 and command-a-reasoning-08-2025 (256k context), and command-r7b-12-2024 (128k).

Cohere's pricing page publishes no per-token API rates for current Command A models; the only listed API prices are legacy command-r-03-2024 at $0.50 / $1.50 and command-r-plus-08-2024 at $2.50 / $10.00 per million tokens. Get a quote before modeling costs.

Retrieval-augmented generation (RAG) is where Cohere earns its place here. Embed 4.0 returns 256, 512, 1024, or 1536-dimensional vectors (1536 default) over a 128k context, with 96 texts per call. The same embed-then-query pipeline shows up in the semantic search plugin built with Strapi and OpenAI.

Rerank 4.0 Pro and Fast have a 32k context length and a hard limit of 10,000 documents per request, though Cohere recommends 1,000 or fewer for performance. Consider Cohere when ranking quality on retrieved chunks matters more than generation.

Two gaps: no user-controlled prompt cache, and the error documentation offers "wait and try again later" rather than a documented SDK backoff, so write your own. The SaaS agreement provides services "AS IS" and "AS AVAILABLE." SDKs cover Python, TypeScript (cohere-ai), Java, and Go. For a typical content pipeline, retrieve the top 100 candidates from a vector store, pass them to Rerank, and feed the top ten to Command; the roundup of vector databases for AI applications covers the storage half.

7. Mistral AI API: Open Weights, Euro Pricing, and Self-Hosting

Mistral's portfolio pairs Apache 2.0 models (Mistral Large 3 as mistral-large-2512, Mistral Small 4 as mistral-small-2603, Ministral 3 14B at 256k, plus 8B and 3B variants) with Mistral Medium 3.5 under a Modified MIT license and the proprietary Codestral (128k). Modified MIT requires companies above $20M monthly revenue to buy a commercial license or use Mistral Studio; on-premise deployment needs a self-deployment agreement.

Documentation pricing is in euros: Mistral Small 4 at €0.12 input / €0.50 output per million tokens, Large 3 at €0.44 / €1.30, Medium 3.5 at €1.25 / €6.40, with cached input at 0.1×. The main pricing page quotes USD and may list different model generations, so pick one source when comparing against USD-priced competitors.

Self-hosting runs on vLLM (Mistral's recommendation), plus TensorRT-LLM, llama.cpp, and Ollama. The @mistralai/mistralai SDK is at version 2 and requires Node.js 18+. Use it when inference has to run inside your own network.

Check your lockfile: advisory MAI-2026-002 reports npm versions 2.2.2 through 2.2.4 and PyPI 2.4.6 were compromised.

Mistral's docs recommend custom structured outputs over JSON mode, the SDK retries 429 and 5xx with (2 ** attempt) + random.uniform(0, 1), and the Enterprise plan offers "custom SLAs" with no published uptime figure.

Integration Patterns Shared Across AI APIs for Developers

FeatureOpenAIAnthropicGeminiBedrockAzure OpenAICohereMistral
SSE streamingSSE streamingSSE streamingLive streamingAsyncIterableFiltered streamingYesSSE streaming
Tool calling and JSON SchemaStructured outputsStructured outputsStructured outputsStructured outputsStructured outputsYesStructured outputs
Prompt cachingDefault cachingExplicit or implicitImplicit defaultExplicit checkpointsDefault cachingNoPrefix sharing
Batch discount50%50%50%50% select models50%Embed Jobs, Batches50%
SDK retry built inBuilt-in retryBuilt-in retryBuilt-in retryBuilt-in retryBuilt-in retryNot documentedBuilt-in retry
Published SLAScale Tier onlyNoneNot covered here99.9% SLA99.9% SLANoneCustom SLA

Keep API keys server-side. OpenAI's guidance is to never expose API keys in browsers, and the OWASP Secrets Cheat Sheet adds encryption at rest, least privilege, and rotation. Prefer IAM roles on AWS and Entra ID on Azure over static keys. Strapi's own guide to storing API keys securely covers environment handling in a Node.js backend.

For 429 and 5xx responses, honor Retry-After first (RFC 9110 allows a date or delay-seconds), then fall back to exponential backoff. Google's recommended backoff formula is wait = min((2^n + random-fraction), maximum-backoff) with 32–64 second caps, and AWS documents full, equal, and decorrelated jitter variants. Non-idempotent requests should not be auto-retried.

If you expect to switch providers, an abstraction layer buys you time. The Vercel AI SDK covers all seven providers here; LangChain.js and LiteLLM are broad multi-provider interfaces. Tool semantics and schema enforcement stay provider-specific in all three. The tutorial on building a Strapi and AI SDK shows that pattern end to end.

Connecting AI APIs to Strapi 5

Strapi 5 (Node.js 22, 24, or 26 per the installation requirements) supports a recommended three-part integration pattern for an AI API. The correct home for the call itself is a custom service, invoked from a custom controller or a Document Service middleware.

The backend request flow passes through global middlewares, route matching, route policies, route middlewares, controllers, services, the Document Service or Query Engine, Document Service middlewares, and the response in that order, so a service holds the provider SDK, and a controller stays thin.

// src/api/article/services/summarize.js
const OpenAI = require('openai');
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

module.exports = () => ({
  async run(documentId) {
    const doc = await strapi.documents('api::article.article').findOne({ documentId });
    const res = await client.responses.create({ model: 'gpt-5.6-luna', input: doc.body });
    return res.output_text;
  },
});

Once a developer configures a webhook endpoint, Strapi webhooks send entry.create, entry.update, entry.publish, media.create, and related events as HTTP POST payloads on every plan; releases.publish requires Growth or Enterprise and review-workflows.updateEntryStage is Enterprise only.

Each entry payload carries a stable 24-character documentId plus a per-version id, and relations and media are always populated. That makes a webhook the natural trigger for an external summarization or embedding job, as the Strapi and n8n guide demonstrates. Write results back through the REST API (flat response, status=draft|published) or the GraphQL API, available when the GraphQL plugin is installed and enabled; the REST versus GraphQL comparison for Strapi 5 helps you pick.

The Strapi MCP server reached GA in 5.49.0 and is free, with no plan restriction stated in the docs. It ships disabled (server.mcp.enabled defaults to false), listens on /mcp over Streamable HTTP with a Bearer admin token, and generates up to eight tools per Collection Type (list, get, create, update, delete, publish, unpublish, discard_draft) and six per Single Type. "Available tools depend on the permissions granted to the Admin token used for the connection." The MCP GA announcement and the guide to Claude Code with MCP cover setup with Claude and Cursor.

Some AI work needs no external API at all. Strapi AI is built into Strapi 5.30+ for Growth plan customers (not available on Enterprise plans) with 1,000 monthly credits included and extra credits at $1.50 per 100. The Content-Type Builder assistant and Media Library alt-text generation (images only) are on by default; AI Translations are off until you enable them and translate one way from the default locale. For self-hosted teams without Growth, the community strapi-plugin-ai-translator plugin supports Strapi 5 and any OpenAI-compatible Chat Completions endpoint.

Pick by the Constraint That Will Actually Bind

  • Cost-sensitive, high-volume text: GPT-5.6 Luna at $0.20/$1.20 or Gemini 2.5 Flash-Lite at $0.10/$0.40, with batch for anything asynchronous.
  • Whole-codebase or archive prompts: Claude Sonnet 5, any Gemini 2.5 or 3 model, or GPT-5.6, all at roughly 1M tokens, with prompt caching to keep repeated prefixes cheap.
  • Contractual uptime and audit scope: Bedrock or Azure OpenAI, the only two here with a standing 99.9% SLA outside a premium tier.
  • Retrieval-heavy search: Cohere Embed 4 and Rerank 4.
  • Data sovereignty with on-premise inference: Mistral's Apache 2.0 models on vLLM.

Whichever you choose, either put the SDK call in a Strapi service and invoke it from a controller or Document Service middleware, or configure a Strapi webhook to trigger an external worker that calls the provider and writes results back through Strapi's API. Pin model IDs in environment config so the next retirement notice is a one-line change.

Paul BratslavskyDeveloper Advocate

Related Posts

8 Vibe Coding Prompt Techniques
Ecosystem·15 min read

8 Vibe Coding Prompt Techniques for Web Development

Learn 8 proven vibe coding prompt techniques to accelerate web development with AI. Transform scattered code requests into production-ready features.

·October 15, 2025
Lessons Learned Vibe Coding
34 min read

Building Faster with V0 and Claude Code: Lessons Learned from Vibe Coding

Building Faster with V0 and Claude Code: Lessons Learned from Vibe Coding

·September 16, 2025
Bolt.new and Strapi
Tutorials·14 min read

Build a Company Website with Bolt.new and Strapi 5 - Part 1

Learn how to use Bolt.new's AI-powered development with Strapi 5 CMS to build a Next.js company website. Step-by-step tutorial with prompts and examples.

·November 12, 2025