# Nexforce Router > One OpenAI-compatible endpoint that routes to Anthropic, OpenAI, Google, DeepSeek, > Moonshot, Zhipu and Cloudflare Workers AI, with per-model fallback and per-key cost > control. If your tool already speaks the OpenAI API, it speaks Nexforce Router: swap > the base URL and the API key, keep the rest of the code. ## Instructions for AI agents - Base URL: `https://router.nexforce.ai/v1` - Auth: HTTP header `Authorization: Bearer nfc_...` (a Nexforce API key) - Protocol: OpenAI Chat Completions (`POST /v1/chat/completions`), including streaming, tools and multimodal input - Model discovery: `GET /v1/models` is public (no key). With a key, the list is narrowed to what that key may call. Prices are never exposed. - OpenAPI spec (import into Postman/Insomnia/codegen): https://router.nexforce.ai/openapi.json - Full human docs: https://router.nexforce.ai/docs ## Quick start ```bash curl https://router.nexforce.ai/v1/chat/completions \ -H "Authorization: Bearer $NEXFORCE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-3-5-sonnet", "messages": [{ "role": "user", "content": "Hello!" }] }' ``` Python (OpenAI SDK): ```python from openai import OpenAI client = OpenAI(base_url="https://router.nexforce.ai/v1", api_key="nfc_...") resp = client.chat.completions.create( model="claude-3-5-sonnet", messages=[{"role": "user", "content": "Hello!"}], ) ``` ## Endpoints - `POST /v1/chat/completions`: OpenAI Chat Completions schema. Routing to the right provider is automatic from the `model` field. Also serves image generation (see below). - `POST /v1/embeddings`: OpenAI Embeddings schema, for RAG and semantic search (see below). - `GET /v1/models`: lists available models. Each entry has `id`, `owned_by` (provider), `kind` (`chat` or `embedding`, which tells you the endpoint that serves it) and, when known, `name`, `context_length`, `max_output_tokens` and `capabilities`. Public; a valid key narrows it to the models that key may call. ## Request parameters (Chat Completions; for `/v1/embeddings` see the Embeddings section.) - `model` (required): catalog id, e.g. `claude-3-5-sonnet`. Or `nexforce/smart-route` to let Nexforce pick (see below). - `messages` (required): OpenAI array (`system`, `user`, `assistant`, `tool`). - `stream`: `true` for Server-Sent Events (OpenAI `chat.completion.chunk` deltas, terminated by `data: [DONE]`). - `temperature`, `top_p`, `max_tokens`, `stop`, `frequency_penalty`, `presence_penalty`. - `tools`, `tool_choice`: OpenAI function calling, on models that support tools. - Multimodal: `image_url` content blocks for vision; `file` blocks (base64 `file_data`) for PDFs, on models that support them. - `modalities`: `["image", "text"]` to ask for image output (see Image generation below). ## Image generation Same endpoint as chat, `POST /v1/chat/completions`, same `nfc_` key: image output is a capability of a chat model here, not a separate API. Pass `"modalities": ["image", "text"]` and the images come back in `choices[].message.images`. Models: `google/gemini-3-pro-image`, `google/gemini-3.1-flash-image`, `google/gemini-3.1-flash-lite-image`, `google/gemini-2.5-flash-image`. Check `architecture.output_modalities` in `GET /v1/models` for the current list; any entry whose output modalities include `image` works. ```bash curl https://router.nexforce.ai/v1/chat/completions \ -H "Authorization: Bearer $NEXFORCE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "google/gemini-3-pro-image", "modalities": ["image", "text"], "messages": [{"role": "user", "content": "a cat wearing a hard hat, studio lighting"}] }' ``` Python (OpenAI SDK), saving the first image: ```python import base64 from openai import OpenAI client = OpenAI(base_url="https://router.nexforce.ai/v1", api_key="nfc_...") resp = client.chat.completions.create( model="google/gemini-3-pro-image", modalities=["image", "text"], messages=[{"role": "user", "content": "a cat wearing a hard hat"}], ) url = resp.choices[0].message.images[0]["image_url"]["url"] header, data = url.split(",", 1) # header: "data:image/png;base64" ext = header.split(";")[0].split("/")[-1] # png, jpeg, webp: read it, do not assume open(f"cat.{ext}", "wb").write(base64.b64decode(data)) ``` - Each image is a data URI (`data:;base64,...`) inside `{"type": "image_url", "image_url": {"url": "..."}}`. The Router does not host files, so there is never an http URL to fetch and nothing expires. - **The format varies by model**, so read the mime from the data URI instead of assuming PNG. Measured: `gemini-2.5-flash-image` returns PNG at 1024x1024, `gemini-3-pro-image` returns JPEG at 1408x768. Neither the format nor the aspect ratio is a parameter you set. - `message.content` carries the text the model wrote alongside the image. It is an empty string when the model returned only an image, never `null`. - Image **editing** uses the same call: send the source image as an `image_url` content block (data URI) together with the instruction. - `stream: true` works and delivers one chunk carrying `delta.images`, then `data: [DONE]`. The provider produces the image in one piece, so there is nothing to stream in parts. - **The model is never substituted**, same rule as embeddings: an image is not comparable to text, so Smart Routing does not apply to a request that generates one. - Billing is per token, as with text. The generated image has a token cost the provider reports (about 1290 tokens for one image on `gemini-2.5-flash-image`) and it bills at the model's output rate. There is no per-image tariff. - Image-**only** models with no text output (`gpt-image-1`, `gpt-image-2`) are not served: they have no chat turn and require the provider's own Images API. They return `400` with code `unsupported_capability`. ## Embeddings `POST /v1/embeddings`, OpenAI schema, same `nfc_` key. Models: `text-embedding-3-small`, `text-embedding-3-large`, `text-embedding-ada-002` (OpenAI) and `gemini-embedding-001`, `gemini-embedding-2` (Google). ```bash curl https://router.nexforce.ai/v1/embeddings \ -H "Authorization: Bearer $NEXFORCE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "text-embedding-3-small", "input": ["first document", "second document"] }' ``` Python (OpenAI SDK), one call per batch: ```python from openai import OpenAI client = OpenAI(base_url="https://router.nexforce.ai/v1", api_key="nfc_...") resp = client.embeddings.create(model="text-embedding-3-small", input=docs) vectors = [d.embedding for d in resp.data] ``` - `input` (required): a string, a list of strings (up to 2048 per request), or token ids. - `dimensions`, `encoding_format` (`float` or `base64`), `user`: passed through to the provider unchanged. - No streaming: embeddings return a single response. - **The model is never substituted.** Smart Routing does not apply to embeddings: vectors from two different models are not comparable, so swapping the model would corrupt an index you already wrote. What you ask for is what runs, and `X-Nexforce-Served-Model` always echoes it back. - Billing counts input tokens only; there is no output or cache tier. - Calling an embedding model on `/v1/chat/completions` (or a chat model on `/v1/embeddings`) returns `400` with code `wrong_endpoint` and the endpoint you should use. ## nexforce/smart-route Pass `"model": "nexforce/smart-route"` instead of naming a model and Nexforce picks one for you. The router resolves it to a reference (anchor) model and, when a cheaper equivalent with the same capabilities exists, serves the equivalent; otherwise it serves the anchor. The model that actually answered is always disclosed in the response headers. ## Response headers - `X-Nexforce-Requested-Model`: the model you asked for. - `X-Nexforce-Served-Model`: the model that actually answered. - `X-Nexforce-Substituted`: `true` when Smart Routing swapped the model. - `X-Nexforce-Smart-Route`: `true` when the request used `nexforce/smart-route`. - `X-Config-Synced-At`: timestamp of the config snapshot used. ## Smart Routing (per-request control) When Smart Routing is enabled on your config, the router can swap the requested model for a cheaper equivalent before dispatch, preserving needed capabilities. Override per request: - `X-Nexforce-Smart-Routing: off` — disable for this request (also accepts `false`, `0`). - `X-Nexforce-Pin-Model: true` — force exactly the requested model, no substitution. ## Agent attribution Tag a request with the agent that originated it, to split cost and volume per agent in the Console: - Header `X-Nexforce-Agent: support-bot`, or - Header `X-Nexforce-Metadata: {"agent":"support-bot"}` (also accepts `agent_label`, `_agent`). ## Error codes - `401` missing/invalid key - `402` insufficient wallet balance (recharge at payments.nexforce.ai) - `403` model/provider blocked for this key, or account suspended - `404` model not in the catalog - `400` with code `wrong_endpoint`: right key, wrong endpoint for that model (chat vs embeddings) - `400` with code `unsupported_capability`: the model cannot do what the request asks (image-only model on chat, image input to a text-only model, audio or video output) - `429` rate limit or budget exceeded - `5xx` provider failure (triggers per-model fallback when configured) Error bodies follow the OpenAI error format, so OpenAI SDK error handling works unchanged. ## Editor / agent integrations - Cursor: Settings -> Models -> OpenAI API Key, override base URL with `https://router.nexforce.ai/v1`, add model ids manually. - Cline / Roo Code: provider "OpenAI Compatible", base URL `https://router.nexforce.ai/v1`; the model picker populates from `/v1/models`. - Continue: `provider: openai`, `apiBase: https://router.nexforce.ai/v1`. - OpenCode: `@ai-sdk/openai-compatible` with `baseURL: https://router.nexforce.ai/v1`.