Frontier and open-weight language models for planning, tool calling, drafting and extraction, with vision where the model supports it.
Every model. One control plane.
Language, image, video, speech, music, transcription and embeddings, from frontier APIs and from open-weight models running on your own GPUs. Your agents ask for a model; the hub decides which preset and which deployment answers.
Models move. Your agents should not.
Nothing about a model is permanent. A provider posts a sunset date, renewal reprices you, a region runs out of capacity, procurement asks for a second source. Each is a configuration change here and a migration project everywhere else.
Retired
A sunset notice lands. You repoint one logical model at a live deployment, and the agents that name it never learn there was a notice.
Repriced
Your rate moves, or something cheaper ships. Point the Balanced variant at the new deployment, watch the metering for a week, and the old one is still a click away.
Throttled
A provider caps you at the worst hour of the week. The same logical model resolves to a second deployment and the queue drains.
The model your agents name is not where it runs
A logical model is what your agents ask for. A variant is how it should behave. A deployment is where it actually runs. Change the bottom layer and the top layer never notices.
Your agents still ask for gemma-4. Only the deployment changed.
One catalog, every modality
Language, image, video, speech, music, transcription and embeddings sit side by side in a single workspace catalog, each one carrying its provider, its identifier, its status and its rate.
31 models
Relative cost bands, not quoted rates. Every model surfaces its own per-1K, per-second or per-minute pricing in the console.
Tune the trade-off, not the code
A variant is a saved preset over a base model. Point an agent at the preset instead of the raw model, and you can retune latency, cost and quality centrally, without touching a single agent.
The default for most agents. Enough headroom for tool calling and multi-step work without paying frontier rates on every turn.
Add a modality, not a vendor
Speech, video and embeddings are governed, metered and deployed exactly like a language model. Giving an agent a voice is a line of configuration, not a new contract, a new key and a new line on the bill.
Generate and edit imagery inside a workflow, with the same governance and metering as every other model in the catalog.
Text- and image-to-video, including avatar presenters. Billed by the second rather than by the token, and surfaced that way.
Give an agent a voice. Low-latency turbo voices for interactive use, higher-fidelity voices for produced output.
Score and sound-bed generation for produced media, available to agents as an ordinary tool call.
Transcribe audio and video across 120+ languages, with speaker diarization. A first-class modality, not an add-on.
Vector representations that power grounded retrieval over your indexed files and knowledge graph.
The next model is a decision, not a migration
A lab ships something better on a Tuesday. You add the deployment, point the Balanced variant at it, and every agent in the workspace is on it that afternoon. The old deployment stays active for as long as it takes to trust the new numbers. Model choice stops being a quarterly project.
Run the old model and the new one against the same work. Settle the argument on your numbers, not somebody else's benchmark.
Give one agent a voice before the business case exists. A week of metering says whether the deployment stays on.
Work you had ruled out because the data cannot leave the building becomes ordinary work, on the same catalog as everything else.
Frontier in the cloud, or nothing leaves the building
The same catalog, the same agents, the same presets. What changes is where inference happens. Run open-weight models on your own GPUs through Ollama or vLLM and no prompt, document or embedding ever crosses your network boundary.
- Managed capacity from OpenAI, Google Vertex and Anthropic
- Specialist providers for image, video, speech and transcription
- Per-token, per-second and per-minute rates surfaced per deployment
- Open-weight language and embedding models served by Ollama or vLLM
- Prompts, documents and embeddings stay inside
- Offline, signed model updates for air-gapped sites
Most deployments run both. Language and embeddings serve from your own GPUs, and the modalities that still need a provider are in the catalog because an administrator put them there. Anything that must not leave the building is not in it.
See how it lands in your data centreEvery model accounted for
The catalog is workspace-scoped and administered, not a free-for-all. Administrators decide which models exist, which are recommended, and which credentials they run under.
At a glance
- Catalog scope
- Models, variants and deployments are configured per workspace
- Status control
- Any model or deployment can be made active or inactive centrally
- Recommended presets
- Mark the variant teams should reach for by default
- Pricing visibility
- Per-1K token, per-second and per-minute rates surfaced per deployment
- Metering
- Usage attributed by model, agent and group across the billing period
- Credentials
- Organization keys override platform defaults, and stay masked
- Local runtimes
- Ollama and vLLM deployments served from your own hardware
- Failover
- Repoint a logical model at another deployment without redeploying agents
A catalog only administrators can change, running under credentials you control.
See Exemplary AI in action
Book a demo and we'll show you purpose-built agents grounded in your own knowledge — deployed on your infrastructure.