Exemplary AIEnterprise
For creatorsDemoBook a demo
Model Hub

Every model. One control plane.

Language, image, video, speech, music, transcription and embeddings, from frontier APIs and from open-weight models running on your own GPUs. Your agents ask for a model; the hub decides which preset and which deployment answers.

Talk to us
On-premise inferenceAny providerSwap without redeploying
In the catalog today
OpenAI·Google·Anthropic·Runware·Deepgram·AssemblyAIOllama and vLLM on your own GPUs
LanguageImageVideoSpeechMusicTranscriptionEmbedding
LanguageImageVideoSpeechMusicTranscriptionEmbedding
Why this layer exists

Models move. Your agents should not.

Nothing about a model is permanent. A provider posts a sunset date, renewal reprices you, a region runs out of capacity, procurement asks for a second source. Each is a configuration change here and a migration project everywhere else.

Retired

A sunset notice lands. You repoint one logical model at a live deployment, and the agents that name it never learn there was a notice.

Repriced

Your rate moves, or something cheaper ships. Point the Balanced variant at the new deployment, watch the metering for a week, and the old one is still a click away.

Throttled

A provider caps you at the worst hour of the week. The same logical model resolves to a second deployment and the queue drains.

How it resolves

The model your agents name is not where it runs

A logical model is what your agents ask for. A variant is how it should behave. A deployment is where it actually runs. Change the bottom layer and the top layer never notices.

Logical modelsWhat agents ask for
GPT-5gpt-5
Gemma 4gemma-4
Claude Sonnet 4.6claude-sonnet-4-6
VariantsHow it should behave
Fast
BalancedRecommended
High quality
DeploymentsWhere it actually runs
Google Vertex MaaSvertex-maas
vLLM · your GPUsvllm
Logical models
GPT-5gpt-5
Gemma 4gemma-4
Claude Sonnet 4.6claude-sonnet-4-6
Variants
Fast
BalancedRecommended
High quality
Deployments
Google Vertex MaaSvertex-maas
vLLM · your GPUsvllm

Your agents still ask for gemma-4. Only the deployment changed.

Animated diagram: a request for the logical model Gemma 4 resolves through its balanced variant to a deployment, which can be switched between a hosted cloud provider and self-hosted GPUs without changing the model the agent asked for.
The catalog

One catalog, every modality

Language, image, video, speech, music, transcription and embeddings sit side by side in a single workspace catalog, each one carrying its provider, its identifier, its status and its rate.

31 models

GPT-5
gpt-5
openai
GPT-5.2
gpt-5_2
openai
Gemini 2.5
gemini-2_5
google
Gemini 3
gemini-3
google
Gemini 3 Flash
gemini-3-flash
google
Gemini 2.5 Flash Lite
gemini-2_5-flash-lite
google
Gemini 3.1 Flash Lite
gemini-3_1-flash-lite
google
Gemini 3.1 Pro
gemini-3_1-pro
google
Gemini 3.5 Flash
gemini-3_5-flash
google
Gemma 4
gemma-4
google
Claude Haiku 4.5
claude-haiku-4-5
anthropic
Claude Sonnet 4.6
claude-sonnet-4-6
anthropic
Claude Opus 4.7
claude-opus-4-7
anthropic
Qwen 3 14B
ollama-qwen3-14b
ollamaSelf-hosted
Gemma 3 4B Vision
ollama-gemma3-4b-vision
ollamaSelf-hosted
Gemma 4 (open weights)
gemma-4-vllm
vllmSelf-hosted
GPT Image 1.5
gpt-image-1_5
openai
Nano Banana
nano-banana
google
Nano Banana Pro
nano-banana-pro
google
Kling Video 2.6 Pro
kling-v2_6
runware
Kling Video 3.0
kling-v3
runware
Kling Video 3.0 Pro
kling-v3-pro
runware
Kling Video O3 Pro
kling-o3-pro
runware
Kling Avatar 2.0 Standard
kling-avatar-2_0-standard
runware
Kling Avatar 2.0 Pro
kling-avatar-2_0-pro
runware
ElevenLabs v3
eleven-v3
runware
ElevenLabs Turbo v2
eleven-turbo-v2
runware
ElevenLabs Music v1
eleven-music-v1
runware
Deepgram Nova-3
deepgram-nova-3
deepgram
AssemblyAI Best
assemblyai-best
assemblyai
Gemini Embedding 001
gemini-embedding-001
google

Relative cost bands, not quoted rates. Every model surfaces its own per-1K, per-second or per-minute pricing in the console.

See what each model actually costs
Variants

Tune the trade-off, not the code

A variant is a saved preset over a base model. Point an agent at the preset instead of the raw model, and you can retune latency, cost and quality centrally, without touching a single agent.

Balanced quality and speedRecommended

The default for most agents. Enough headroom for tool calling and multi-step work without paying frontier rates on every turn.

SpeedEconomyQualityIllustrative
Modalities

Add a modality, not a vendor

Speech, video and embeddings are governed, metered and deployed exactly like a language model. Giving an agent a voice is a line of configuration, not a new contract, a new key and a new line on the bill.

Reasoning and generation

Frontier and open-weight language models for planning, tool calling, drafting and extraction, with vision where the model supports it.

Image generation

Generate and edit imagery inside a workflow, with the same governance and metering as every other model in the catalog.

Video generation

Text- and image-to-video, including avatar presenters. Billed by the second rather than by the token, and surfaced that way.

Text to speech

Give an agent a voice. Low-latency turbo voices for interactive use, higher-fidelity voices for produced output.

Music generation

Score and sound-bed generation for produced media, available to agents as an ordinary tool call.

Speech to text

Transcribe audio and video across 120+ languages, with speaker diarization. A first-class modality, not an add-on.

Embeddings

Vector representations that power grounded retrieval over your indexed files and knowledge graph.

What it buys you

The next model is a decision, not a migration

A lab ships something better on a Tuesday. You add the deployment, point the Balanced variant at it, and every agent in the workspace is on it that afternoon. The old deployment stays active for as long as it takes to trust the new numbers. Model choice stops being a quarterly project.

Run the old model and the new one against the same work. Settle the argument on your numbers, not somebody else's benchmark.

Give one agent a voice before the business case exists. A week of metering says whether the deployment stays on.

Work you had ruled out because the data cannot leave the building becomes ordinary work, on the same catalog as everything else.

Sovereignty

Frontier in the cloud, or nothing leaves the building

The same catalog, the same agents, the same presets. What changes is where inference happens. Run open-weight models on your own GPUs through Ollama or vLLM and no prompt, document or embedding ever crosses your network boundary.

Hosted frontier models
  • Managed capacity from OpenAI, Google Vertex and Anthropic
  • Specialist providers for image, video, speech and transcription
  • Per-token, per-second and per-minute rates surfaced per deployment
openaivertex-maasanthropicrunware
Cloud
Your GPUs, air-gapped
  • Open-weight language and embedding models served by Ollama or vLLM
  • Prompts, documents and embeddings stay inside
  • Offline, signed model updates for air-gapped sites
ollamavllm
On-premise
Your GPUs, air-gapped
  • Open-weight language and embedding models served by Ollama or vLLM
  • Prompts, documents and embeddings stay inside
  • Offline, signed model updates for air-gapped sites
ollamavllm
On-premise
Hosted frontier models
  • Managed capacity from OpenAI, Google Vertex and Anthropic
  • Specialist providers for image, video, speech and transcription
  • Per-token, per-second and per-minute rates surfaced per deployment
openaivertex-maasanthropicrunware
Cloud

Most deployments run both. Language and embeddings serve from your own GPUs, and the modalities that still need a provider are in the catalog because an administrator put them there. Anything that must not leave the building is not in it.

See how it lands in your data centre
Governance

Every model accounted for

The catalog is workspace-scoped and administered, not a free-for-all. Administrators decide which models exist, which are recommended, and which credentials they run under.

At a glance

Catalog scope
Models, variants and deployments are configured per workspace
Status control
Any model or deployment can be made active or inactive centrally
Recommended presets
Mark the variant teams should reach for by default
Pricing visibility
Per-1K token, per-second and per-minute rates surfaced per deployment
Metering
Usage attributed by model, agent and group across the billing period
Credentials
Organization keys override platform defaults, and stay masked
Local runtimes
Ollama and vLLM deployments served from your own hardware
Failover
Repoint a logical model at another deployment without redeploying agents

A catalog only administrators can change, running under credentials you control.

SOC 2 Type IIAICPA
HIPAACompliant
GDPRReady
ISO 27001Certified

See Exemplary AI in action

Book a demo and we'll show you purpose-built agents grounded in your own knowledge — deployed on your infrastructure.

Book a demo
Exemplary AIEnterprise

Agentic AI for the enterprise — installed inside your network.

  • Sovereign deployment
  • Purpose-agnostic agents
Book a demo

Build

  • Agents
  • Model Hub
  • Knowledge Graph
  • File store

Connect

  • Integrations
  • Chatbots
  • Browser Extension

Govern

  • Management
  • Groups
  • Analytics

Company

  • About Us
  • Exemplary for creators
  • Book a demo

Use cases

By industry
  • Healthcare
  • Government
  • Financial services
  • Legal
  • Media
By department
  • Customer Support
  • IT & Engineering
  • Human Resources
  • Finance & Procurement
  • Compliance & Risk

© 2026 Exemplary AI. All rights reserved.

  • SOC 2 Type II
  • HIPAA
  • GDPR-ready
  • ISO 27001