ENTERPRISE AI GATEWAY · GOVERNANCE CONTROL PLANE · SELF-HOSTED
Your Enterprise Needs AI. Your AI Needs ShieldON.
As models multiply, agents begin to act, and AI connects to more business systems, governance is no longer optional. ShieldON gives you ultimate transparency to keep risk, cost and access under control as AI scales.
Demo console screen: the Overview page with request metrics and a live requests table. Static sample data for presentation only.
Architecture
Every app and agent on one side. Every model on the other. Shieldon in between.
Apps, agents and tools inside the enterprise reach cloud and local models through Shieldon. Sensitive data is redacted before it leaves, each request goes to the model that fits it, every trip lands on the ledger, and your own AI agent can change the policy.
Inside the enterprise boundary, apps, agents, tools and assistants call cloud and local models through the Shieldon AI Gateway, and responses come back the same way. Guardrails redact sensitive data before a request leaves; routing sends simple requests to local models and complex ones to frontier models, with fallbacks; every request is recorded with its key, route and tokens; usage produces token-saving suggestions; and the enterprise's own AI agent reads usage, sets policy and publishes a release through the API.
LLM INTEGRATIONS
Connect a provider once. Switch, mix and fail over without touching application code.
Put OpenAI, Anthropic, Gemini, Bedrock, Alibaba Cloud Bailian, DeepSeek and thirty more behind the same OpenAI-compatible endpoint. Provider keys are stored encrypted inside Shieldon and shared only with the Workspaces you choose — applications only ever see a Shieldon key. Chat, Responses, Embeddings, Images and Audio are supported per model.
- OpenAI
- Azure OpenAI
- Microsoft Foundry
- GitHub Models
- Anthropic
- Google Gemini / Vertex AI
- AWS Bedrock
- Alibaba Cloud Bailian
- Zhipu AI
- DeepSeek
- Moonshot AI
- MiniMax
- Groq
- Together AI
- Fireworks AI
- OpenRouter
- xAI
- Perplexity
- Cerebras
- DeepInfra
- Nebius
- SambaNova
- SiliconFlow
- Novita AI
- Jina AI
- Voyage AI
- Cloudflare Workers AI
- Hugging Face
- Ollama
- vLLM
- LocalAI
- Databricks
- Generic OpenAI-compatible
- OpenAI
- Azure OpenAI
- Microsoft Foundry
- GitHub Models
- Anthropic
- Google Gemini / Vertex AI
- AWS Bedrock
- Alibaba Cloud Bailian
- Zhipu AI
- DeepSeek
- Moonshot AI
- MiniMax
- Groq
- Together AI
- Fireworks AI
- OpenRouter
- xAI
- Perplexity
- Cerebras
- DeepInfra
- Nebius
- SambaNova
- SiliconFlow
- Novita AI
- Jina AI
- Voyage AI
- Cloudflare Workers AI
- Hugging Face
- Ollama
- vLLM
- LocalAI
- Databricks
- Generic OpenAI-compatible
01 · WHO IS CALLING THE API?
Know which app, team and environment is behind every call — before it reaches a model.
Each application gets its own Gateway API Key, and that key carries policy: which models it may use, which context every request must include, and which networks it may call from. Rotate it without breaking anything. The upstream provider key never leaves Shieldon.
Illustrative gateway responses
Demo console screen: a Gateway API key named billing-assistant is created, given four Metadata variables, and restricted to 203.0.113.0/24; a request from another address is rejected.
Create Gateway API key
Give the key a name and decide what it is allowed to call. The application gets one secret; you keep the provider credentials.
Make every request identify itself
Decide what every call must declare — application, environment, team, project, customer, user, feature or session — and Shieldon rejects calls that don't. Those same fields become the basis for cost attribution later.
Only from the networks you expect
Restrict a key to your own network ranges, and calls from anywhere else are turned away before they touch a model.
Organizations and Workspaces keep departments, products and environments apart, each with its own members and roles.
02 · WHAT INFORMATION IS BEING SENT TO THE LLM?
Card numbers, ID numbers and API keys stop at the boundary. Prompts are never stored by default.
One safety boundary for every application, instead of a separate check in every codebase. Shieldon inspects what goes in and what comes out, catches personal data and credentials, and either flags the call or blocks it — your choice, per policy. Only the decision is recorded, never the matched content. And by default, prompts, outputs and embeddings are not stored at all.
- Personal data: email addresses, phone numbers, national ID numbers, payment card numbers, US Social Security numbers
- Secrets: private keys, session tokens, cloud and SaaS credentials, passwords pasted into a prompt
- Your own rules: keywords, regular expressions, tool-call allowlists, or your own policy service via a secured webhook
- Start in Monitor mode to see what would be caught, then switch to Block when you are confident.
Demo console screen: a request containing a payment card number is blocked by the Sensitive data protection profile; a second request passes. The activity record stores only a reason code.
Blocked before it ever reaches the provider. Shieldon rejects the request outright rather than quietly rewriting it, and records only the reason, never the value.
PRIVACY DEFAULT
- prompt text
- model output
- tool arguments
- webhook bodies
- embedding vectors
Need to keep content for a short time to debug an issue? That is an explicit, separately agreed policy: encrypted at rest, and every look at it is audited.
Demo console screen: Request details req_48e5aa — Logical model shieldon-chat, Provider ID openai-prod, Status Succeeded, Latency 812 ms. Captured payloads: Content capture was disabled or no payload was retained for this request.
03 · HOW SHOULD THE BEHAVIOR BE CONTROLLED?
Switch providers, cap spend and roll back in minutes. No code change, no hotfix.
Applications call a stable model name; you decide what sits behind it. Route to one provider, fall back to another when it fails, spread traffic by weight, or pick a target by rule. Rate limits and budgets are enforced before a provider is ever called. Every change goes live as a reviewed release you can roll back in one click.
Demo console screens: a fallback route absorbs a 429; a limit policy blocks at 600 RPM; a Credits budget blocks at its limit and is adjusted; a release is published and rolled back.
Illustrative route attempts · req_9b21c4
- attempt 1 · @openai-prod/gpt-4.1
Fallback that callers never see
When your primary provider rate-limits or times out, Shieldon retries and moves on to the next target you approved. The application sees one successful response and one request ID — nothing to catch, nothing to redeploy.
Test failover with your own providers and models before relying on it.
Limits per model, with a plan for when things go wrong
Cap requests, tokens and concurrent calls per model. And if the limiter itself is ever unavailable, you have already chosen whether safety or availability wins.
Credits budgets that actually stop
Set a budget per Workspace or per application key, by day, week or month in your own timezone. When it runs out, calls stop or a warning fires — your call. Every top-up is logged, so nobody quietly raises the ceiling.
Every change reviewed. Every release reversible.
Validate the draft, see exactly what will change compared with what is live, approve, publish. If something looks wrong afterwards, roll back to the last known-good release in one click. The gateway only ever runs a published release.
04 · HOW CAN THE SYSTEM IMPROVE OVER TIME?
Every call has an owner. Every cost has a home. Every change has a name on it.
Shieldon keeps two records for you: a usage and cost ledger for every request, and an audit log for every change to the control plane. Together they answer the finance question today — which app, team or customer drove this month's spend — and they are the foundation for a human-approved improvement loop that is on our roadmap.
Chargeback by app, team or customer, without the spreadsheet.
Every request carries the context you require — application, environment, team, project, customer, user, feature, session. Shieldon records it alongside tokens, the model actually used, retries and errors, so you can break cost down any way finance asks. Image and audio pricing is not yet fully normalized across providers.
- application
- environment
- team
- project
- customer
- user
- feature
- session_id
curl "https://gateway.example.com/v1/chat/completions" \ -H "Authorization: Bearer $SHIELDON_GATEWAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "shieldon-chat", "messages": [{"role":"user","content":"Hello"}], "metadata": { "application": "support-bot", "environment": "production" } }'That row is everything Shieldon kept: who, how much, which route, what it cost. Not the prompt.
A change log nobody can edit.
Who created, changed, published, rolled back or rotated what — when, and from where. Every entry is permanent, so an incident review starts with facts instead of guesses.
- Observe usage
- identify patterns
- propose improvements
- validate and evaluate
- approveHuman approval required
- release
- measure outcomes
Demo console screens: a request's metadata is recorded in the usage ledger; the audit log lists published, rotated, rolled back and updated events.
REQUEST PIPELINE
Every call takes the same path. Every step is a place you can set policy.
Authenticate, check the sender, load the published policy, inspect the input, choose a route, check limits and budget, call the provider, fall back if needed, inspect the output, meter the cost, record the audit trail. Nothing skips a stage, and the chapters above map onto it in order.
- Ingress
- Request IDChapter 04, How does it improve?
- Gateway API key authChapter 01, Who is calling?
- Org/Workspace authorizationChapter 01, Who is calling?
- Published policy snapshotChapter 03, How is it controlled?
- API adapter
- Guardrails (input)Chapter 02, What is being sent?
- Router→ @openai-prod/gpt-4.1Chapter 03, How is it controlled?
- Rate/Token/Concurrency/Budget pre-checkChapter 03, How is it controlled?
- Provider adapter
- Retry/Fallback429 → @bailian/qwen-maxChapter 03, How is it controlled?
- Stream normalization
- Guardrails (output)Chapter 02, What is being sent?
- Usage & cost meteringChapter 04, How does it improve?
- Async usage logs / auditChapter 04, How does it improve?
- Response
WHO THIS IS FOR
Built for the teams who have to answer for AI in production.
PILOT IN YOUR ENVIRONMENT
From a provider key to your first governed call in five steps, on one Linux host.
Quick Setup walks you through it on a single page: connect a provider, create an application key, publish, make a real test call, integrate. Then hand your teams one endpoint.
curl "https://gateway.example.com/v1/chat/completions" \ -H "Authorization: Bearer $SHIELDON_GATEWAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"shieldon-chat","messages":[{"role":"user","content":"Hello"}],"metadata":{"application":"support-bot","environment":"production"}}'Your applications only ever use this key. The provider's own key stays inside Shieldon.
Single Linux host · Docker Compose. Nginx connects to Gateway (Go), Control Plane (Go), Web console + Docs. Gateway (Go) connects to PostgreSQL, Redis. Control Plane (Go) connects to PostgreSQL, Redis, S3-compatible storage (optional).
Everything runs in your environment: one entry point, a handful of containers, standard PostgreSQL and Redis, optional object storage — shipped as multi-arch images with an offline install bundle.
BEYOND THE MODEL
Govern the capabilities agents use, not just the models they call.
Codex, Claude Code, OpenClaw and the agents your own teams build will all need to pass through one gateway eventually. Our direction is to govern everything they use — Skills, Tools, MCP servers, CLI workflows and APIs — with the same four questions: is it safe, what does it cost, can we see it, and does it get better over time.
1. Governance foundation
Know what you have: one identity model and an inventory of every agent, Skill, Tool, MCP server and API, with owners, data classification and audit.
2. Policy and control
Set the rules: safety policies and spending controls that apply across every agent, with approvals, capability restrictions and release management.
3. Continuous improvement
Get better on purpose: recommendations for prompts, policies, Skills and use cases drawn from real usage, evaluated and rolled out under control.
4. Enterprise operating layer
Run AI like the rest of the enterprise: discover and register shadow-IT agents, integrate more runtimes, expand deployment options, extend where customers need it most.
Safety
Know what crosses the model and tool boundary, keep secrets and personal data inside, and put high-risk actions behind approval.
Finance / FinOps
Attribute every dollar to a person, team, app, agent, Skill, Tool, model and Provider; enforce budgets; spot anomalies early.
Visibility / Observability
One connected trail from the person to the outcome: which agent, which Skill, which Tools and MCP servers, which APIs and models, which policies applied.
RSI
Real usage suggests better prompts, policies and use cases; every suggestion is versioned, evaluated and approved by a person.
Self-hosted · Docker Compose · Docs included in the deployment · Sign in at app.shieldon.ai