Votal Shield — Rafay integration guide
This is the partner handoff for enabling Votal Shield AI guardrails on the Rafay platform, multi-tenant. Pair it with the machine-readable contract:
- OpenAPI 3.1 spec:
docs/assets/openapi-partner.json(curated partner subset — the guard, policy, and telemetry endpoints only, never the admin surface). Import this into Rafay’s codegen / API explorer.
1. Model in one paragraph
Every Rafay customer = one Shield tenant, identified by its own API key. Rafay provisions a tenant + key per customer (section 4), stores the key, and at request time calls Shield PreCall on the user’s input and PostCall on the model’s output, passing that customer’s key. Shield resolves the tenant from the key, applies that tenant’s configured guardrails, and returns an allow/block/redact verdict. Policy is per tenant, so one customer’s rules never touch another’s. Shield never needs to store Rafay’s model keys or see anything beyond the text being screened.
Rafay customer app ──► Rafay AI serving path
│ (1) PreCall: POST /guardrails/input ─┐
│ block? → return the block to user │ X-API-Key:
▼ │ <customer tenant key>
LLM / model ├─► Votal Shield
│ │ (hosted API, or
│ (2) PostCall: POST /guardrails/output ─┘ in-cluster container)
▼ block/redact? → use sanitized_output
response to user
2. Base URL (two deployment shapes, same API)
- Hosted (fastest):
https://api.guardrails.votal.ai - In-cluster (data residency): the Shield guardrail container deployed in the customer’s cluster; same paths, Rafay sets the base URL per deployment. See Appendix A for the Kubernetes/Helm reference.
Everything below is relative to the chosen base URL.
3. Authentication
Send the customer’s tenant key on every guard call. Header priority:
X-API-Key: <tenant-key>← preferred (avoids collision with upstream proxies)X-Tenant-Key: <tenant-key>Authorization: Bearer <tenant-key>
The key is hashed and mapped to exactly one tenant; the tenant’s policy is applied
automatically. A caller cannot override the tenant — any tenant_id in the
body/headers that disagrees with the key is rejected (IDOR defense).
REQUIRED deployment setting for multi-tenant isolation: Deploy the Rafay-facing Shield with
SHIELD_GUARD_REQUIRE_KEY=enforce. The default isoff, under which a call with a missing/invalid key fails open (screens with no tenant policy) instead of being refused. In a multi-tenant deployment that is a silent isolation gap — set it toenforceso a bad key returns401 missing_tenant_key/invalid_tenant_key.
4. Onboarding a customer (“enable guardrails”)
Enabling guardrails for a Rafay customer = provision a Shield tenant + key and store the key against that customer. Two ways:
4a. Rafay provisions via the admin API (recommended)
These admin endpoints are privileged and intentionally NOT in the partner OpenAPI (
openapi-partner.jsonpublishes only tenant-scoped surfaces). Votal issues Rafay the admin key and these provisioning endpoints separately as a partner capability. Do not expect them in the imported spec.
Rafay holds a Shield admin key (X-Admin-Key, issued to Rafay once, out of
band — not a tenant key). Per customer:
# create the tenant
curl -X POST "$BASE/v1/admin/tenants" \
-H "X-Admin-Key: $RAFAY_ADMIN_KEY" -H "content-type: application/json" \
-d '{"tenant_id":"rafay-cust-acme","name":"ACME Corp","api_keys":[]}'
# mint that tenant's key (store the returned key against the customer in Rafay)
curl -X POST "$BASE/v1/admin/tenants/rafay-cust-acme/api-keys" \
-H "X-Admin-Key: $RAFAY_ADMIN_KEY" -H "content-type: application/json" \
-d '{"label":"rafay-prod","scope":"guard"}'
List/revoke a customer’s keys: GET / DELETE /v1/admin/tenants/{tenant_id}/api-keys.
4b. Self-service (customer mints their own)
If the customer has a Shield portal session or an existing key, they can mint:
POST /v1/tenant/me/api-keys (returns the plaintext key once — store it then).
The admin key is the one high-privilege secret in this integration. It mints and revokes tenants, so Rafay must hold it like any provisioning credential (vault, not in customer-visible config). Tenant keys are per customer and low-blast-radius by comparison.
5. Let the customer choose which guardrails are on
Two options — pick one, or offer both:
- Embed the Shield tenant portal (SSO) so the customer configures policy in Shield’s own UI. Lowest build cost for Rafay.
- Surface it in Rafay’s UI via the policy API (authed with the customer’s key):
GET /v1/tenant/me/policies→ current input/output guardrails + custom policiesPUT /v1/tenant/me/policies→ replace the input/output guardrail config- Natural-language / OWASP-style custom rules:
GET|POST /v1/tenant/me/custom-policies/,POST /v1/tenant/me/custom-policies/{id}/enable(and/disable). A custom policy body is:name,description,prompt(20–2000 chars),action(pass|warn|redact|block),stage(input|output),confidence_threshold(0.5–1.0),priority(1–1000),multi_turn.
New/changed policy takes effect on the next guard call for that tenant — no redeploy.
6. Runtime: the guard calls
6a. PreCall — screen the user’s input
POST /guardrails/input
Request (only message is required; the tenant’s configured guardrails run
automatically — you do not list them per call):
{
"message": "the user's prompt text",
"messages": [{"role":"user","content":"..."}],
"session_id": "rafay-session-123",
"user_role": "analyst",
"agent_key": "rafay-app-x"
}
curl -X POST "$BASE/guardrails/input" \
-H "X-API-Key: $CUSTOMER_KEY" -H "content-type: application/json" \
-d '{"message":"my SSN is ..."}'
Response:
{
"safe": false,
"action": "block",
"guardrail_results": [
{"guardrail":"pii-detection","passed":false,"action":"block",
"message":"message contains a Social Security Number","details":{...},"latency_ms":41.2}
],
"inference_time_ms": 44.0
}
- PASS:
safe=true,action="pass", everyguardrail_results[].passed=true→ forward the prompt to the model. - BLOCK:
safe=false,action="block"→ do not call the model; return the block to the user. The reason is the failing result’smessage. actionis the highest-severity outcome across guardrails, rankedpass < log < warn < redact < block. Onredact, a guardrail may have rewritten content — see PostCall for the sanitized payload pattern.
6b. PostCall — screen the model’s output (and tool results)
POST /guardrails/output
Request (output required):
{
"output": "the model's response text",
"context": {"tool_name":"search","user_role":"analyst","stage":"output"}
}
Response adds two conditional fields to the same schema:
sanitization— audit of any data-policy redaction appliedsanitized_output— present only when the payload was modified; forward this to the user instead of the original.
So PostCall handling: if action=="block" → suppress the output; else if
sanitized_output is present → return it; else return the original.
6c. File uploads
POST /guardrails/file — same verdict schema plus a file block; use it when a
customer uploads a document to an AI tool through Rafay.
Also in the partner spec (beyond prompt screening)
The curated spec also publishes the agentic / tool surfaces Rafay may want for
agent workloads: agent registry (/v1/agents/registry), tool-call RBAC
(/v1/shield/tool/check, /v1/shield/tool/output, /v1/agents/tools/policies),
data policies (/v1/data-policies/...), and MCP gateway upstreams
(/v1/tenant/me/mcp-gateway/upstreams). Same tenant-key auth.
Optional: one-call gateway (not in the partner spec)
If Rafay prefers Shield to also make the model call, an OpenAI-compatible endpoint
(POST /v1/shield/chat/completions) does input-screen → model → output-screen in
one request. It is not published in the partner subset — ask Votal to enable
it if Rafay wants Shield to own the model call rather than the split
PreCall/PostCall above.
7. Telemetry (per customer)
GET /v1/tenant/me/telemetry— (in the partner spec) tenant-scoped agent/chat telemetry, keyed by the caller’s own key; filterslimit,offset,agent_key,status(pass|warn|redact|mask|block),tool_name,q,since,until. This is the partner-appropriate telemetry call.GET /v1/tenant/me/usage,GET /v1/tenant/me/audit,GET /v1/tenant/me/guardrails/metrics— usage, audit, and effectiveness (all in the partner spec).- An admin-scoped
GET /v1/shield/decisions/{tenant_id}also exists for cross-tenant enforcement queries, but like the admin provisioning endpoints it is not in the partner subset — use/v1/tenant/me/telemetryfrom Rafay.
Rafay can surface these in its own dashboard, or link customers to Shield’s Telemetry view.
8. Failure modes & latency (state these to customers, don’t SLA them)
- Store/Redis outage: Shield fails open (screens with a warning rather than taking the customer’s app down). If the deployment must fail closed, raise it with Votal — it’s a deploy posture, not a per-call flag.
- Timeouts: Rafay should set a client timeout on the guard call and decide
its own fail-open vs fail-closed if Shield is unreachable. Recommended: fail
closed on
PreCallfor regulated tenants, fail open otherwise — Rafay’s choice per customer. - Latency (typical, measured — not an SLA): Tier-1 keyword/regex
<5 ms, Tier-2 sentiment/topic~150 ms, Tier-3 adversarial/PII~500 ms. The response reports actualinference_time_ms/ per-resultlatency_msevery call. Absolute latency varies with load and model placement; quote ranges, not a single number.
9. End-to-end Rafay steps (checklist)
- Receive from Votal: the admin key (
X-API-Key-styleX-Admin-Key), a sandbox tenant + key, the base URL(s), anddocs/assets/openapi-partner.json. - Import the OpenAPI into Rafay’s API tooling; generate a client.
- Deploy posture: set
SHIELD_GUARD_REQUIRE_KEY=enforceon the Rafay-facing Shield (hosted config confirmed by Votal, or in the in-cluster Helm values). - Onboarding hook: when a Rafay customer enables guardrails, call §4a to create the tenant + mint the key; store the key in Rafay’s secret store keyed to the customer.
- Policy UX: embed the Shield portal (§5) or build the policy screens on
/v1/tenant/me/policies+/custom-policies. - Serving path: wire §6a PreCall before the model and §6b PostCall after;
honor
actionandsanitized_output. - Observability: pull §7 telemetry into Rafay’s dashboard per customer.
- Validate against the sandbox tenant: a benign prompt passes, a policy-
violating prompt returns
action=block, and the decision appears in telemetry. - Go-live: repeat for a pilot customer, confirm isolation (customer A’s key never sees customer B’s policy or decisions), then roll out.
10. What Votal still needs to confirm before sending
- Issue Rafay an admin/provisioning key and a sandbox tenant + key.
- Confirm the deployment shape (hosted vs in-cluster) and, if in-cluster, hand over the Helm chart + model sizing.
- Confirm
SHIELD_GUARD_REQUIRE_KEY=enforceis set on the Rafay-facing deployment. - Decide whether policy config is portal-embed or API in Rafay’s UI (or both) so section 5 can be trimmed to the chosen path.
Appendix A — In-cluster deployment (Kubernetes / Helm)
For data residency, the Shield guardrail server runs inside the customer’s cluster; Rafay points its serving path at the in-cluster Service instead of the hosted URL. The API in sections 3–7 is identical — only the base URL changes.
What Votal ships: the container image (registry path provided per partner) and a packaged Helm chart on request. The manifests below are a reference Rafay/the customer can apply directly or fold into their own chart — they are not a published chart in this repo. (A
deploy/helm/shield-identitychart exists for the identity plane; the guardrail data plane is shipped as an image today.)
A.1 Pick an image (model placement)
| Image | Model | Node | Use when |
|---|---|---|---|
Dockerfile.cloud (app only, port 80) |
calls a remote model endpoint Votal provides | CPU | lightest in-cluster footprint; model traffic may leave the cluster |
Dockerfile (app + in-cluster vLLM, ports 80 app / 8000 model) |
local vLLM, nothing leaves the cluster | GPU | strict residency — the model runs in the customer’s cluster too |
Both expose the same guard API on port 80. The GPU image additionally runs vLLM on 8000 (internal to the pod).
A.2 Shared dependency: Redis (tenant store + policy cache)
Multi-tenancy needs a Redis the Shield pods share — it holds tenant→key mappings and the policy cache. Use a managed Redis or a serverless one:
REDIS_URL=redis://<host>:6379/0, orUPSTASH_REDIS_REST_URL+UPSTASH_REDIS_REST_TOKENfor serverless.
A.3 Config (the env that matters)
| Env | Value | Why |
|---|---|---|
SHIELD_GUARD_REQUIRE_KEY |
enforce |
required — else a bad key fails open (isolation gap) |
REDIS_URL (or UPSTASH_*) |
your Redis | tenant store + policy cache |
SHIELD_ADMIN_KEY |
secret | authorizes tenant provisioning (section 4a) |
LLM_MODEL_NAME |
model id | which guardrail model to use |
| (cloud image) model endpoint + token | from Votal | points the app at the remote model |
(GPU image) VLLM_PORT=8000, SHIELD_MAX_MODEL_LEN |
as sized | local vLLM; confirm GPU/VRAM with Votal |
A.4 Reference manifest (cloud-model image, CPU)
apiVersion: v1
kind: Secret
metadata: { name: shield-guardrail, namespace: shield }
stringData:
REDIS_URL: "redis://redis.shield.svc:6379/0"
SHIELD_ADMIN_KEY: "<admin-key-from-votal>"
SHIELD_LLM_TOKEN: "<model-endpoint-token-from-votal>"
---
apiVersion: apps/v1
kind: Deployment
metadata: { name: shield-guardrail, namespace: shield }
spec:
replicas: 2 # stateless app; scale horizontally
selector: { matchLabels: { app: shield-guardrail } }
template:
metadata: { labels: { app: shield-guardrail } }
spec:
containers:
- name: shield
image: <registry>/votal/shield-guardrail-cloud:<tag> # Votal provides
ports: [ { containerPort: 80 } ]
env:
- { name: SHIELD_GUARD_REQUIRE_KEY, value: "enforce" }
- { name: LLM_MODEL_NAME, value: "<model-id-from-votal>" }
- { name: REDIS_URL, valueFrom: { secretKeyRef: { name: shield-guardrail, key: REDIS_URL } } }
- { name: SHIELD_ADMIN_KEY, valueFrom: { secretKeyRef: { name: shield-guardrail, key: SHIELD_ADMIN_KEY } } }
- { name: SHIELD_LLM_TOKEN, valueFrom: { secretKeyRef: { name: shield-guardrail, key: SHIELD_LLM_TOKEN } } }
readinessProbe: { httpGet: { path: /health, port: 80 } }
resources:
requests: { cpu: "500m", memory: "1Gi" }
limits: { cpu: "2", memory: "2Gi" }
---
apiVersion: v1
kind: Service
metadata: { name: shield-guardrail, namespace: shield }
spec:
selector: { app: shield-guardrail }
ports: [ { port: 80, targetPort: 80 } ]
Rafay’s serving path then calls http://shield-guardrail.shield.svc:80/guardrails/input.
A.5 GPU (fully on-prem model) deltas
Swap the image for the app+vLLM one and give the pod a GPU; the model never leaves the cluster:
containers:
- name: shield
image: <registry>/votal/shield-guardrail:<tag> # app + vLLM
ports: [ { containerPort: 80 } ] # 8000 is pod-internal
env:
- { name: SHIELD_GUARD_REQUIRE_KEY, value: "enforce" }
- { name: LLM_MODEL_NAME, value: "<model-id>" }
- { name: VLLM_PORT, value: "8000" }
- { name: SHIELD_MAX_MODEL_LEN, value: "<confirm with Votal>" }
resources:
limits: { nvidia.com/gpu: 1 } # GPU/VRAM sizing: confirm with Votal
# schedule onto a GPU node pool (nodeSelector/taints per Rafay's cluster)
replicas for the GPU variant is bounded by available GPUs; keep the CPU
cloud-model variant if you need to scale the app out independently of the model.
A.6 Helm-ify (optional)
If Rafay prefers Helm, parameterise the above as values.yaml
(image.repository/tag, model.name, guard.requireKey, redis.url,
replicaCount, gpu.enabled) over the same Deployment/Service/Secret templates,
or ask Votal for the packaged chart. The identity-plane chart at
deploy/helm/shield-identity/ is a structural example (it is not the
guardrail chart).
A.7 In-cluster checklist
- Image registry path + model endpoint/token +
SHIELD_ADMIN_KEYreceived from Votal. - Redis reachable in-cluster;
REDIS_URL/UPSTASH_*set. SHIELD_GUARD_REQUIRE_KEY=enforceset.- Service reachable from Rafay’s serving path; base URL pointed at it.
- GPU node pool present (GPU image only); VRAM/model sized with Votal.
- Smoke test: a benign prompt passes, a policy-violating prompt returns
action=block, and/v1/tenant/me/telemetryshows the decision.
A.8 OpenShift (Rafay runs OpenShift)
Ready-to-apply OpenShift manifests live in the repo’s openshift/ directory:
| Manifest | Deploys |
|---|---|
openshift/redis.yaml |
Redis (tenant store + policy cache) — the dependency from A.2 |
openshift/shield-guardrail-vllm.yaml |
The guardrail data plane with in-cluster vLLM (GPU) — Deployment + Service + Route + model-cache PVC |
openshift/shield-admin.yaml |
(optional) the admin/tenant portal (Route, port 8080) |
The vLLM manifest is a thin wrapper — the image’s entrypoint
(scripts/start_vllm.sh) already launches vLLM on :8000, waits for it, then
runs the guard API; the manifest only schedules that image on a GPU with the
volumes, Route, and env it needs. Model runs in-cluster (full residency).
oc new-project shield || oc project shield
oc apply -f openshift/redis.yaml
# edit the Secret (REDIS_URL, SHIELD_ADMIN_KEY, optional HF token) first, then:
oc apply -f openshift/shield-guardrail-vllm.yaml
oc get route shield-guardrail -o jsonpath='{.spec.host}' # Rafay's base URL
OpenShift specifics baked into the manifest (and why):
- App on port 8080, not 80 — the restricted SCC can’t bind privileged ports.
/dev/shmas an in-memoryemptyDir— vLLM/NCCL need more than the 64 MB default.- Model-cache PVC (
HF_HOME) so the 4B model isn’t re-pulled on restart. Routetimeout 120s — Tier-2/agentic calls exceed the 30s default.- GPU toleration +
nvidia.com/gpu: 1— needs the NVIDIA GPU Operator installed. - fp8 by default (L40S/L4/H100); set
VLLM_QUANTIZATION=noneon A100/V100/T4. - SCC: if the vLLM image needs root,
oc adm policy add-scc-to-user anyuid -z shield-guardrail -n shield.
For the no-GPU option (app calls a remote model), use the llm-shield-cloud
image with SKIP_VLLM=true + LLM_BACKEND_URL instead — same Deployment shape,
no GPU, no model-cache PVC.