Early access

Cost and reliability
controls for your
LLM traffic

One key per team, hard budget caps, Slack alerts, provider fallbacks, semantic cache and Grafana dashboards. Bring your own provider keys.

app.py
client = OpenAI(
    base_url="https://gateway.velar.run/v1",
    api_key="vg_payments_prod_…",   # one key per team
)

Change two lines. Keep your models, keep your provider accounts.

Why it matters

5%

of LLM calls in production error out, and 60% of those errors are rate limits.

52%

of companies have nobody who owns AI spend.

93%

exceed their AI budget.

Sources: Datadog LLM Observability, 2025 · FinOps Foundation State of FinOps, 2026

What you get

Six controls. One base_url.

Hard budget caps

Monthly or daily budgets per team and app. Over the cap, the call is rejected with a clear error, not a surprise invoice.

One key per team

Every call attributed to team, app, model and environment. Reports per dimension, per day.

Alerts that say what happened

Slack, email or webhook at 80% burn, on error spikes and on rate-limit storms — with the team, app and top models in the message.

Fallbacks and retries

Provider fallback chains, retries with backoff and per-key rate limits, configured by someone who has run production inference.

Semantic cache

Exact-match and semantic caching per app. Repeated questions never reach the provider.

Grafana dashboards you keep

Spend by team, error and rate-limit map, latency percentiles, cache hit rate, budget burn-down. Exportable JSON.

How it works

Three steps to the first dashboard.

1

Point your base_url at Velar

OpenAI-compatible. No SDK change, no code change beyond configuration.

2

Set budgets per team

Give each team or app a key and a monthly cap. Pick where alerts go.

3

Watch it in Grafana

Dashboards and alerts from day one. A monthly cost review on Scale.

Bring your own provider keys — you keep paying OpenAI, Anthropic or Google directly. Logging is off by default. Need prompts to stay inside your cloud? See the In-your-cloud plan.

Works with LLM, speech and image APIs: OpenAI · Anthropic · Google · AWS Bedrock · Azure OpenAI · ElevenLabs · Deepgram · Cartesia · Replicate · fal

Pricing

Flat monthly fee. Never a percentage of your tokens.

Team

$199per month
  • Up to 5 team/app keys
  • Budgets with hard caps
  • Slack and email alerts
  • Exact and semantic cache
  • Fallback chains
  • Grafana dashboards
  • 30-day log retention
  • 1M requests/mo
  • Email support
Request early access

Scale

$599per month
  • Unlimited keys
  • SSO (Cloudflare Access)
  • Audit export
  • 90-day log retention
  • 10M requests/mo
  • Monthly 30-minute cost review
  • Slack support
Request early access

In your cloud

from $2,000per month
  • LiteLLM + Postgres + Grafana deployed in your VPC via Helm/Terraform
  • Operated by Velar
  • SOC 2 / SBP evidence pack
  • Bridge to Managed Reliability & Cost (from $6,000/mo)
Request early access

Early access: the first 10 teams get 50% off for 6 months.

Early access

Request early access

Tell us what you run. You will hear from us within 24 hours.

Providers you use
LLM
Speech & image

FAQ

Common questions

Do you see my prompts?
Logging is off by default: Velar stores metadata (team, model, tokens, cost, latency), not content. You can turn on redacted or full logging per team. If prompts cannot leave your cloud, the In-your-cloud plan runs the same control plane inside your VPC.
How much latency does it add?
The gateway runs at Cloudflare's edge. The budget for added latency is under 50 ms at p95; cached responses return in milliseconds.
What happens if Velar is down?
Revert base_url to your provider and your traffic flows as before. Nothing else to undo. Fail-open per app is a setting.
When can I start?
Early access opens in cohorts of up to 10 teams. You will hear from us within 24 hours of requesting access, and the first cohort starts when the design-partner group is full.