Skip to main content
ai gateway

Anthropic

curated models
3
credential
API key

How does ZopNight track Anthropic spend?

ZopNight registers 3 Anthropic models through the AI Gateway: claude-sonnet and claude-opus as strong tiers, claude-haiku as the cheap tier. Each maps to the vendor's `-latest` alias, so a model upgrade lands without reconfiguration. Spend is metered per request against per-team budgets.

Models registered by default

Anthropic models the gateway registers.
Model Tier Maps to
claude-sonnetstronganthropic/claude-3-5-sonnet-latest
claude-haikucheapanthropic/claude-3-5-haiku-latest
claude-opusstronganthropic/claude-3-opus-latest

Claude models aliased to -latest

Three, each mapped to a -latest alias rather than a pinned version:

  • Strong: claude-sonnetclaude-3-5-sonnet-latest, claude-opusclaude-3-opus-latest
  • Cheap: claude-haikuclaude-3-5-haiku-latest

Aliasing to -latest means a vendor upgrade reaches your traffic without a config change. That is convenient and it is a tradeoff: the model behind a name can change under you, so pin explicitly if you need reproducibility.

Your Claude prompts never reach ZopNight

ZopNight never sits in your request path. The gateway runs on hosted LiteLLM; ZopNight configures it, meters it and enforces the budget, but your prompts and completions do not traverse our infrastructure. That distinction is why per-request cost attribution is possible without us reading request bodies.

Per-team Anthropic keys with USD caps

Each team gets a virtual key with a hard USD budget and its own model allow-list. When a key hits its cap it stops working. This is a limit, not an alert. Keys rotate and revoke independently, so a leaked key is contained to one team’s budget rather than the org’s.

Routing Claude traffic between strong and cheap

Three postures decide which model serves a request: Cost-saver, Balanced and Quality-first. Requests are classified into four semantic categories (coding, analysis and reasoning, extraction and classification, chit-chat) and routed accordingly, so cheap models take the traffic that does not need an expensive one.

A semantic cache sits in front, with strictness settings from strict (0.98 similarity) to loose (0.90). On repetitive workloads the cache is usually a larger saving than the routing itself.

faq · ai gateway

Anthropic: common questions

Which Claude models does ZopNight register?

Three: claude-sonnet and claude-opus in the strong tier, claude-haiku in the cheap tier.

Can I pin a specific Claude version?

The defaults map to the vendor's -latest aliases, so an upgrade reaches your traffic with no config change. That is convenient and it is a tradeoff, because the model behind a name can shift under you. Pin explicitly when you need reproducibility.

Does ZopNight see my Claude prompts?

No. The gateway runs on hosted LiteLLM, which ZopNight configures, meters and budgets, but prompt and completion bodies do not traverse ZopNight infrastructure.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

417 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

417 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·