
Specialized decision-tree inference engine delivering predictable JSON outputs, deterministic workflow evaluations, and ultra-low per-token costs for automated business logic.
Decisions API Beta focuses directly on deterministic reasoning loops, structured schema validation, and instant branch resolution. Unlike sprawling general-purpose LLMs that introduce excessive token overhead during plain classification tasks, this specialized engine processes multi-variable business rules and JSON payloads with strict conformance to developer-defined TypeScript and OpenAPI specifications.
Architected for backend microservices, real-time fraud scoring, support triage automation, and complex state machine orchestrations, the gateway guarantees predictable latency profiles even under peak burst loads.
| Model Version | Decisions-Engine-v1.4-Beta |
|---|---|
| Maximum Context Window | 128,000 Tokens (128K) |
| Throughput Quota | Up to 25,000 RPM / 2,000,000 TPM |
| Authentication Mode | Bearer Token via Hardware-Isolated KMS |
Select your required monthly token tier to provision high-throughput API keys with guaranteed SLA routing.
The model uses constrained decoding at the sampling layer, ensuring every generated token strictly conforms to the supplied JSON schema definition without needing post-processing or retry logic.
Typical median latency sits around 110ms to 140ms for standard classification and branching payloads under 4K input tokens when routed through our global edge cluster.
Yes, throughput quotas adjust automatically based on account billing tiers, with dedicated rate limiter policies to prevent unexpected throttling.