Cirrascale Inference Platform: Multi-Vendor AI Routing Across NVIDIA, AMD, Qualcomm and Tenstorrent


Cirrascale Cloud Services released the production version of its Cirrascale Inference Platform on September 15, 2026. The managed inference stack can route workloads across NVIDIA, AMD, Qualcomm and Tenstorrent accelerators, while supporting open-source models, customer-owned private models and closed model ecosystems including Google Gemini delivered on-premises through Google Distributed Cloud.

The platform is available across Cirrascale's U.S. and international regions. Cirrascale positions it as a serverless enterprise inference layer for search, chatbots, copilots, agentic services, coding assistants, document intelligence and video generation, with a common web console for model deployment and governance.

For infrastructure teams, the main architectural change is hardware abstraction. Cirrascale says requests can be routed to a selected model and accelerator without application-code changes when switching among supported hardware. That creates a practical multi-vendor deployment option for organizations evaluating inference cost, capacity and accelerator availability across different silicon stacks.

What the production release includes

The production platform combines several functions that are often deployed separately:

Capability Production release
Accelerator support NVIDIA, AMD, Qualcomm and Tenstorrent
Model types Open-source, customer-private and supported closed ecosystems
Private Gemini Google Gemini through Google Distributed Cloud
Request routing Model and accelerator selection layer
Fine-tuning Private-data fine-tuning inside the customer's environment
Applications Private chat and knowledge-base integration
Operations Web console, usage visibility and spend controls
Agent controls Governance guardrails for agentic workloads
Deployment footprint Cirrascale U.S. and international regions

Cirrascale says its optimization layer is designed to maximize throughput and stabilize latency as demand changes. A standardized independent benchmark set was not published with the production announcement; accelerator-to-accelerator performance and cost should be evaluated against the customer's actual model, context length, concurrency and latency target.

Multi-vendor routing is the key infrastructure feature

AI inference infrastructure increasingly spans hardware with different memory capacities, software ecosystems, availability profiles and economics. A deployment platform that can target several accelerator families gives operators another layer at which to manage those differences.

Cirrascale's model and hardware selection layer is designed to choose an available accelerator for a workload and move supported requests among NVIDIA, AMD, Tenstorrent and Qualcomm hardware without requiring the application to be rewritten for each change.

That capability can matter in three common situations:

  • Capacity planning: a team can use more than one accelerator supply pool instead of tying every service to one hardware family.
  • Workload matching: different models or latency profiles can be assigned to hardware selected for the deployment requirement.
  • Procurement flexibility: organizations can evaluate cost per token and availability across accelerator vendors while keeping a common inference-management layer.

The public production material leaves the detailed routing policy—inputs, weighting rules and the complete model-and-accelerator compatibility matrix—to deployment evaluation. Teams should validate the exact combinations they intend to deploy.

Private Gemini adds a different deployment path

Cirrascale also supports Google Gemini on-premises through Google Distributed Cloud. The arrangement is aimed at organizations that want access to Google's closed model ecosystem while keeping the deployment inside private infrastructure operated with Cirrascale.

Cirrascale announced this Google Distributed Cloud integration earlier in 2026 and now includes it in the production inference platform alongside open and customer-owned models. This gives the platform a broader model boundary than a conventional GPU cloud focused only on deployable open weights.

For regulated or data-sensitive environments, the relevant deployment questions include where prompts, retrieved documents, fine-tuning data, logs and model outputs are processed and retained. Those controls should be mapped directly to the organization's specific Cirrascale and Google Distributed Cloud configuration.

Fine-tuning and enterprise controls

Cirrascale says customers can fine-tune models on private data while keeping that data inside their environment. The platform also includes a private chat interface connected to enterprise knowledge bases, team-level AI spending controls and governance features for agentic workloads.

The company describes the platform as aligned with HIPAA, SOC 2 and FedRAMP requirements where applicable. Compliance remains deployment-specific: organizations still need to map the service configuration, data flows and contractual controls to the requirements that apply to their workload.

How to evaluate it against a single-vendor GPU stack

Operational flexibility is the central evaluation case for Cirrascale's platform. A useful proof of concept should run the same production-shaped workload through the accelerator options under consideration and record:

  • input and output tokens per second;
  • time to first token and end-to-end latency;
  • p95 and p99 latency under concurrency;
  • cost per million tokens or equivalent workload unit;
  • model-loading and scaling behavior;
  • maximum practical context and batch size;
  • routing behavior during capacity pressure;
  • observability and spend-control coverage;
  • data residency and network path requirements.

These measurements expose whether hardware abstraction translates into a real cost, availability or operational advantage for the intended workload.

Availability and pricing

The Cirrascale Inference Platform is available now across the company's U.S. and international regions. The production announcement directs customers to request a demonstration and does not publish a single inference-platform price schedule.

Cirrascale separately advertises flat-rate monthly billing and no ingress or egress data fees for its cloud services. Commercial evaluation should use a workload-specific quote because accelerator choice, deployment topology, managed services and private-model requirements can materially change total cost.

Sources

Bottom line

The production release turns Cirrascale's earlier inference-platform preview into an available multi-vendor service. Its differentiating technical proposition is a common inference layer spanning NVIDIA, AMD, Qualcomm and Tenstorrent hardware, combined with private-model support, fine-tuning, governance controls and an on-premises Gemini path through Google Distributed Cloud.

The next decision for infrastructure teams is empirical: benchmark the same representative workload across the available accelerator choices and compare latency, throughput, cost, capacity and data-governance requirements. Cirrascale's routing layer is most valuable when those measurements show that workloads genuinely benefit from moving across more than one hardware stack.