CoreWeave Forge Connects AI Inference, Observability, Training and Evaluation


CoreWeave launched CoreWeave Forge on September 30, 2026 as a development layer that connects model and agent inference, production observability, data curation, training and evaluation. The company says Forge works across models, frameworks and cloud environments, including workloads that live outside CoreWeave Cloud.

The release turns several parts of CoreWeave's expanding software portfolio into one development workflow. Teams can run a model or agent, inspect production traces, convert failures into datasets and evaluations, improve the system, and test the next version before deployment. CoreWeave says MasterClass and Canva are already building on Forge.

For AI engineering teams, the practical value is workflow continuity. Production traces and evaluation results can feed the next training or agent iteration without moving the evidence between separate systems, while teams can keep model and infrastructure choices open.

What Forge includes

CoreWeave describes the platform as a five-stage loop: run, observe, curate, improve and evaluate. The launch brings together capabilities including inference, Agent Lens, Registry, Notebooks, Training, Sandboxes and ARIA.

Stage Forge role Practical use
Run Serverless and dedicated inference Serve open models, fine-tuned checkpoints and production agents
Observe Production traces and Agent Lens Find failure clusters, capability gaps and operational trends
Curate Dataset and evaluation construction Turn production evidence into reusable test and improvement data
Improve Training and agent iteration Fine-tune models or change prompts, tools and agent behavior
Evaluate Evaluation suites and judges Compare candidate changes before promotion

CoreWeave's Serverless Inference documentation says requests can be traced and evaluated through native observability features. Dedicated Inference connects live inference with training traces and evaluation results, giving teams a path from production behavior back into the next model version.

Agent Lens connects production failures to regression tests

A notable new component is CoreWeave Agent Lens, available in public preview on CoreWeave's hosted cloud. It analyzes agent traces to identify and group failures, then lets teams turn those clusters into evaluation sets.

CoreWeave also documents integration with coding agents through Skills and MCP. With appropriate permissions, an agent can retrieve a failure cluster, inspect representative traces, propose a change, make test calls, run an evaluation set and summarize the result for human review. Approval and promotion remain controlled by the team.

CoreWeave says Agent Lens insights and online LLM judges are free through the end of 2026. Dedicated and on-premises Agent Lens deployments are planned after the hosted-cloud preview.

Multi-cloud is part of the design

Forge is designed as a software layer that can span CoreWeave and external infrastructure. CoreWeave says customers can use any model or framework and operate across other clouds and private data centers, with portable improvement data and open interfaces between components.

That matters for teams whose inference, training and enterprise data already span multiple environments. Forge can act as the improvement layer while workloads remain distributed. Individual Forge services retain their own hosting and availability boundaries.

The platform also preserves a direct path into CoreWeave's GPU infrastructure for teams that want the development layer and compute stack together. CoreWeave's September 30 infrastructure announcements included production availability of NVIDIA Vera Rubin NVL72 systems, illustrating the broader strategy of coupling model-development software with accelerated infrastructure.

Where Forge fits against a conventional MLOps stack

A conventional production AI stack often combines separate systems for experiment tracking, model serving, tracing, dataset management, evaluation and training. Forge's main architectural proposition is to keep the artifacts generated by those stages connected.

This is especially relevant to agents, where a production failure may involve a model response, tool call, prompt, retrieved context and multi-step trace. Preserving that evidence through diagnosis, evaluation and regression testing can reduce the manual work needed to reproduce failures and validate a fix.

The platform remains usable with external models, frameworks and infrastructure, so adoption can begin at the workflow layer without an immediate migration of every compute workload.

Availability and deployment considerations

CoreWeave says everyone can start building with Forge now, with a 30-day Forge Pro trial and credits across product lines. Agent Lens is specifically described as a public preview on the hosted CoreWeave cloud.

Teams evaluating Forge should map which parts of their existing stack they intend to replace or connect, then test whether production traces, datasets and evaluation suites remain portable enough for their operating model. The strongest fit is likely to be organizations running a continuous model or agent improvement cycle where production evidence regularly feeds new evaluations and training work.

Bottom line

Forge is CoreWeave's move from supplying AI infrastructure and individual development products toward an integrated model-and-agent lifecycle. Its differentiator is the feedback loop: production behavior becomes evidence for evaluation and improvement, while the resulting model or agent version returns to serving in the same environment.

The launch is also strategically significant for CoreWeave. A connected software layer can sit above its GPU cloud while remaining useful to teams with workloads on other infrastructure, giving the company a broader role in AI development than compute capacity alone.

Sources