Scality launches AI Inference Factory to bring enterprise AI on-premises
Scality’s new open-code software stack enables enterprises to run validated open-weight AI models on their own
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
SAN FRANCISCO, Oct. 08, 2026 (GLOBE NEWSWIRE) — Scality, a global leader in data infrastructure software for AI storage at scale, today announced the availability of Scality AI Inference Factory, an open-code software stack for deploying and operating AI inference on enterprise-owned infrastructure. The solution gives enterprises, government agencies and neo-cloud providers a supported alternative to cloud-based AI services, without having to assemble, integrate and maintain the entire software stack themselves.
As AI moves from experimentation into production, organizations are looking for more predictable costs, greater control over model lifecycles and increased data sovereignty. For most organizations, the answer will be hybrid, with the most critical AI processes running on-premises. Running inference locally can address these requirements while keeping sensitive data within infrastructure that organizations own and operate.
Scality AI Inference Factory brings together validated open-weight models, a disaggregated inference-serving layer, a control plane for authentication, metering, routing and scheduling, and Scality AI Data Infrastructure (ADI), which uses policy-governed autonomous operations to manage massive datasets across the AI lifecycle. Scality validates, ships and maintains the integrated stack.
“With AI moving into mission-critical production environments, organizations need greater control over where inference runs, how their models are managed and what happens to their data,” said Jérôme Lecat, CEO at Scality. “For 15 years, Scality has built data infrastructure that thousands of customers around the globe rely on to operate 24/7. AI Inference Factory brings that experience to on-premises AI, giving organizations the reliability and sovereignty they need to run critical AI workloads on their own terms.”
Cloud AI challenges drive interest in on-premises inference
Enterprises scaling AI in production are confronting challenges with the cloud-based inference model. With per-token pricing, the more useful an AI workflow becomes, the more it costs, producing bills that are difficult to forecast and impossible to cap. Model versions and quantization can change at the provider’s discretion, affecting the workflows built on them. Organizations with sensitive or regulated data face sovereignty concerns over where their prompts, documents and proprietary code are processed, under which jurisdiction, and who can restrict access to them.
Scality AI Inference Factory is designed for organizations that want to run critical AI workloads on infrastructure they control while retaining flexibility across models, hardware and software. The stack supports specialized chatbots, AI-assisted software development, agentic applications and other inference-intensive workloads.
Scality AI Inference Factory combines:
- Validated open-weight models maintained as part of the supported stack;
- Disaggregated inference serving that separates prefill from decode so each can scale independently;
- A control plane that authenticates and meters requests, routes them to the appropriate GPUs holding the relevant context and schedules workloads against SLA targets; and
- Scality ADI, which provides high-performance object storage for models, enterprise data and inference state.
Scality validates and maintains the complete stack as a supported solution, helping organizations keep pace with rapidly changing models, serving technologies and security requirements without having to integrate and maintain the individual components themselves.
“As AI adoption expands and AI agents become more prevalent, organizations are evaluating which IT environments are best suited to run inference. On-premises deployments are coming into focus for a variety of reasons, including cost optimization, data security and privacy, and regulatory requirements. IDC research shows that running inference on-premises is a credible option for a growing set of enterprise and public sector workloads, provided the infrastructure can hold and serve model state efficiently at scale. Scality is among the vendors addressing that requirement across the inference and storage layers.” — Nataliya Yezhkova, Vice President, Storage and Data Management, Enterprise Infrastructure, IDC
Scality ADI extends AI inference beyond GPU memory constraints
Scality ADI serves as the shared storage layer for AI Inference Factory, extending key-value (KV) cache beyond limited GPU high-bandwidth memory (HBM). This allows inference infrastructure to preserve and retrieve model context instead of requiring GPUs to recompute it, improving GPU utilization and reducing inference costs.
This is particularly important for reasoning models and agentic workloads, where KV cache can quickly exceed available GPU memory. Scality ADI provides a shared, multi-petabyte cache that GPUs read at latency the same order of magnitude as GPU memory, allowing them to retrieve existing context and remain focused on inference.
Scality testing demonstrated that ADI can reduce GPU consumption while maintaining the same level of performance:
- 1.9-second load time for Gemma-3 27B, almost 10x faster than local NVMe, with the model streamed in parallel across the cluster over RDMA;
- 166 ms warm time-to-first-token on a 14K-token context restore from ADI, only 83 ms behind HBM;
- 14x faster KV cache retrieval than recomputation on a 14K-token context and up to 72x faster on a 439K-token context;
- A KV cache more than 80x larger than a single GPU’s memory, keeping up to 1,000 concurrent sessions resumable without recomputation;
- No measurable impact on token generation, as context is restored before the first token; and
- 97% of network line rate for data transfer between GPUs and storage.
The architecture separates prefill from decode, with ADI acting as the shared cache between GPU pools. This allows available decode GPUs to access existing context without tying it to a specific GPU server. Separating prefill from decode has documented gains: DistServe (OSDI 2024) measured up to 7.4 times more requests served within the same latency targets. Adding a shared KV cache pool on top, Moonshot AI’s Mooncake reported 75 percent more requests on Kimi’s production traffic.
As Giorgio Regni, Scality CTO, puts it: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with the GPUs staying busy the whole time. Storage is no longer the reason to keep the KV cache inside the GPU server.”
Regni continues: “That opens the door to disaggregated serving. One pool of GPUs handles prefill, processing the prompt and writing the KV cache to ADI. A second pool handles decode, reading the cache back and generating tokens. With Scality ADI as the shared cache, any decode GPU can pick up any context, and there is essentially no size limit. Prefill no longer interrupts decode, and both pools run at full load.”
For a technical deep dive into the AI Inference Factory architecture and test results, visit Scality’s companion blog.
Open architecture provides flexibility across models, frameworks and hardware
Scality AI Inference Factory is designed to preserve customer choice across the AI infrastructure stack. It is compatible with open-source harnesses and agentic frameworks such as OpenCode, Hermes, Goose, LangGraph and Pydantic AI, and runs validated open-weight models such as Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek. The software deploys on standard servers from Dell, HPE, Lenovo and Supermicro.
Every software component Scality ships is delivered as open code, enabling customers to inspect how inference state is stored and moved, and to submit contributions for Scality review. For organizations running sensitive, regulated or sovereign AI workloads, this provides transparency into the software handling their models and data rather than requiring them to rely on a closed external service.
Scality ADI integrates with standard AI stacks and deploys at multi-petabyte scale on NVMe, with RDMA access to HDD for virtually unlimited KV cache capacity. A single namespace automatically tiers data across TLC, HDD and, optionally, tape, allowing organizations to align storage performance and economics with different stages of the AI data lifecycle.
Availability
Scality AI Inference Factory is available now as a software license or as a fully managed service. Scality is demonstrating the solution and its reference architecture at Scality Day on October 8 in Paris. To learn more, visit the Scality AI Inference Factory web page and detailed companion blog.
About Scality
Scality builds data infrastructure software for enterprise AI, cyber resilience, and sovereign control at multi-petabyte to exabyte scale. Its platform reduces operational burden while aligning the right storage media, performance, and protection to each workload through human-approved lifecycle policies. Built on CORE5 cyber resilience and open-code principles, Scality software delivers extreme performance, operational simplicity, and sustainable economics. The world’s most demanding enterprises and government organizations depend on Scality to power AI initiatives, defend critical data, and build infrastructure that lasts for decades. Recognized as a leader by Gartner, Scality is where AI remembers, learns, and thinks. Follow us on LinkedIn.
Media contact
Erin Jones
Avista Public Relations for Scality
805.440.6587
scality@avistapr.com
A photo accompanying this announcement is available at https://www.globenewswire.com/NewsRoom/AttachmentNg/83b57162-b956-456e-9995-27c24a20aa45


