Building sovereign defense AI: How DSO built a custom mixture-of-experts (MoE) model.
Settore pubblico
DSO National Laboratories partnered with Mistral to build a custom large language model (LLM) using Forge, Mistral’s end-to-end training system.

Key takeaways:
Accelerated multilingual mixture-of-experts (MoE) model development
Built on Forge, Mistral’s end-to-end system for training custom frontier models
"Partnering with Mistral gave us direct access to expertise in MoE architecture, helping us accelerate our model development timeline." — Dr Hai Leong Chieu, Distinguished Member of Technical Staff, DSO National Laboratories (DSO)
DSO National Laboratories is Singapore’s largest defense R&D organization, with more than 1,800 engineers and scientists working on capabilities for Singapore’s national security. As AI became increasingly important to defense operations and mission requirements grew more complex, DSO needed to train its own multilingual large language models from scratch that could be deployed entirely on premise. To accelerate development, DSO partnered with Mistral to co-develop a specialized multilingual MoE model using Forge, Mistral’s end-to-end system for developing and training custom frontier models.
The case for self-reliant AI.
Singapore’s defense establishment faces an increasingly complex operational environment, with teams required to analyze vast amounts of data under significant time pressure. For Dr. Chieu, who leads Natural Language Processing and LLM research at DSO’s Information Division, the mission was to build self-reliant AI that could operate independently of any external vendor or supply chain disruption.
“Our strategy has always been to develop the capability to train our own large language models from the ground up, “ Dr. Chieu explained. "We want to be in a position to deliver performant, on-prem AI solutions to our customers."
DSO's team had been pre-training its own models and tracking open source research for years. The foundation was strong, but mastering the MoE architecture, which DSO had identified as critical to the next generation of LLMs, required practical, at-scale experience the team had not yet built.
Why Mistral and why MoE.
DSO chose Mistral for its track record. Mistral had already built and deployed MoE models in production, at scale, giving DSO direct access to experience that mattered for this effort.
"We recognised early on that MoE architecture was where the field was heading,” Dr. Chieu recalled. “Mistral had already demonstrated real-world success with this architecture, and that track record was exactly what we needed in a collaborator."
The two teams co-developed a MoE model using Mistral’s model development stack, later integrated into Forge. Training data included language content sourced via AI Singapore (AISG). Both teams worked in a shared development environment, which supported faster iteration and feedback.
More than the technical setup, Dr. Chieu valued the knowledge transfer. "The pace of research in this field means that what gets published often lags behind what actually works in practice," he said. "Collaborating with Mistral gave us a direct line to that knowledge, helping us focus on proven approaches that deliver results."
From custom model training to edge AI.
The custom MoE model delivered leading performance in its class on Southeast Asian language benchmarks, meeting the performance threshold for DSO’s operational use cases. Beyond the numbers, the collaboration advanced DSO’s internal capabilities, both in what the team built and in what they can now build independently.
Building on that foundation, DSO and Mistral have since launched a second phase of work. "Our next challenge is building small, high-performing reasoning models that can run efficiently at the edge on autonomous systems operating in the field,” Dr. Chieu said. “It is an ambitious goal, and one we are confident in pursuing together with Mistral."
The work reflects a broader trajectory, from establishing core LLM training capabilities to developing AI for demanding, resource-constrained environments.




