Aolani, a Singapore-founded neocloud, has announced the launch of the Aolani Token Factory, a managed inference platform designed to enable organizations to deploy and scale AI models on a pay-per-token basis without the need to provision or manage underlying GPU infrastructure. This move positions Aolani as the first Singapore-founded neocloud to offer production-grade, managed inference at scale.
The launch comes at a time when global AI companies are expanding operations in Singapore and businesses worldwide are investing in AI to drive tangible outcomes. The demand for production-grade inference infrastructure is accelerating rapidly. Aolani Token Factory aims to close the accessibility gap, providing AI-native companies and enterprises a compliant and high-performance path from experimentation to production-scale deployment.
The platform operates on a per-token metering model, allowing customers to pre-purchase credits and pay based on token consumption, rather than investing in capital-intensive GPU infrastructure. Aolani manages the entire inference stack, including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization. This enables customers to scale consumption without continuously provisioning additional infrastructure.
At launch, the platform supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand the model catalogue based on customer demand. Customers can also deploy their own models through OpenAI-compatible APIs. For enterprise clients with strict compliance and data residency requirements, dedicated capacity and data isolation options are available.
The Aolani Token Factory is tailored for three core production use cases: AI agents for high-volume inference and workflow automation, enterprise AI applications like internal copilots and knowledge assistants, and coding agents for code generation, completion, testing, and review.
Sea Xu, Applied AI Research Lead at Aolani, emphasized the platform's design: "The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia's AI ecosystem evolves and grows rapidly, it is our goal to ensure that the infrastructure serving it keeps pace."
Nicholas Chia, Chief Executive Officer at Aolani, highlighted the strategic importance: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute."
Interested parties can register interest at Aolani Token Factory. This development underscores the growing trend of AI infrastructure services that lower barriers to entry, enabling more organizations to leverage advanced AI capabilities.

