Singapore-based neocloud provider Aolani has announced the launch of the Aolani Token Factory, a managed inference platform designed to let organizations deploy and scale AI models on a pay-per-token basis without managing GPU infrastructure. The company claims this makes it the first Singapore-founded neocloud to offer production-grade managed inference at scale, targeting a growing demand for efficient AI deployment in the region.
The platform addresses a critical bottleneck for AI adoption: the high cost and complexity of provisioning and managing GPU infrastructure. By offering per-token metering, customers can prepurchase credits and pay based on usage, avoiding capital-intensive investments. Aolani handles the entire inference stack, including GPU allocation, model serving, orchestration, scheduling, and workload optimization, allowing companies to scale consumption flexibly.
At launch, the Token Factory supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand based on demand. Customers can also deploy custom models via OpenAI-compatible APIs. For enterprises with strict compliance and data residency requirements, dedicated capacity and data isolation options are available.
The platform targets three primary use cases: AI agents for high-volume inference and workflow automation, enterprise applications like internal copilots and knowledge assistants, and coding agents for code generation and review. This broad applicability positions it as a versatile solution for various industries.
Sea Xu, Applied AI Research Lead at Aolani, emphasized the platform's high-performance inference stack and its role in keeping pace with Southeast Asia's rapidly evolving AI ecosystem. 'We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge,' she said.
Nicholas Chia, CEO of Aolani, highlighted the strategic importance of the launch: 'Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity.'
The Token Factory could significantly lower barriers to AI adoption in Asia, where many companies are exploring AI but lack the infrastructure expertise or resources. By shifting from capital expenditure to operational expenditure, it enables faster experimentation and deployment, potentially accelerating innovation across sectors. For global AI companies expanding in Singapore, it offers a compliant and high-performance path to production.
As AI inference demand grows, managed services like this may become critical for scalability. Aolani's move could also intensify competition among cloud providers in the region, benefiting customers with more options and better pricing. For business leaders, this development signals a trend toward more accessible AI infrastructure, enabling them to focus on applications rather than hardware.
Interested parties can register interest at the company's website. The launch marks a milestone for Aolani and the regional AI ecosystem, promising to reshape how organizations access AI compute.
