AI might often be the icing, but you still have to bake the cake
Manoj Prasanna Kumar, CTO of Singtel group’s Digital Infrastructure company, talks to Contributing Editor Annie Turner about pioneering GPU-as-a-Service and AI’s role in its operation, plus the company’s wider strategy to move up the value chain and create a virtuous circle. 
One of the largest communications technology groups in Asia, Singtel is renowned for being one of the most innovative telcos in the world. Manoj Prasanna Kumar is CTO of the group’s digital infrastructure arm, Digital InfraCo, and he is responsible for technology for all business lines. Digital InfraCo operates AI-ready data centres (DCs), AI clouds, the Paragon orchestration platform, satellites and subsea cables globally.
Singtel’s pioneering, home-grown, award-winning Paragon platform builds software to help telcos monetise their networks, such as through network slicing, edge and mobile edge compute, managing network quality of service and providing digital experiences for consumer customers. Recently, Paragon has evolved to also orchestrate GPUs and AI services to cater for the mission-critical AI workloads of enterprises and public sector customers.
In March 2024, Digital InfraCo announced it would launch a GPU-as-a-Service (GPUasaS) powered by NVIDIA’s H100 Tensor Core GPU clusters which it was already running in some data centres in Singapore. Also, it would be among the first in the region to deploy NVIDIA’s GB200 Grace Blackwell (GB) Superchips, which delivers 30X faster real-time large language model (LLM) inference than their predecessor.
Federating GPU capacity
Since the GPU cloud, known as RE:AI, was launched in October 2024, all the AI chips “are getting steady and high demand” according to Kumar. Digital InfraCo is the only company in Singapore to have deployed them along with cutting-edge liquid cooling technology. RE:AI is powered by its Paragon platform. While this means it cannot leverage GB capacity from hyperscalers within Singapore, it can federate capacity with hyperscalers that have deployed GB chips elsewhere for global customers that do not specifically require GPUs inside Singapore.
“This is where our software platform, Paragon, creates a differentiation,” Kumar says, “It runs our GPU control plane, allowing customers to discover capacity, not just from hyperscalers – we have other cloud partners with GPU capacity globally. We do this [federation] continuously to drive more business to our partners and ensure we can accommodate our customers’ needs as much as possible. It’s a mutual win. We provide a single window of access to global capacity to customers [while] acting as business catalysts to our partners.”
The rapid evolution of GPUs
Since the launch of GPUasaS, the evolution of GPU technology has accelerated: “Last year, we were very proud that we launched the GB200 Superchips, which consume about 155KW and are the most complex in the industry; earlier this year, NVIDIA announced chips that consume 600kW,” Kumar says.
He continues, “As power consumption increases triple fold, the power-consumption problem is three- times more complex, and the reaction time is many times shorter. This requires focus on greater automation to ensure that when deploying such high power-consuming equipment that generates high heat, the facilities are as robust as possible to ensure there’s no failure in the infrastructure.”
Digital InfraCo has deployed closed-loop automation in AI cloud operations to improve operational resilience but “You can never say you’re 100% covered by a fully closed-loop system because at any time we can discover something else,” Kumar notes.
“With tasks that are well oiled and repetitive, automation can help, largely in reducing human tasks. But for new things, automation – even AI – is data in, data out. Without similar instances in the past, AI is not magic box to learn how to predict and handle an incident.”
He stresses that, “We definitely need people to apply domain knowledge continuously, to ensure that AI doesn’t make any mistakes, wrong assumptions or have any bias based on what it sees in the training data. Automation is used with controlled adoption of AI, depending on the confidence levels concerning the incidents.”
Powering ahead
Kumar continues, “We see more complexity and benefits from AI operating in the data centre and GPU cloud environments…GPUs are high-risk infrastructure because of the amount of heat that they produce. Every modern GPU consumes well over 10kW per server; if you stack four or five servers in a rack, that’s about 50kW power consumption in one rack” and that is rising.
Digital InfraCo was the first to use liquid cooling to dissipate heat from GPUs and its success has boosted the liquid cooling market. Even so, if poor cooling flow occurs for any reason, GPUs can overheat or burn out, creating a considerable physical risk within data centres and problems for customers.
Hence reaction time has to be fast, ahead of an incident or when an incident happens. “This makes the closed loop system much more important; because of the much shorter reaction times and the SLAs – running time must be as high as possible to ensure customers don’t lose money,” Kumar says.
The company is using simulations to establish how systems behave in certain circumstances regarding heat generation and systems’ stability at various levels of utilisation. “We use a lot of automation to configure threshold-based reactions, to dial down the clock speed of GPUs, or shut down GPUs temporarily, reboot them to cool them, or increase the flow rate of liquid into the GPUs to remove heat effectively,” Kumar states.
Symbiosis of AI software and GPU infra
“GPU cloud is definitely a key focus area for us,” he adds. “We are expanding continuously with more and more newer models of GPUs, expanding our customer outreach, catering to the AI requirements of more customers within the country. We are also actively looking at regional expansion plans to see where we could easily replicate what we did in Singapore and cater for the sovereign demands of those markets as well.”
Kumar observes, “Three years ago, LLMs were not even a thing; today it has redefined everybody’s life. There is not a single person who doesn’t use ChatGPT to do their day-to-day job. Three years from now, there could be many more, better things and as more complex technologies emerge, they need more complicated underlying infrastructure to run them in a scalable manner. I think they are a symbiotic evolution – the evolution of AI software and underlying GPU hardware infrastructure.
Moving up the value chain and virtuous circles
The GPUasaS is part of Digital InfraCo’s wider strategy, to “continuously to move up the stack and up the enterprise value chain,” Kumar explains. “We provide customers with connectivity globally, in terms of subsea cables and satellites which is the foundational layer. We then moved up into AI DCs, offering liquid cooled data centres to cater to customers who want to bring their own AI chips. We progressed to build our own AI cloud RE:AI…to offer a turnkey AI Infrastructure-as-a-Service and AI-as-a-service to customers.”
Now Digital Infraco is in the throes of moving up another level to build “apps, models and platforms through our ecosystem partners to offer them as vertical, standardised solutions in combination with the GPUs and the network,” he says. “That is the next logical evolution for us, to have a play in the application space [to offer] popular apps that have been adopted by vertical customers in other markets and cross-pollinate them into Singapore.”
Kumar continues, “The more applications we help customers adopt, the more of their problems we can solve. As more customers use those apps, the more the underlying GPUs and consequently the underlying data centres are used. It’s symbiotic throughout the stack.”
There has been much talk at recent conferences about using agentic AI for market research. Kumar says while AI is useful, it is not always enough. “Like I said, AI is garbage in, garbage out. AI needs high-quality training data and access to facts to make recommendations and suggestions with high accuracy. Using AI for searching and summarizing facts for market research will yield only results based on what’s published and accessible to the AI models.
“AI may not know facts about companies or products or systems without significant online presence and would probably miss many such products and partners that interest us because they are in stealth mode. AI is often the icing on the cake, but we still have to bake the cake for AI with facts to make it palatable,” he concludes.
