Artificial intelligence at home: is a superworkstation better, or a 15cm box?
The competition between small supercomputers using NVIDIA’s AI-dedicated processors and workstations equipped with the RTX 5090 is about much more than sheer processing power; it is primarily about control over data, costs and in-house expertise
There is a question that has been coming up repeatedly in IT meetings at many Italian companies over recent months – from medium-sized software houses to engineering firms, and including R&R&D departments in the manufacturing sector: is it better to continue renting cloud computing power to experiment with artificial intelligence, or has the time come to buy it and keep it in-house?
Cloud computing bills that are rising faster than expected, privacy officers asking where the company data fed to AI actually ends up, and technical teams wanting to be able to use a model without having to wait for a GPU to become available in the cloud are all excellent reasons to conclude that it is better to do it yourself, but building an AI infrastructure in-house is far from a trivial task, and the hardware required is very expensive. This is why many companies have returned to experimenting, and there are essentially two paths to follow: setting up workstations with the latest-generation Nvidia graphics cards (the famous RTX 5090s, which are, however, extremely expensive) or using computers based on Nvidia’s Spark platform – small boxes with 128GB of RAM and a processor designed to deliver AI inference from the Blackwell family. We tested an Asus Ascent GX10 for a few weeks and realised just how different these options are from one another, and that there isn’t one solution that’s always better than the other.
The Elephant and the Cheetah
To understand the difference between the two approaches, an image is more helpful than a table of technical specifications: the ASUS GX10 is like an elephant with an ironclad memory, capable of keeping a mental map of all the waterholes on the savannah at once, but not particularly quick when it comes to moving from one to the next; the workstation with the RTX 5090 is the cheetah that burns up the track, but only over short distances. The main difference lies in the RAM installed. For AI models to run smoothly, they must all fit within the same memory, as close as possible to the processing unit. The architecture of the Nvidia GB10 SoC, which forms the basis of the Asus Ascent GX10 (as well as many similar systems from other manufacturers), boasts a full 128GB of RAM that can be used to host AI models. Systems based on the RTX 5090, on the other hand, can only utilise the 32GB of ultra-fast DDR7 RAM on board the graphics card. If the model is too large, it won’t fit, and the race doesn’t even get started.
In fact, the ASUS Ascent GX10 is designed to allow two identical units to be daisy-chained via the NVIDIA ConnectX-7 network and to handle even larger models, such as Llama 3.1 with 405 billion parameters – something that would be impossible for a single consumer workstation, however well-equipped it may be. On the other hand, with its 32 GB of memory, the RTX 5090 is still capable of running models such as Llama 3 (with 70 billion parameters) at Q4–Q6 quantisation levels, DeepSeek R1 in Q4, and video generation models such as Wan 14B, provided one is willing to accept certain compromises to make them fit.
In practical terms, for businesses this means that if the aim is to build an in-house virtual assistant based on a high-end open-source model, without compromising on the quality of the responses, the GX10 (or a cluster of two units) offers greater flexibility. If, on the other hand, the project involves the fine-tuning of smaller models, the generation of images or videos, or tasks requiring intensive and repetitive computing cycles — in other words, brute force rather than capacity — the RTX 5090 proves to be significantly more responsive.
