Products

Artificial intelligence at home: is a superworkstation better, or a 15cm box?

The competition between small supercomputers using NVIDIA’s AI-dedicated processors and workstations equipped with the RTX 5090 is about much more than sheer processing power; it is primarily about control over data, costs and in-house expertise

Da sinistra a destra: GeForce RTX 5090 e Asus Ascent GX10

5' min read

Translated by AI
Versione italiana

5' min read

Translated by AI
Versione italiana

There is a question that has been coming up repeatedly in IT meetings at many Italian companies over recent months – from medium-sized software houses to engineering firms, and including R&R&D departments in the manufacturing sector: is it better to continue renting cloud computing power to experiment with artificial intelligence, or has the time come to buy it and keep it in-house?

Cloud computing bills that are rising faster than expected, privacy officers asking where the company data fed to AI actually ends up, and technical teams wanting to be able to use a model without having to wait for a GPU to become available in the cloud are all excellent reasons to conclude that it is better to do it yourself, but building an AI infrastructure in-house is far from a trivial task, and the hardware required is very expensive. This is why many companies have returned to experimenting, and there are essentially two paths to follow: setting up workstations with the latest-generation Nvidia graphics cards (the famous RTX 5090s, which are, however, extremely expensive) or using computers based on Nvidia’s Spark platform – small boxes with 128GB of RAM and a processor designed to deliver AI inference from the Blackwell family. We tested an Asus Ascent GX10 for a few weeks and realised just how different these options are from one another, and that there isn’t one solution that’s always better than the other.

Loading...

The Elephant and the Cheetah

To understand the difference between the two approaches, an image is more helpful than a table of technical specifications: the ASUS GX10 is like an elephant with an ironclad memory, capable of keeping a mental map of all the waterholes on the savannah at once, but not particularly quick when it comes to moving from one to the next; the workstation with the RTX 5090 is the cheetah that burns up the track, but only over short distances. The main difference lies in the RAM installed. For AI models to run smoothly, they must all fit within the same memory, as close as possible to the processing unit. The architecture of the Nvidia GB10 SoC, which forms the basis of the Asus Ascent GX10 (as well as many similar systems from other manufacturers), boasts a full 128GB of RAM that can be used to host AI models. Systems based on the RTX 5090, on the other hand, can only utilise the 32GB of ultra-fast DDR7 RAM on board the graphics card. If the model is too large, it won’t fit, and the race doesn’t even get started.

In fact, the ASUS Ascent GX10 is designed to allow two identical units to be daisy-chained via the NVIDIA ConnectX-7 network and to handle even larger models, such as Llama 3.1 with 405 billion parameters – something that would be impossible for a single consumer workstation, however well-equipped it may be. On the other hand, with its 32 GB of memory, the RTX 5090 is still capable of running models such as Llama 3 (with 70 billion parameters) at Q4–Q6 quantisation levels, DeepSeek R1 in Q4, and video generation models such as Wan 14B, provided one is willing to accept certain compromises to make them fit.

In practical terms, for businesses this means that if the aim is to build an in-house virtual assistant based on a high-end open-source model, without compromising on the quality of the responses, the GX10 (or a cluster of two units) offers greater flexibility. If, on the other hand, the project involves the fine-tuning of smaller models, the generation of images or videos, or tasks requiring intensive and repetitive computing cycles — in other words, brute force rather than capacity — the RTX 5090 proves to be significantly more responsive.

The cost – not just in euros

In terms of price, the two products fall into different price brackets, though they are not that far apart. The Ascent GX10 is available online from around €3,600 (a price subject to considerable fluctuation), which is, all things considered, a reasonable cost for the purposes for which it is designed: delivering privacy-preserving AI systems to small offices and providing a testing ground for those designing larger systems. A workstation with an RTX 5090 starts with the cost of the graphics card itself (which, at the time of writing, is a minimum of €4,400) and is then supplemented by the rest of the PC, taking into account power consumption that exceeds 1 kWh.

But the financial cost is only half the story. The other half concerns what is known within the company, using a term that is somewhat overused but effective, as the ‘extended total cost of ownership’: set-up time, the skills required to maintain the system, electricity consumption and the physical space occupied.

Here, the GX10 plays a card that, for many companies, carries more weight than one might think: its compact size. The device measures just 15 x 15 x 5 centimetres, fits easily on a desk next to the keyboard and, although it does tend to get warm under heavy use, it isn’t noisy.

The real reason why companies are turning to on-premises AI

At the moment, something is happening that has rarely been seen before: costs are no longer the driving force behind technological choices. What matters now, with a legislature that takes such a hard line on issues of sovereignty and a geopolitical situation that has literally destroyed everyone’s trust in one another, is the desire for control. With the Ascent GX10, ASUS promises a fully on-premises AI system, with predictable costs, enterprise-grade stability and no exposure of data to external parties. It is, of course, an advertising slogan, of course, but it taps into a need felt by many companies, especially those handling sensitive data (health, financial, industrial data covered by confidentiality agreements) and which cannot afford to send confidential information to a cloud API, however secure the supplier’s contractual promise may be. The same applies, with slight variations, to the workstation equipped with an RTX 5090: those who choose it often do so not only for the computing power, but because they want a development environment that does not depend on a stable internet connection, a renewable cloud subscription, or a provider that could change its contractual terms from one day to the next – and which, ultimately, can see what your code does.

It must be said, for the sake of completeness, that neither system truly replaces cloud infrastructure when it comes to training models from scratch on a massive scale, or managing workloads spread over several days: for that, as the system integrators themselves point out, H100 or H200-class GPUs in a data centre remain essential. But one day we will get there too, because the Internet has opened up the world and government interests have taken full advantage of it without any limits.

Copyright reserved ©

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti