Invisix, a semiconductor metrology company, has raised €20 million in an oversubscribed seed round. The company develops soft x-ray metrology platforms for high-volume, non-destructive measurement of complex structures within advanced chips. This technology is crucial for improving the production of semiconductors destined for AI and High-Performance Computing, addressing the growing challenges posed by device miniaturization and complexity.
Nvidia presented its DGX Spark roadmap at Computex 2026, targeting laptops and desktop PCs. The plan outlines three future generations, including the Rubin platform, which will integrate LPDDR6 memory, and the subsequent Rosa Feynman. This initiative extends Nvidia's AI computing capabilities to client devices, focusing on performance and efficiency for local workloads, including the RTX Spark line.
Computex highlighted the expanding influence of AI, bringing CPUs and ASICs to the hardware forefront alongside GPUs. This evolution presents new opportunities and challenges for on-premise deployment strategies, prompting CTOs and architects to carefully evaluate trade-offs between flexibility, efficiency, and TCO for AI workloads, from managing Large Language Models (LLMs) to large-scale inference.
Meta is reportedly planning a significant expansion in the AI hardware sector, with the introduction of a pendant and a roadmap for smart glasses. This move suggests an acceleration towards edge AI devices, potentially capable of processing complex workloads directly on the device. The initiative could redefine user interaction with artificial intelligence, bringing advanced capabilities outside traditional data centers and opening new frontiers for data sovereignty.
Executives from leading memory manufacturers gathered in Taiwan ahead of Computex 2026. This meeting highlights the increasing importance of high-performance memory for the evolution of Large Language Models and on-premise deployment strategies. Decisions regarding VRAM and bandwidth will directly impact TCO and data sovereignty for companies developing local AI stacks.
Eight major PC manufacturers have committed to integrating the Nvidia-MediaTek RTX Spark platform into their upcoming laptops. This move marks a significant step towards the widespread adoption of "AI agent laptops" by fall, shifting artificial intelligence capabilities directly onto client devices and opening new prospects for edge AI processing.
Nvidia announced the RTX Spark Superchip at Computex 2026, a new platform targeting laptops and desktop PCs. Integrating an Arm CPU and a Blackwell GPU with 128GB of unified memory, Nvidia aims to transform Windows into an "agentic AI OS." This solution promises to bring advanced artificial intelligence capabilities directly to local devices, offering new opportunities for on-premise processing and data sovereignty.
MediaTek is strategically shifting its focus towards distributed artificial intelligence, targeting edge devices such as smart glasses, AI-powered PCs, and home servers. This move reflects a vision where AI computation moves beyond traditional cloud data centers, favoring solutions closer to the user to ensure greater control, reduced latency, and data sovereignty. The company aims to capitalize on this significant transition in the computing landscape by developing silicon optimized for local inference.
Intel has announced the launch of its Xeon 6+ processor series, previously known as Clearwater Forest, starting June 1st. Concurrently, the company is introducing the new Intel Ethernet E835 network card. These new hardware components are crucial for on-premise infrastructures, offering significant updates for AI workloads and data sovereignty requirements. Further details on Crescent Island and Diamond Rapids are anticipated.
At Computex, Intel revealed new details about its Crescent Island AI GPU, highlighting a configuration with up to 480 GB of LPDDR5X memory. This capacity aims to address memory shortages, which are crucial for deploying Large Language Models (LLM) on self-hosted infrastructures. The company also provided updates on its Xe3P inference accelerator, strengthening its hardware offering for AI workloads.
Skymizer introduces HTX301, a new hardware accelerator designed to optimize Large Language Model (LLM) inference directly on-premises. The solution focuses on a "decode-first" architecture, aiming to improve efficiency and reduce latency in local deployments. This approach addresses the growing need for companies to maintain control over data and operational costs, offering an alternative to cloud-based solutions for intensive AI/LLM workloads.
AMD has announced the global release of the Radeon RX 9070 GRE, an RDNA 4 architecture-based GPU previously available only in China. Priced at $549 and set to launch on June 2, this graphics card is strategically positioned between the RX 9060 XT and RX 9070 models, offering a new option for users seeking a balance between performance and cost in the GPU segment.
Jensen Huang, Nvidia's CEO, will take the stage at Computex 2026 and GTC Taipei on May 31 for a highly anticipated keynote. This event represents a crucial moment to understand Nvidia's upcoming directions in the artificial intelligence landscape, with significant implications for on-premise deployment strategies, LLM hardware, and the infrastructure decisions faced by CTOs and IT architects.
Running Large Language Models (LLMs) in self-hosted environments presents significant challenges, especially when GPU VRAM is insufficient. A user experienced this issue with a 21GB Gemma 26B model on an AMD RX6600XT GPU, forcing the model to 'spill' into system RAM. This scenario raises crucial questions about the CPU/GPU workload distribution mechanism and the impact of PCIe bus and RAM speed on inference performance, a key consideration for those evaluating on-premise deployments.
A leak reveals details about Nvidia's upcoming N1X and N1 processors. Specifications indicate the adoption of 16-channel DDR5 memory, with bandwidth expected to exceed 500 GB/s. These figures, if confirmed, suggest a significant step forward in processing capabilities, with implications for intensive workloads such as Large Language Models (LLM) and on-premise inference, where memory access speed is crucial for performance.
Ahead of its official Computex launch, specifications for Nvidia's N1/N1X System-on-Chip (SoC) have leaked. The new Arm-based SoC is expected to feature up to 20 cores, with standard configurations including 10- and 12-core variants. These details provide an early look at Nvidia's future processing solutions, potentially relevant for on-premise and edge computing deployments where efficiency and control are paramount.
The emergence of processors like the Snapdragon X Elite marks a turning point for on-device AI, shifting the processing of Large Language Models and other AI functionalities directly to client devices. This evolution offers new opportunities for data sovereignty and reduced latency, laying the groundwork for a more distributed AI architecture less reliant on centralized cloud infrastructures.
Thermal management is a critical challenge in high-density AI hardware on-premise deployments. A user has developed a DIY cooling solution for a DGX Spark cluster, addressing overheating issues caused by the forced proximity of the units. The project, which includes a 3D-printed case and an automatic ventilation system, highlights the ingenuity required to optimize local infrastructure while maintaining cost control and data sovereignty.
Steven Sinofsky, a former Microsoft executive, recently shared a significant memory: the moment Windows first ran on Arm hardware with an Nvidia Tegra chip. This event, dating back to 2010, was an early attempt to explore new architectures for the operating system. This retrospective offers insights into the challenges and opportunities that shaped Windows' evolution and the processor landscape, particularly the rise of Arm in computing.
The recent return of the Orpheus II ISA soundcard, driven by niche demand for DOS and legacy Windows systems, offers a valuable insight. This phenomenon highlights how the need for specific hardware, optimized for well-defined workloads, is equally crucial in the context of Large Language Models. For CTOs and infrastructure architects, choosing on-premise solutions requires careful evaluation of hardware specifications to ensure data sovereignty and optimal TCO.