The surge in AI server demand is creating ripples in the supply chain: orders for power management ICs (PMICs) are spilling over to additional suppliers, signaling bottlenecks. A key signal for anyone planning on-premise deployments.
Shanghai Orient Computing Core Technology, founded by a Tsinghua-schooled chip industry veteran, is developing 3D AI processors to reduce China’s reliance on foreign GPUs. The move comes amid US export restrictions and the race for technological sovereignty.
AI lab Anthropic is exploring custom processors with Samsung as a potential manufacturing partner. While still informal, the move signals a push to diversify beyond Nvidia hardware, with implications for on-premise LLM deployments concerning TCO and data sovereignty.
A developer crafted a CUDA patch for llama.cpp that lets DeepSeek V4 Flash run with a one-million-token context on a single RTX 5090, slashing VRAM requirements from roughly 256 GB to just 31 GB while reaching prefill speeds up to 263 tokens per second. Validated through needle-in-haystack tests, the achievement marks a turning point for on-premise deployment of ultra-long-context models.
Samsung has reached over 70% production yield for its next-generation HBM4E memory, raising the stakes against SK Hynix and Micron. The milestone indicates manufacturing maturity that could expand bandwidth availability for AI accelerators, a critical resource for LLM inference and training. For teams evaluating on-premise infrastructure, a healthier supply chain directly affects hardware TCO and deployment constraints.
The Japanese chipmaker is refocusing its semiconductor investments on two booming sectors: AI server processing and electric mobility. The move underscores the growing convergence of high-performance computing and vehicle electrification.
Anthropic has entered talks with Samsung Electronics to explore manufacturing a custom AI chip. The project is at an early stage, with no decisions yet on purpose, power, or server integration. The move fits a broader industry shift toward vertical integration among leading AI players, potentially impacting on-premise LLM deployments: better efficiency is possible, but questions remain about whether such hardware will be available to enterprise customers.
Anthropic is reportedly discussing a custom chip with Samsung for its LLMs, shortly after OpenAI’s similar move with Broadcom. The trend toward proprietary silicon could reshape TCO and data sovereignty for on-premise AI deployments, while adding integration complexity.
A record-breaking investment reshapes the memory supply chain: the South Korean giant bets on NAND and DRAM to sustain AI infrastructure demand. Implications for on-prem cluster operators, spanning HBM, TCO, and bottleneck management.
An aggressively priced CPU air cooler with high noise levels. For those building local inference rigs or on-premise workstations, the trade-off between cost and quiet operation becomes critical.
Intel posted initial GCC compiler patches for AI Compute Extensions (ACE), the new instruction set co-developed with AMD to accelerate AI workloads on x86. The cross-vendor successor to Intel's AMX, ACE targets matrix multiplication for machine learning. The move brings native on-premise inference acceleration one step closer without relying on dedicated GPUs.
Official product pages for the Core Ultra 270K Plus and 250K Plus show recommended prices up to $50 higher. The move signals cost pressures and affects builders of workstations for local LLM inference.
The Norwegian deep-tech company got funding led by Nysnø Climate Investments, Sandwater, and Emerald to bring to market ever smaller and more efficient motors, a signal for robotics and on-device AI.
Raja Koduri's startup Oxmiq Labs has closed a $35M Series A to scale OxCore, a licensable GPU architecture that lets chipmakers build custom AI silicon without a full multi-year design program. Total raised now stands at $60M.
The Korean company announced a new NAND fab in Cheongju, targeting first-half 2029 production. The investment reflects how AI is driving demand not only for high-bandwidth memory (HBM) but also for fast, dense storage to handle growing datasets and on-premise workloads.
Chairman Yu Yingtao's resignation marks a new chapter for Chinese ICT vendor H3C, which is accelerating its push into AI servers — a sign that demand for on-premise LLM infrastructure is reshaping vendor strategies, amid sovereignty and supply chain concerns.
The Norwegian deeptech company secures a round led by Nysnø, Sandwater and Emerald Technology Ventures to expand production of frameless motors based on FiberPrinting technology, targeting robotics, aerospace and medical device markets.
A user managed to fit two RTX 3090 GPUs inside an open-frame Thermaltake Core P3 case by 3D-printing a bracket to tilt the radiator. Beyond the striking visuals, the build can locally run models like Qwen 27B. For those evaluating on-premise deployment, it’s a reminder that powerful self-hosted LLM setups are within reach — with a bit of physical tinkering and 48 GB of combined VRAM to handle mid-size model inference.
The Korean foundry advances its 2nm roadmap as demand for AI chips grows. The shift promises gate-all-around transistors, better energy efficiency and density, crucial for next-gen silicon dedicated to training and inference, with direct implications for on-premise computing.
The Taiwanese manufacturer is diversifying into server racks, leveraging the North American powersports cycle and Vietnam’s EV transition. Its entry into a market crucial for AI infrastructure could impact supply chains and hardware costs for on-premise deployments.