On June 30, 2026, Etched finally lifted the veil on its progress since its founding in 2022: the startup looking to challenge Nvidia’s dominance in AI hardware announced a $5 billion valuation and commitments exceeding $1 billion for complete inference systems. The chip, successfully manufactured by TSMC earlier this year, sits at the heart of what Etched calls “frontier inference clusters” – a bundle that includes the processors, custom-designed racks, and software, built to make inference of large language models (LLMs) faster, cheaper, and more energy-efficient.

The chip that emerged from silence

Behind the announcement lies a story of resilience. Co-founders Gavin Uberti and Robert Wachen dropped out of Harvard to become Thiel fellows and spent 2023 trying to convince investors that the future of AI would require specialized chips, not just Nvidia’s general-purpose GPUs. Every major fund they pitched passed, and the company – the founders recount – was operating month to month, close to running out of cash. Today the scene is radically different: Etched has raised a total of $800 million, including a previously undisclosed $500 million round closed in December at that $5 billion valuation. The investor lineup features VentureTech Alliance, Jane Street, Hudson River Trading, Two Sigma, Ribbit Capital, and a group of heavyweight angels like Andrej Karpathy, Geoffrey Hinton, Fei-Fei Li, Arthur Mensch, and Scott Wu, alongside billionaires Stanley Druckenmiller and Peter Thiel.

The gold rush in custom inference

The interest is no accident. Inference – the process that turns a prompt into a response – is today the most expensive bottleneck for companies serving AI models at scale. Every request burns compute and money, and traditional GPUs, however versatile, are not optimized for this specific workload. Etched is not alone in this race: Cerebras staged the year’s standout IPO, Groq just raised $650 million, while the cloud hyperscalers Amazon, Google, and Microsoft are all building their own in-house AI chips. Even OpenAI unveiled its first custom processor, built by Broadcom. The bet is that silicon designed from scratch for inference can slash operational costs and improve latency, offering an alternative to costly GPU clusters.

What it means for those hosting their own LLMs

The shift toward specialized inference chips is not just for cloud giants. For organizations evaluating on-premise deployment of their LLMs – driven by data sovereignty, controlled latency, or predictable costs – systems like Etched’s could rewrite the Total Cost of Ownership (TCO) calculation. A dedicated inference appliance, with optimized energy consumption and focused performance, can make self-hosting economically viable even for large models, keeping sensitive data within the corporate perimeter. Of course, trade-offs exist: specialized hardware is less flexible than a GPU for mixed workloads or fine-tuning, and requires a mature software ecosystem. AI-RADAR, which closely tracks the evolution of local stacks, offers analytical frameworks at /llm-onpremise to navigate such decisions.

An ecosystem in turmoil

Etched’s rise signals a profound shift in AI infrastructure. After years of unchallenged GPU dominance, the AI hardware space is filling with heterogeneous solutions, each tailored to a segment of the pipeline: training, fine-tuning, inference. For those designing their AI strategy, the availability of inference-specific options could reshape data center architecture and the distribution of workloads between cloud and on-premise. In this ferment, the news of a billion dollars in contracts even before commercial availability is a strong signal: the industry is ready to bet on a future where inference is no longer the Achilles’ heel of advanced models.