Modal Labs, a New York-based startup providing compute infrastructure for AI workloads, has chosen London for its new European office. The Marble Arch location can house up to forty people, and the team is expected to be in place by early September, sources say.

Founded in 2021, Modal has carved out a precise niche: instead of chasing the GPU-hungry training of Large Language Models, its platform is optimized for inference – running already-trained models. This focus sets it apart from many peers and addresses a pressing enterprise need: putting AI into production with low latency, predictable costs, and increasingly, guarantees about data residency.

The London move fits a pattern. In recent months, the UK capital has seen offices or expansions from heavyweights like OpenAI and Anthropic, and startups like Cursor and Cohere. The axis is shifting: after the mega-cluster build-out for training, the market is seeking points of presence closer to end users and European regulations. For companies weighing cloud deployments, but also hybrid or on-premise architectures, having a vendor with local staff – and potentially regionally replicable infrastructure – changes the conversation around compliance, latency, and Total Cost of Ownership.

Modal arrives backed by a $355 million funding round closed in May, pushing its valuation to $4.65 billion (merely eight months earlier it was $1.1 billion), led by Redpoint Ventures and General Catalyst. The company, with about 170 employees across New York, San Francisco, and Stockholm, also recently made security headlines when an “unauthorized agent” from OpenAI allegedly breached Modal’s systems and those of Hugging Face.

The choice to double down on London aligns with CEO Erik Bernhardsson’s statement: “Modal has been transatlantic since day one; we’re doubling down with our new London office as part of our commitment to helping European companies achieve scale.” It signals that inference – the point where models deliver real value – can no longer transit transatlantic routes for all sensitive data. For teams managing on-premise AI pipelines or bound by data residency rules, the physical presence of an operator like Modal shifts incentives: more than just faster support, it opens the door to hybrid architectures where data stays nearby and latency isn’t a drag on productivity.