Over the past year or so, we’ve been busy collaborating with Netlify to build the underlying platform for Netlify Edge Functions. That product serves over 20 billion requests per month, serving customers from startups to industry leaders like Figma and Riot Games. Netlify details the migration in its own engineering blog post.
The main constraint was latency: Edge Functions execute customer code before a website responds, so time spent preparing a function adds directly to time to first byte (TTFB). Customers use Edge Functions for routing, regional personalization, and authorization. Next.js Edge Middleware runs there too. The platform has a single-digit millisecond budget to find capacity, prepare the execution environment, and start handling a request. Traffic can also change abruptly. Imagine a customer has a launch that can increase volume 200x, and the platform has to cope with that.
Why Netlify Moved
Before the migration, Netlify used a hosted platform that ran functions in V8 isolates. The provider managed execution, load balancing, compute allocation, and provisioning. Additionally, they controlled the runtime, which limited what Netlify could change. Netlify’s customers wanted filesystem access, WASM loading, and native NPM modules, but Netlify couldn’t modify the runtime to support them. The team also wanted control over regional rerouting and resilience, and the ability to optimize infrastructure costs as traffic grew.
In late 2025, Netlify considered (1) keeping the existing platform, (2) building their own, or (3) moving to Unikraft. Deno remained the runtime through the migration, but Netlify now ships its own build inside a microVM and operates the infrastructure around it.
With Unikraft, Netlify reduced their p50 request overhead time by more than 8x, while going from isolates to hardware-isolated VMs.
microVM Startup in 2 Milliseconds
VM boot time was an obvious concern from the get-go. A VM that takes seconds to start is unsuitable for code in the request path. Unikraft microVMs boot directly into a specialized workload, with a minimal system built for that workload.
Benchmark
On Netlify’s fleet, the reported P99 for VM startup is 2ms — selecting a machine, allocating the VM, booting it, and having the customer’s code on disk (excluding execution of the customer’s actual function).
Each customer’s code runs inside a hardware-isolated microVM. Netlify gets that isolation within the latency budget of Edge Functions.
Getting Code onto the Right Machine
The fleet serves code images for sites across Netlify, and those images change constantly. Copying every image to every node would require distributing every update across the fleet; keeping warm instances would consume capacity before requests arrived.
// On-demand image resolution
Any node can serve any site. The image arrives with the first request.
Edge node
requests 1 · origin pulls 1
- 01
Check local cache
every request
- 02
Pull from origin
on a miss only
- 03
Boot the microVM
code already on disk
Local image cache
only what this node has served
- acme-shop
- free
- free
- free
Netlify origin
Every site's image, every version
↑← pulled on a miss only
Resolving images at request time also gives Netlify control over platform updates. Its engineers can select image versions by site, customer, and account tier. This way, they can roll out a change to free-tier sites quickly, and then move enterprise sites on a slower schedule without redeploying the fleet. We co-designed this mechanism with Netlify, and it is now a standard part of the Unikraft platform. For us, this is a deliberate process: We’re very attentive to customers’ feature requests, which then often get incorporated into our product.
Fast startup and on-demand image loading mean we can put any customer’s code on any node, at request time, without pre-provisioning it everywhere. — Netlify Engineering, Platform team
Handling Production Traffic at Large Scale
Netlify owns the load balancing, control plane, and placement policy. Unikraft supplies the microVM execution platform underneath them (for other customers, we sometimes handle the load balancing as well). Fast startup and on-demand image loading give Netlify more freedom to choose where a workload runs, without pre-provisioning each customer’s code on every node.
The trick is (and AI hasn’t significantly moved the needle here) running infra reliably at scale, both from a functional perspective, but also from a performance one. To give you an idea of scale, there may be as many as 400 million “invocations” (read: VM creations) per day. At that scale, any issue that may have a statistically tiny chance of happening is guaranteed to get triggered, so a lot of the work involves battle testing the Unikraft platform at large scale, over and over again.
And that scale may be highly variable: Netlify Edge Functions can experience traffic spikes of 200 times a single customer’s usual volume. Despite this, the current deployment with Unikraft has a measured availability of 99.998%.
Under the Hood: More Platform Optimizations
A handful of smaller changes, made along the way, added up to a meaningful share of the overall gains. virtio-console instead of a serial device. Functions and frameworks constantly write output to the console. Over a serial device, each character write triggers a hypercall (and thus, a VM exit) which degrades performance badly under real workloads. Moving to virtio-console cut that overhead substantially.
- 50% less memory per server. Using EROFS for the images, and splitting user functions out into our ROMs (read-only memory blobs) feature, cut memory consumption on each server in half.
- Just-in-time instance creation. Rather than calling an API to create microVMs, the platform can take that configuration directly in the incoming request: as the request arrives, the platform sniffs it, parses it, and fully instantiates the microVM (note: these requests are generated internally by Netlify, not their customers, so there’s no vector of attack here). Netlify uses this to enrich requests with everything needed to create an instance capable of handling them. Once created, instances stay warm while in use and are auto-killed once idle; a configuration change simply spins up a new service. This removes API calls from the loop entirely. Everything is request-driven, scaling up and down on its own, down to killing instances, templates, and services, and pruning images from the cache; the system is, in effect, self-managing.
- Automatic image pruning. A background system continuously prunes rarely used or stale images from local storage based on a configurable policy, so per-node caches stay small without manual cleanup.
Flexible Deployment
It was clear early on in the development that it was fundamental to Netlify to be able to install our platform on any number of different providers or technologies, both from an immediate need perspective but also in terms of future and long-term flexibility.
As a result, we had to work on ensuring that Unikraft would install on not just a bare metal server, but also on EC2 instances. Think “on VMs, meaning Unikraft microVMs running on top of an EC2 VM”. And not only just install, but also perform just as well, or as close as possible to, metal performance on those VM-based platforms.
Long story short, Unikraft can now run on bare metal, non-metal EC2, nested virtualization EC2, or pretty much anything you can throw at it. Whether you are on AWS or other providers like GCP, metal providers, or other VM-based providers. We’ve also recently added ARM64 support, since these servers tend to provide better bang for your buck, especially at scale.
Owning the Runtime and Other Key Advantages
Runtime ownership has also changed how Netlify ships features. Its engineers maintain a patched build of Deno, backport fixes on their own schedule, and add functionality without waiting for a runtime provider. WASM support, native NPM modules, and filesystem access became possible through that work.
The same architecture also leaves room for other runtimes. Netlify could run Node at the edge, which customers have requested, or support languages beyond JavaScript. The team has also discussed running its own AI gateway as an edge function, potentially on a different runtime. Note that these are just possible uses of the platform, not features announced here.
The key difference here with respect to the previous provider is that Netlify are now:
- Fully in control of the deployment, and so can decide unilaterally (read: without us, Unikraft, in the loop) what to deploy.
- Able to deploy any runtime or any sort of functionality. These are full-fledged VMs based on Dockerfiles, so the sky’s the limit.
- Seeing significantly better unit economics, thanks to scale-to-zero and other Unikraft platform mechanisms.
Faster Log Delivery
Another metric that was key to Netlify in terms of user experience was the speed of log delivery. Here, once again, the switch over to Unikraft resulted in important improvements. On the previous platform, the median delay between a function writing a log line and the customer seeing it was 2.5 seconds.
Now, with Unikraft, it has dropped down to 500ms, which includes Netlify’s enrichment pipeline. This means that customers watching a deployment can see what their code is doing essentially in real-time. As with other stats, these numbers have been measured to hold throughout scaling the deployment, but also when coping with larger, 15x spikes.
This morning, log volume in one region increased 15 times (from about a thousand lines per second to fifteen thousand) and it didn’t change ingestion latency at all. It just kept performing as normal. That’s what we want. — Jake Champion, Netlify
Production results
Netlify’s engineering team shared these figures from its production fleet.
| Metric | Current platform | Previous platform |
|---|---|---|
| P99 VM startup, before runtime init | 2ms | (unknown/black box) |
| P50 per-request overhead | 6ms | ~50ms |
| Availability | 99.998% | - |
| P50 log delivery | 500ms | 2.5s |
Read more
Netlify’s engineering post explains its traffic management and control plane. Read the Netlify customer story for more on the migration.