$ ldd --version # how does your code actually run?

Runtimes

Two questions, one page. First, how does a single program run, from source code down to the CPU? Then, how does it do many things at once, with threads, async, futures and the runtimes that schedule them? Read straight through or jump around: each section builds on the one before it.

// execution-models

Execution Models

Every language needs a way to turn source code into running instructions. That "way" is the execution model, and it fundamentally determines startup time, peak performance, memory usage, and portability. The six families below cover essentially every language on the PL Timeline.

Native (AOT)

Compiled ahead-of-time to machine code that the CPU runs directly. No layer in between: fastest, smallest, but tied to one platform.

languages

C, C++, Rust, Go, Zig, Fortran, Swift, OCaml

JIT (Just-In-Time)

Runs from bytecode, then recompiles hot paths into optimised native code while the program is running. Slow to warm up, near-native once hot.

languages

Java (HotSpot), C# (.NET), JavaScript (V8), Julia, LuaJIT, PyPy

Bytecode VM

Compiled to portable bytecode that a virtual machine executes, providing GC and a sandbox. Write once, run anywhere the VM exists.

languages

Java (JVM), C# (CLR), Erlang (BEAM), Python (CPython), Ruby (YARV)

Interpreter

Reads and executes the source statement by statement, no compile step. Instant startup and a live REPL, but slowest for heavy CPU work.

languages

Python, Ruby, PHP, Perl, sh / Bash, R, SQL

WebAssembly

Compiled to a portable binary that runs inside a sandbox (browser or Wasmtime). Near-native speed; I/O only through explicit host imports.

languages

Rust, C, C++, Go, Zig, AssemblyScript

Transpiler

Translated to another high-level language (usually JS or C) rather than machine code, then that language's toolchain takes over.

languages

TypeScript → JS, Elm → JS, PureScript → JS, Cython → C

// diff --side-by-side

Side-by-Side Comparison

A quick overview of the key trade-offs between execution models. Every choice is a compromise: native compilation gives you peak speed but locks you to a target; interpreters start instantly but pay for it every instruction.

ModelStartupPeak perf.MemoryPortabilityMemory mgmt
Native (AOT)FastHighestLowestRecompile per targetManual / ownership
JITSlow (warm-up)Near-nativeHighVM per platformGC (tracing)
Bytecode VMMediumGoodMediumVM per platformGC (tracing)
InterpreterInstantLowestMediumInterpreter per platformGC (ref-counting / tracing)
WebAssemblyFastNear-nativeLowUniversal (sandboxed)Depends on source lang
TranspilerDepends on targetDepends on targetDepends on targetSame as targetSame as target

// from source to execution

The Big Picture

To tie it together: all execution models are variations on the same pipeline. The key question is when compilation happens and how many layers sit between your source code and the CPU.

Source code.rs / .c / .py / .jsAOT compilergcc, rustc, goMachine codeELF / Mach-O / PEBytecode + JITJVM, V8, CLROptimised codeat runtimeInterpreterCPython, MRI, BashExecuted livestatement by stmtCPUhardware executionx86 / ARM / RISC-V
AOT-compiled languages (Rust, C, Go) take the top path: everything is resolved before execution. JIT languages (Java, JS) defer optimisation to runtime. Interpreters (Python, Bash) skip compilation entirely.

So far, one program running start to finish. But servers and apps must juggle thousands of things at once, and most of the time each one is just waiting, on a socket, a disk, or a timer. The rest of this page is about how a runtime manages all that waiting: threads, async, futures, and the schedulers that drive them.

// std::thread vs async fn

Synchronous vs Asynchronous (Rust)

Rust is unusual in giving you both models as first-class citizens with zero runtime by default. Synchronous code uses real OS threads and blocking calls: simple and perfect for CPU work. Asynchronous code uses async/.await plus an executor (Tokio, smol…) that polls tasks, cooperatively multiplexing millions of them onto as few as one thread: the right tool when you're I/O-bound. Threads are not inherent to async; a multi-thread executor is just an added optimization. Neither model is "better"; they solve different problems.

Synchronous (threads)Asynchronous (async / await)
Concurrency unitOS thread (1:1 with a kernel thread)Future: a poll-based state machine driven by an executor
Spawned withstd::thread::spawntokio::spawn / smol::spawn
Threads usedOne OS thread per task1..N, optional: a single thread is enough
Cost per unit~1–8 MB stack + kernel bookkeeping~a few hundred bytes, no own stack
Practical countThousandsMillions
SchedulingPre-emptive, by the OS kernelCooperative: the executor polls, tasks yield at .await
Blocking a callFine, only that one thread waitsDangerous: stalls the whole executor (use spawn_blocking)
I/OBlocking syscalls (read / write)Non-blocking + epoll / kqueue / io_uring
CPU-bound workIdeal (threads, rayon)Poor: cooperative tasks starve; offload with spawn_blocking or rayon
Many connectionsLimited by thread countExcellent: the reason async exists
CancellationHard: no safe way to kill a threadEasy: just drop the future
ComplexitySimple, direct, no function colouringasync colouring, Send + 'static, Pin, lifetimes
Ecosystemstd, rayon, crossbeamtokio, async-std, smol, futures

Same task, both ways

Spawn 4 workers, each computes i * i, then collect the results

Synchronous : std::thread

use std::thread;

fn main() {
    let mut handles = Vec::new();
    for i in 0..4 {
        // each worker gets its own OS thread (~MBs of stack)
        handles.push(thread::spawn(move || i * i));
    }
    // join blocks the main thread until each finishes
    let results: Vec<i32> = handles
        .into_iter()
        .map(|h| h.join().unwrap())
        .collect();
    println!("{:?}", results); // [0, 1, 4, 9]
}

Asynchronous : tokio::spawn

#[tokio::main]                       // sets up the async runtime
async fn main() {
    let mut handles = Vec::new();
    for i in 0..4 {
        // each worker is a lightweight task (~hundreds of bytes)
        handles.push(tokio::spawn(async move { i * i }));
    }
    // .await yields instead of blocking the thread
    let mut results = Vec::new();
    for h in handles {
        results.push(h.await.unwrap());
    }
    println!("{:?}", results); // [0, 1, 4, 9]
}

What the sync version does

  • Asks the OS for 4 real threads, each with its own multi-MB stack. The kernel schedules them pre-emptively across cores.
  • join() blocks the main thread until each worker returns. While it waits, that thread does nothing else.
  • For CPU work this is exactly right: 4 threads can use 4 cores in true parallel. For 10,000 workers it collapses: thread creation and context switches dominate.

What the async version does

  • Creates 4 tiny tasks (state machines of a few hundred bytes) that an executor polls. This can run on a single thread; the multi-thread pool is an optional optimization layered on top, not what makes it async. No new OS threads per worker.
  • .await yields instead of blocking: on a not-ready result the task returns control to the executor, which polls another. A stray blocking call here would freeze every task sharing that thread.
  • For pure i * i this adds overhead for no gain. The payoff appears with thousands of I/O-bound tasks (sockets, DB queries) idling at once.

Same output, different machine behaviour. Both print [0, 1, 4, 9], but the sync version spends its cost on threads and stacks (great for CPU), while the async version spends it on a scheduler and state machines (great for waiting on I/O). The rule of thumb: threads when you're compute-bound, async when you're I/O-bound and highly concurrent.

// Poll::Pending vs Poll::Ready

What async really is (the Future trait)

We keep saying "the executor polls a task", but what is the task? A future is a value that stands for a computation that has not finished yet: a job you can hold in a variable, pass around, and repeatedly ask "are you done?". It is the description of the work, not a thread busy doing it. In Rust every async fn and async block compiles into exactly such a value: a type that implements the Future trait.

That trait is tiny. It has one method, poll, which returns Poll::Ready(value) when the work is finished or Poll::Pending when it is not. And futures are lazy: constructing one runs none of its code. Until something calls poll, a future is inert data, which is why an async fn you never await simply does nothing.

"An async block is a value", made concrete

Control flow makes the point sharply. Inside a plain block, return (and ?) exits the whole function. Inside an async block they do not: they resolve the Future the block produces. So the two snippets below look almost identical but behave completely differently.

Plain block : returns from example()

// A plain block: `return` exits the whole function.
fn example() -> i32 {
    let x = {
        return 5; // returns from example() immediately
    };
    // unreachable: the block diverges, so x has the never type `!`
}

Async block : returns a Future

// An async block: `return` resolves the block's Future.
async fn example() -> i32 {
    let x = async {
        return 5; // sets this Future's output to 5, does NOT exit example()
    };

    // x is a Future<Output = i32>. It has run no code yet (lazy).
    x.await // driving it here is what finally produces 5
}

On the left, return 5 ends example() on the spot, and x never exists (it takes the never type !). On the right, x is a Future<Output = i32> that has executed nothing; only x.await drives it and yields the 5. The future is a value you hold, and awaiting is what runs it.

A mental model: the future is a recipe, the runtime is the cook

Writing a recipe cooks nothing; it is just a plan with steps. A cook has to pick it up and follow it. Some steps say "put it in the oven and wait": instead of standing there, the cook sets a timer and starts another dish, returning only when it dings. Rust async is exactly this, term for term.

Recipe card
the future (a lazy plan)
Following the next step
poll()
"in the oven, wait"
.await returns Pending
Setting a kitchen timer
registering a Waker
The timer dings
wake(): re-poll
Dish plated
Poll::Ready(value)
Cooking other dishes meanwhile
one thread, many tasks

A future is a resumable state machine

The magic is what the compiler does with your async fn. It rewrites the body into an enum with one variant per pause point, and every .await is a pause point. Each call to poll tries to advance from the current state; if it reaches an .await that is still Pending, it records "I stopped here" inside the enum and returns. The next poll resumes at exactly that spot.

Because that resume-point lives in a small struct rather than on a thread's call stack, futures are stackless: a task is just a few hundred bytes of enum, not a multi-megabyte stack. That is the whole reason one thread can juggle millions of them.

Your async fnStates of the futureasync fn handle(sock) -> usize { let n = read(sock).await; store(&buf[..n]).await; n}pause #1pause #2Start.await #1 · Reading.await #2 · StoringDone · Ready(n)
How the code becomes states. The function entry is Start; each .await line (pink) becomes a state the future can pause in, Reading then Storing; the final expression is Done, carrying Ready(n).
Executorcalls future.poll(cx) repeatedlypoll()Pending / Ready(n)The future = enum Handle (a few hundred bytes). Its current variant marks where it is paused.Startnot polled yet.await #1reading socket.await #2storing dataDoneReady(n)pollread readystore readypoll → Pendingpoll → Pending
Now with the future itself in view. The whole enum Handle is the future: one value the executor drives by calling poll(). Its current variant is the state, and each .await is a state it can rest in: while the resource is not ready, poll returns Pending (the pink loops); when it becomes ready the machine advances one state, until Done returns Ready(n).

And that machine is not a metaphor: it is quite literally the enum the compiler generates for you.

// You write this:
async fn handle(mut sock: TcpStream) -> usize {
    let n = sock.read(&mut buf).await;   // pause point #1
    store(&buf[..n]).await;              // pause point #2
    n                                    // final value
}

// The compiler rewrites it into roughly this state machine.
// Each .await becomes a state the future can be "parked" in.
enum Handle {
    Start   { sock: TcpStream },              // not polled yet
    Reading { fut: ReadFuture },              // parked at .await #1
    Storing { fut: StoreFuture, n: usize },   // parked at .await #2
    Done,
}

// poll() looks at the current state, tries to drive the inner
// future, and either advances to the next state, returns
// Poll::Pending (saving where it stopped), or Poll::Ready(n).
One thread, over timetime →Future AFuture BThreadA's I/O ready → wake()B's I/O ready → wake()createdrunningPending · parked, waker setrunningReady(T) ✓ donecreated, waiting its turnrunningPending · parked, waker setrunningReady ✓poll(A)→ Pendingpoll(B)→ Pendingidlefree for other workpoll(A)→ Ready ✓poll(B)→ Ready ✓running (being polled)Pending (parked, waker set)Ready (done)created / queuedOnly one future runs at a time; the dotted line shows which future the thread is polling right now.
The same run, now following the two futures themselves. Each future moves through a lifecycle: created, running while the thread polls it, Pending and parked with a waker while its I/O is not ready, then running again once woken, then Ready. The thread lane below shows it only ever drives one future at a time: when poll(A) returns Pending it parks A and polls B, staying busy instead of blocking.

At the bottom of every such chain sits a leaf future, one that actually talks to the outside world (a socket, a timer) instead of awaiting another future. This is where Pending and the Waker come from. Here is one written by hand so the contract is visible:

use std::future::Future;
use std::pin::Pin;
use std::task::{Context, Poll};
use std::time::Instant;

// A hand-written future: an async fn compiles to something like this.
struct Delay { deadline: Instant }

impl Future for Delay {
    type Output = ();

    // The executor calls poll(). We either finish or ask to be polled later.
    fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<()> {
        if Instant::now() >= self.deadline {
            Poll::Ready(())                 // done: hand back the value
        } else {
            // Not ready. Hand the executor a Waker so it can re-poll us
            // once the timer fires, instead of busy-looping.
            register_timer(self.deadline, cx.waker().clone());
            Poll::Pending
        }
    }
}

Notice the Waker. When a leaf returns Pending it first stashes the waker it was handed in the Context. The waker is the future's way of saying "call poll on me again when there is news". The runtime keeps it, and once the resource is ready it invokes wake() to put the task back on a run queue. No busy-waiting, and no thread parked per future.

Async is only worth it over an async interface

Polling only helps if the thing you wait on can answer "not ready yet, I'll signal you". A blocking syscall like read / write gives no such handle: it just parks the calling thread until data arrives. To wait on many blocking calls at once your only option is many threads, and the kernel schedules them. That is concurrency, but not necessarily parallelism: those threads may all share one core and simply take turns.

An async interface flips this around, and there are two families. Readiness-based ones (epoll, kqueue) let one thread register thousands of descriptors and ask the kernel which are ready, then you run the non-blocking read yourself. Completion-based ones (io_uring, IOCP) go further: you submit the whole operation and the kernel hands back the finished result. Either way, that signal is what the executor's reactor uses to fire the waker and re-poll a task. So epoll is not "the" async interface, just the most common one on Linux; without some async interface underneath, "async" buys you nothing over threads.

InterfacePlatformModelHow you reach it in Rust
epollLinuxReadinessmio, auto-selected by Tokio's enable_io() (so, by #[tokio::main])
kqueuemacOS / BSDReadinessmio, Tokio's default on those platforms (same enable_io())
select / pollPOSIX (legacy)Readinesslibc directly; largely superseded by epoll / kqueue
IOCPWindowsCompletionmio, Tokio's default on Windows (again via enable_io())
io_uringLinux (modern)CompletionOpt-in, a different runtime: tokio-uring, glommio, or compio
Readiness: the kernel says "fd 7 is ready", then you run the non-blocking read yourself. Completion: you submit the operation and the kernel hands back the finished result. On stock Tokio you never pick one, enable_io() selects the platform default (epoll / kqueue / IOCP) through mio; io_uring means choosing a different runtime.

Here is that interface in action. Follow one future that is waiting to read from a TCP socket, and watch how the reactor uses epoll to keep the thread free instead of blocking it inside a read:

⑥ waker.wake() → re-queue the task, then poll() againExecutorschedules + polls tasksFuture.poll()returns Pending / Ready(T)Reactormap: fd → Waker① poll()eventually Ready(T)② PendingKernel · epollthe async interfaceepoll_ctl(fd, EPOLLIN) · watch fdepoll_wait() · block, return ready fdsTCP socket · fd 7peer sends bytes→ becomes readable③ register interest⑤ ready: [fd 7]④
The poll loop, over a real async interface. One future is waiting to read from socket fd 7; the numbered steps show how the reactor drives epoll so that no thread ever sits blocked inside a read.
  1. ① poll(). The executor calls poll on the task's future to drive it forward.
  2. ② Pending. No bytes yet, so the future returns Pending. On the way out it hands the reactor its file descriptor (fd 7) and a Waker, then steps aside so the thread is free.
  3. ③ Register interest. This is where the async interface is used: the reactor calls epoll_ctl(ADD, fd 7, EPOLLIN), telling the kernel "signal me when fd 7 is readable". One reactor registers thousands of descriptors this way.
  4. ④ Readiness. Bytes arrive from the peer, so the kernel flags fd 7 as readable.
  5. ⑤ epoll_wait returns. The reactor's single epoll_wait() call, which was blocked waiting on all registered fds at once, wakes up and returns the ready list containing fd 7.
  6. ⑥ Wake and re-poll. The reactor looks up the Waker registered for fd 7 and calls wake(), which re-queues the task. The executor polls it again, and this time poll reads the bytes and returns Ready(T).

With a blocking read() there is no step ③ or ⑤: the thread simply sleeps inside the syscall until data arrives, so serving N connections needs N threads. The async interface is exactly what lets a single thread wait on all of them and only touch the ones that are ready.

One question the diagram leaves open: who creates that epoll instance, and when? You never call epoll_create yourself. It is created once, when the runtime starts up, by its I/O driver (the reactor). In Rust that is exactly what enable_io(), implied by #[tokio::main], switches on; the one instance is then shared by every task, and individual descriptors are registered into it lazily the first time each resource is polled. This is not a Rust quirk: Node's libuv, Python's asyncio selector and Go's netpoller each create theirs the same way, at event-loop start-up. The Tokio section below shows the exact call; first, how the same idea looks in other languages.

Futures are not a Rust idea

The concept goes back to the 1970s (Baker and Hewitt named "futures", Friedman and Wise "promises") and almost every mainstream language ships one today: JavaScript's Promise, Python's coroutine, C#'s Task, Java's CompletableFuture. What changes from language to language is two things: whether the async value is lazy (does nothing until driven) or eager (already running), and whether the runtime is built into the language or something you choose. Rust is the unusual one on both counts.

Only two independent questions separate one language's async value from another's, so every language lands in one of four boxes:

Runtime built into the language
Runtime you add yourself
Lazy
value
Python asyncioKotlin

sits inert until the built-in loop drives it

RustC++20

sits inert until the executor you picked polls it

Eager
value
JavaScriptC#Go

starts running the instant it is created

uncommon
Vertically: does the value start on its own (eager) or wait to be driven (lazy)? Horizontally: does the language ship the loop, or do you bolt one on? These are orthogonal, which is why "how is a Rust future different from a Python one?" is a bit of a trick question: they share the same lazy row. As values they behave identically; they differ only on the horizontal axis, Python bundles asyncio while Rust makes you pick Tokio or smol.

That is the whole "zoom out": once two async values are both lazy, they run identically underneath. Each is created inert, each hands out a one-shot resume handle (Rust's Waker, Python's done-callback, Kotlin's Continuation, C++20's coroutine_handle) that means "call me to resume this", and each is driven by a loop that sleeps on the OS async interface rather than busy-waiting. The eager designs differ in one respect only: the value starts the moment it is created. The table below turns every one of these traits into a column, and the Lazy or eager row draws it as a small timeline so you can see the difference at a glance.

But if there is a loop, isn't it just spinning? (why every "Busy-wait?" row says No)

The word "loop" misleads. A busy loop would be loop { if ready() { break } }: one core pinned at 100%, asking "ready yet?" millions of times a second and never letting the work start any sooner. An executor never does that. Its loop has two phases, and the second one sleeps:

  1. 1. Drain the ready tasks. Poll each task a waker has flagged as ready, running it up to its next .await. This phase is real CPU work and only touches tasks that can make progress.
  2. 2. When the queue is empty, block. The executor makes one call, epoll_wait() (or it parks the thread), and the kernel puts the thread to sleep at ~0% CPU until a file descriptor is ready or the nearest timer expires. No spinning: the thread is genuinely idle.
Busy-wait (spin loop)core pinned at 100%; work never starts soonerready? ready? ready? ready? ready? ready? ready? ready? ready?CPU 100% the entire wait, all wastedPark + waker (what runtimes actually do)thread sleeps in epoll_wait until the OS wakes itasleep in epoll_wait · ~0% CPUwake()poll once → Ready

The waker is what removes the need to poll while waiting: a task hands the reactor a "call me when my resource is ready" callback and steps aside, so nothing touches it until the reactor fires that callback and re-queues it. A timer works the same way: sleep(2s) does not spin, it registers a deadline in a timer wheel, and the executor passes the nearest deadline as the timeout to epoll_wait, so the thread wakes exactly once, when the timer is due. That is why the table below can answer Busy-wait? No for every language: the mechanism is always park-and-be-woken, never spin.

Comparison table: futures across languages

This is the per-language table: what the async value is in each language and how it behaves, lazy or eager, its resume handle, whether a runtime is built in, and the syntax. It stops at the language boundary. For how a concrete runtime actually schedules those futures (Tokio vs smol vs glommio vs an event loop), jump to Comparison table: concurrency runtimes.

Pick languages to compare their model and syntax side by side. The first snippet runs the same task everywhere: an async double(x) returning x * 2, called on 10 and 20 concurrently and summed to 60. The second shows the async-block structure from above (a return inside an inner async value) in each language, so you can see how the same idea is spelled and whether it is lazy or eager.

Rust JavaScript Python Go
Descriptionasync fn compiles to a lazy Future; you choose the runtime (here Tokio).Calling double() returns an eager Promise; the event loop is built in.async def builds a lazy coroutine; asyncio drives it.No async keyword; a goroutine plus channel play the future's role.
AbstractionFuture (poll-based state machine)callbacks -> Promise -> async fncoroutine (asyncio)goroutine (no async syntax)
Lazy or eagercreatedidlepollruns Lazy: nothing runs until polledcreatedruns nowrunning Eager: a Promise starts immediatelycreatedidlepollruns Lazy: runs when awaited or scheduledcreatedruns nowrunning Eager: starts on go f()
Resume handleWaker: wake() re-queues the taskmicrotask / job callbackdone-callback + call_soonruntime parks / unparks the goroutine (gopark / goready)
Runtime built inNo: bring your own executor (Tokio, smol)Yes: event loop baked into the engine / hostStdlib loop, swappable (uvloop)Yes: scheduler is part of the runtime
Scheduling / I-OExecutor polls; reactor on epoll / kqueue / io_uringMacrotask + microtask queues; libuv / host pollerselector / epoll; GIL caps parallelismnetpoller does non-blocking I/O, M:N under the hood
Busy-wait?No: parks in epoll_waitNo: host poller sleepsNo: blocks in selector.selectNo: netpoller blocks in epoll
Syntaxasync fn / .awaitasync fn / awaitasync def / awaitgo f() (blocking-looking code)

Same task, each language

An async double(x) returning x * 2, run on 10 and 20 concurrently, then summed to 60

Rust

use tokio::join;

// async fn returns a Future: lazy, does nothing until driven
async fn double(x: i32) -> i32 {
    x * 2
}

#[tokio::main]
async fn main() {
    // join! polls both futures concurrently, then awaits them
    let (a, b) = join!(double(10), double(20));
    println!("{}", a + b); // 60
}

JavaScript

// calling double() starts the Promise immediately (eager)
async function double(x) {
  return x * 2;
}

// Promise.all awaits both concurrently
const [a, b] = await Promise.all([double(10), double(20)]);
console.log(a + b); // 60

Python

import asyncio

# calling double() builds a coroutine: lazy, awaited later
async def double(x: int) -> int:
    return x * 2

async def main() -> None:
    # gather schedules and awaits both concurrently
    a, b = await asyncio.gather(double(10), double(20))
    print(a + b)  # 60

asyncio.run(main())

Go

// no async/await: a goroutine + channel stand in for a future
func double(x int, out chan<- int) {
    out <- x * 2 // starts running immediately (eager)
}

func main() {
    out := make(chan int, 2)
    go double(10, out)
    go double(20, out)
    fmt.Println(<-out + <-out) // 60
}

Return inside an async value

The inner async unit is its own value: return (or a channel send in Go) resolves it, not the outer function. You only get its result by awaiting / receiving. Watch which languages are lazy vs eager.

Rust

async fn example() -> i32 {
    let x = async {
        return 5; // resolves THIS Future, not example()
    };

    // x is a Future<Output = i32> that has run nothing yet (lazy)
    x.await // awaiting drives it and yields 5
}

JavaScript

async function example() {
  const x = (async () => {
    return 5; // resolves THIS Promise, not example()
  })();

  // x is a Promise, but note: it already started running (eager)
  return await x; // await yields 5
}

Python

async def example() -> int:
    async def inner() -> int:
        return 5  # resolves THIS coroutine, not example()

    x = inner()      # a coroutine object; runs nothing yet (lazy)
    return await x   # awaiting drives it and yields 5

Go

func example() int {
    ch := make(chan int, 1)
    go func() {
        ch <- 5
        return // returns from the goroutine, not example()
    }()

    // the channel receive is the "await": it produces 5
    return <-ch
}

// tokio::spawn(async move { ... })

Async & Concurrency Runtimes

A future does nothing on its own, and every language just met has the same problem. Something has to poll those tasks: a concurrency runtime that schedules async work, manages I/O multiplexing, and distributes it across cores. Some languages bake one in (Go, Erlang); others let you choose (Rust, Python). The differences matter when you're building servers, embedded systems, or anything that waits on I/O.

Rust

Stackless coroutines via async/await. The compiler generates state machines; a user-chosen executor polls them.

  • TokioGeneral-purpose multi-threaded, work-stealing scheduler. The default choice for servers.
  • smolLightweight, single-dependency executor. Good for CLIs and small services.
  • embassyAsync executor for embedded / bare-metal (no_std). Runs on microcontrollers.
  • glommioThread-per-core, io_uring-based. Optimised for high-throughput storage and networking.
  • compioCompletion-based (io_uring / IOCP). Cross-platform alternative to glommio.

JavaScript / TypeScript

Single-threaded event loop with async/await. I/O callbacks are queued; the microtask queue runs Promises.

  • V8 (Chrome, Node.js, Deno, Bun)JIT-compiled engine with libuv (Node) or native event loop (Deno/Bun).
  • SpiderMonkey (Firefox)Mozilla's engine with baseline + IonMonkey JIT tiers.
  • JavaScriptCore (Safari, Bun)Apple's engine, also used by Bun for fast startup.

Python

asyncio event loop with async/await. The GIL limits true parallelism in CPython; use multiprocessing or a native extension for CPU work.

  • asyncio (stdlib)Default event loop. selector-based on Unix, IOCP on Windows.
  • uvloopDrop-in replacement built on libuv. 2-4x faster than the default loop.
  • TrioStructured-concurrency-first library: nurseries enforce task lifetime rules.
  • AnyIOCompatibility layer that works on top of asyncio or Trio.

Go

M:N scheduling with goroutines (lightweight green threads) and channels. No async/await syntax needed: every function is implicitly non-blocking.

  • Go runtimeBuilt-in scheduler multiplexes goroutines across OS threads. netpoller handles I/O.

Java / Kotlin (JVM)

Platform threads mapped 1:1 to OS threads, plus (since Java 21) virtual threads via Project Loom for M:N scheduling.

  • Platform threadsClassic OS threads, heavyweight (~1MB stack each).
  • Virtual threads (Loom)JVM-managed lightweight threads, millions per process. Blocking calls auto-yield.
  • Kotlin coroutinesStructured concurrency with suspend/resume, dispatched to thread pools.

C#

Task-based asynchronous pattern (TAP) with async/await. The runtime schedules continuations on the thread pool or a synchronisation context.

  • .NET ThreadPool + Task schedulerWork-stealing thread pool. Task.Run dispatches CPU work; I/O operations use IOCP.

C++ (C++20 coroutines)

co_await / co_return turn a function into a lazy coroutine. The language ships no scheduler, so a library provides the executor and I/O source.

  • Asio / io_contextThe de-facto async I/O library; run() drives an epoll / IOCP loop that resumes coroutines.
  • stdexec (senders/receivers)The proposed std::execution model: composable schedulers and thread pools.
  • cppcoro / libunifexCoroutine task types, when_all, async generators; libunifex adds sender-based scheduling.

Erlang / Elixir

Actor model on the BEAM VM: each process is an isolated lightweight unit with its own heap, communicating solely via message passing.

  • BEAM VMPreemptive scheduler with per-process reduction budgets. Millions of processes, no shared state.
  • OTPFramework of supervisors, gen_servers and behaviours. Fault tolerance by design.

How the same async block gets scheduled

The cards above list which runtimes exist per language; this shows how they actually run your tasks. Take one identical async task and hand it to different runtimes: almost every one falls into one of four scheduling shapes. That shape, not the syntax, is what decides whether a task can hop threads, whether it can be interrupted, and whether it must be Send.

One thread, one queue

A single thread runs a cooperative loop: take the next ready task, run it up to its next await, repeat. No data races and almost no overhead, but one blocking task stalls all the others and it only uses one core.

JS engines · Python asyncio · Tokio current-thread · embassy

Work-stealing pool (M:N)

A few worker threads (about one per core) each own a run queue. An idle worker steals half of a busy sibling's tasks, so a task can resume on a different thread than it started, which is exactly why tasks must be Send.

Tokio multi-thread · .NET ThreadPool · smol

Thread-per-core (share-nothing)

One thread pinned per core, and a task never leaves the core it started on. No locks, no stealing and no Send bound, usually paired with io_uring for predictable tail latency.

glommio · compio · Seastar

Preemptive green threads

No async keyword at all: the runtime multiplexes lightweight threads across cores and can pause one mid-run at a safepoint, so a single long computation cannot starve everything else.

Go goroutines · Erlang / Elixir BEAM · JVM virtual threads

Scheduling shapeTask can move threads?Preemptive?Send bound?Example runtimes
One thread, one queueNo: only one thread existsNo: cooperative, yields at awaitNot neededJS engines, Python asyncio, Tokio current-thread
Work-stealing pool (M:N)Yes: stolen by an idle workerNo: cooperativeYes: tasks cross threadsTokio multi-thread, .NET ThreadPool, smol
Thread-per-coreNo: pinned to its coreNo: cooperativeNo: Rc / RefCell are fineglommio, compio, Seastar
Preemptive green threadsYes: scheduler moves themYes: paused at safepointsn/a: no async keywordGo, Erlang / Elixir BEAM, JVM virtual threads
The first three shapes are cooperative: a task only yields when it hits an await, so a tight CPU loop with no await can hog its thread. The last is preemptive: the runtime can pause a task even mid-computation, which is why Go and Erlang stay responsive under CPU-heavy load without any await points. The Tokio section below zooms into the work-stealing shape in detail.

// who calls poll(), and on which thread?

Inside a runtime : Tokio's scheduler

A runtime is the thing that actually polls your futures. Tokio ships two: the current-thread scheduler, which runs everything on a single thread (proof that the thread pool is optional), and the default multi-thread scheduler, which starts one worker thread per core. The interesting question is how the multi-thread version keeps every core busy without a central lock everyone fights over.

First, what is a "thread pool"?

A thread pool is a small, fixed set of OS threads created once at start-up, usually one per CPU core, that sit idle waiting for work. Rather than spawn a brand-new thread per task (expensive, and unbounded if tasks keep arriving), you drop tasks into a queue and the pool's threads pull them off, run them, and come back for the next one. So the same handful of threads are reused for millions of tasks. Crucially, the pool exists only to spread ready tasks across cores for real parallelism; it is an optimisation, not part of what makes async work. Those exact same futures run correctly on the single-threaded current-thread scheduler, just capped at one core. That is why talking about a thread pool when contrasting sync and async is a little misleading: async is about polling and cooperative scheduling, and it works on one thread; the pool is a separate, optional choice about how many cores to use.

Per-worker local queue

Each worker owns a small fixed-size (256) run queue of ready tasks. It pushes and pops from its own end with almost no synchronisation, so the common path is contention-free.

Global injection queue + LIFO slot

Tasks spawned from outside a worker, or overflow from a full local queue, land in a shared injection queue. A per-worker LIFO slot holds the just-woken task so it runs next, keeping message-passing latency low.

Work stealing

A worker whose queue runs dry first checks the injection queue, then picks a random sibling and steals about half of its tasks. Idle cores pull work toward themselves instead of a dispatcher pushing it.

Global injection queueremote spawns + local-queue overflowWorker 1 (core 0)one OS thread, in a looplocal run queue (256)tasktasktaskLIFO slot (just-woken)poll(task)thread runs the taskpopReady ✓ done · else Pending ↓then loops for the next taskWorker 2 (core 1)one OS thread, in a looplocal run queue (256)tasktasktaskLIFO slot (just-woken)poll(task)thread runs the taskpopReady ✓ done · else Pending ↓then loops for the next tasksteal ~half when idleReactor / driver: mio → epoll / kqueue / IOCPone epoll_wait watches all fds; readiness fires a Wakerworker → reactor: poll returned Pending, register the Wakerreactor → worker: resource ready, requeue the task
The multi-thread scheduler. Each worker is one OS thread running a loop: pop a ready task from its local queue, call poll(task), then move to the next. If poll returns Ready the task is done; if it returns Pending the task is parked and its Waker is registered with the shared reactor (pink). When a worker's queue empties it steals ~half a sibling's tasks. The reactor runs one epoll_wait over all descriptors and, the moment a resource is ready, requeues that task onto a worker (green), where it gets polled again.

Where the epoll instance actually comes from

The reactor in the diagram is Tokio's I/O driver, and it is created once, when you build the runtime. You never call epoll_create: enable_io() does, through mio (the crate that abstracts epoll / kqueue / IOCP behind one API). #[tokio::main] just expands to this:

#[tokio::main]
async fn main() {
    // ...your async code...
}

// expands to roughly:
fn main() {
    let rt = tokio::runtime::Builder::new_multi_thread()
        .enable_io()    // start the I/O driver -> creates ONE epoll instance
        .enable_time()  // timer driver, for sleeps / timeouts
        .build()
        .unwrap();

    rt.block_on(async {
        // every task on this runtime shares that single epoll instance
    });
}

// enable_io() eventually calls, exactly once:
//   mio::Poll::new()  ->  epoll_create1(EPOLL_CLOEXEC)   // on Linux
// The reactor owns that Poll. When all run queues are empty, a worker
// parks on it via epoll_wait; the readiness it returns fires the Wakers.

That single epoll instance is shared by every task on the runtime. Individual descriptors are added to it lazily with epoll_ctl the first time each resource is polled, and removed when the resource is dropped. When every worker's run queue is empty, one worker parks on the driver and calls epoll_wait; the readiness it returns is what fires the wakers that put tasks back on the queues. Forget enable_io() and the first socket you await panics with "there is no reactor running", because the instance was never created.

Where does blocking go?

A blocking call on a worker thread would stall every task queued behind it, so Tokio keeps a separate, much larger blocking pool for spawn_blocking and for file I/O (which has no portable async interface). The async workers stay free to keep polling. This is also why the pool is an optimisation, not the essence of async: the same futures run correctly on the single-threaded current-thread scheduler, just without cross-core parallelism.

How other Rust runtimes schedule differently

Tokio is one point in a design space. Runtimes differ along a few axes: how threads map to cores, whether a task can migrate between threads (work stealing), whether tasks must be Send, and which async interface drives the reactor. That last choice is really "how are async resources scheduled": readiness (epoll / kqueue), completion (io_uring), or raw hardware interrupts.

RuntimeThreadsSteals?Send?I/O driverHow it schedules
Tokio · multi-threadM:N, ~1 worker/coreYesYesmio (epoll/kqueue/IOCP)per-core run queue + global injection + LIFO slot; an idle worker steals ~half a sibling's tasks
Tokio · current-thread1 threadNoNo (with LocalSet)mio (epoll/kqueue/IOCP)one cooperative queue, no parallelism; proof the pool is optional
smol / async-executorM:N, threads you spawnYesYes (LocalExecutor: no)async-io (epoll/kqueue)small global queue + per-worker queues; you pick the thread count
glommiothread-per-core, share-nothingNo (task pinned to its core)No (Rc / RefCell ok)io_uring (Linux)latency-aware task queues per core; no cross-core locks or stealing
embassy (embedded)1 per executor, no_stdNoNohardware IRQ + timers (no epoll)statically allocated tasks; idle means the CPU sleeps (WFI); interrupt-priority executors
Two big families. Work-stealing M:N (Tokio, smol) shares tasks across a thread pool, so tasks must be Send; great for general servers. Thread-per-core (glommio, compio) pins each task to one core and shares nothing, so there are no locks and no Send bound, usually paired with io_uring; great for storage-heavy or tail-latency work. Embedded (embassy) drops the OS entirely: the "reactor" is the interrupt controller.

// same task, different runtimes

Same Pattern, Different Runtimes

You have seen one runtime from the inside; now compare several from the outside. Pick up to four and put them head to head: first the same spawn & collect task as above (4 workers computing i * i) to see the bare syntax, then a real-world outbox-relay worker: a listener watches for new work (e.g. database LISTEN/NOTIFY), signals a wakeup, and N relay workers pick it up and process it concurrently. Each snippet has a short note on what makes that runtime distinctive.

Comparison table: concurrency runtimes

This is the per-runtime table: given those futures, how does each concrete runtime schedule the work and which primitives does it hand you (scheduler, I/O backend, spawn bounds, wakeup, fan-out, parallelism)? It deliberately skips "what a future is", that is the language-level story in Comparison table: futures across languages. Think of it as the same async value, viewed one layer lower.

Python asyncio Rust / Tokio Go (goroutines) Rust / smol
DescriptionSingle-threaded event loop. asyncio.Event for signalling, create_task for fan-out, gather to join.Multi-threaded work-stealing scheduler. tokio::sync::Notify for wakeup, tokio::spawn for fan-out, JoinSet to collect handles.M:N scheduler built into the runtime. No async/await: goroutines block transparently. Channels and sync primitives for coordination.Lightweight single-dependency executor. Same async/await, but uses smol::Timer and event_listener::Event instead of Tokio primitives.
SchedulerSingle-threaded event loop (selector / IOCP)Multi-threaded work-stealing (configurable thread count)M:N scheduler, goroutines multiplexed on OS threadsSingle or multi-threaded (smol::Executor or global)
I/O backendepoll / kqueue / IOCP via selectorsepoll / kqueue / IOCP (mio)netpoller (epoll / kqueue / IOCP)epoll / kqueue / IOCP (polling crate)
Spawn boundsNone: all tasks share one threadSend + 'static (tasks cross threads)None: any func, runtime handles itSend + 'static (LocalExecutor lifts it)
Wakeup / signalasyncio.Event (set / wait / clear)tokio::sync::Notify (notify_waiters / notified)Channels (chan struct{}) or sync.Condevent_listener::Event (notify / listen)
Fan-outasyncio.create_task + asyncio.gathertokio::task::JoinSet or tokio::spawngo func() + sync.WaitGroupsmol::spawn + futures_lite combinators
ParallelismNo (GIL). Use multiprocessing for CPU work.Yes, tasks distributed across a thread poolYes, GOMAXPROCS goroutines run in parallelYes, with smol::Executor on multiple threads
Best forI/O-bound services, rapid prototyping, scriptingGeneral-purpose servers, APIs, proxiesMicroservices, CLIs, network infrastructureCLIs, small services, libraries wanting minimal deps

Spawn & collect

Same task as the sync/async example: spawn 4 workers, each computes i * i, then collect the results

Python asyncio

One event loop on one thread. async def marks a coroutine; gather runs them concurrently but never truly in parallel (the GIL). The simplest mental model of the five.

import asyncio

async def worker(i: int) -> int:
    return i * i                       # a trivial async task

async def main() -> None:
    # create_task schedules them; gather awaits all, in order
    tasks = [asyncio.create_task(worker(i)) for i in range(4)]
    results = await asyncio.gather(*tasks)
    print(results)                     # [0, 1, 4, 9]

asyncio.run(main())                    # starts + drives the loop

Rust / Tokio

#[tokio::main] bootstraps a multi-thread runtime. spawn moves each task onto a thread pool, so they can run in real parallel; the trade-off is that tasks must be Send + 'static.

#[tokio::main]                       // sets up the async runtime
async fn main() {
    let mut handles = Vec::new();
    for i in 0..4 {
        // each worker is a lightweight task (~hundreds of bytes)
        handles.push(tokio::spawn(async move { i * i }));
    }
    // .await yields instead of blocking the thread
    let mut results = Vec::new();
    for h in handles {
        results.push(h.await.unwrap());
    }
    println!("{:?}", results); // [0, 1, 4, 9]
}

Go (goroutines)

No async or await keywords at all. go f() starts a goroutine; a WaitGroup joins them. Blocking calls yield the thread transparently.

package main

import (
	"fmt"
	"sync"
)

func main() {
	results := make([]int, 4)
	var wg sync.WaitGroup
	for i := 0; i < 4; i++ {
		wg.Add(1)
		go func(i int) { // no async keyword, just "go"
			defer wg.Done()
			results[i] = i * i
		}(i)
	}
	wg.Wait()            // join all goroutines
	fmt.Println(results) // [0 1 4 9]
}

Rust / smol

No runtime macro: block_on drives the futures directly. You assemble exactly the pieces you need, keeping the dependency footprint tiny.

fn main() {
    // block_on drives the futures; no #[main] macro needed
    smol::block_on(async {
        let mut handles = Vec::new();
        for i in 0..4 {
            handles.push(smol::spawn(async move { i * i }));
        }
        let mut results = Vec::new();
        for h in handles {
            results.push(h.await);   // smol tasks return the value directly
        }
        println!("{:?}", results);   // [0, 1, 4, 9]
    });
}

Outbox relay worker

A real-world pattern: 1 listener signals a wakeup, N workers relay in parallel

Python asyncio

Wakeup uses asyncio.Event: you must clear() it yourself after each wait(). Fan-out is create_task + gather; everything shares one thread, so no locks are needed.

import asyncio

class OutboxRelayWorker:
    def __init__(self, worker_count: int = 4):
        self._worker_count = worker_count

    async def _listen_loop(self, wakeup: asyncio.Event) -> None:
        while True:
            await asyncio.sleep(1)      # stand-in for pg LISTEN/NOTIFY
            wakeup.set()                # wake every waiter at once

    async def _relay_loop(self, wakeup: asyncio.Event) -> None:
        while True:
            await wakeup.wait()         # suspend until signalled
            wakeup.clear()              # must reset the flag manually
            print("relaying outbox batch...")

    async def run(self) -> None:
        wakeup = asyncio.Event()
        # create_task schedules coroutines on the single event loop
        tasks = [
            asyncio.create_task(self._listen_loop(wakeup), name="listen"),
            *[
                asyncio.create_task(self._relay_loop(wakeup), name=f"relay-{i}")
                for i in range(self._worker_count)
            ],
        ]
        await asyncio.gather(*tasks)    # join all tasks

asyncio.run(OutboxRelayWorker().run())

Rust / Tokio

Notify is shared via Arc because tasks land on different threads. notified() needs no manual clear. JoinSet owns the spawned handles and lets you join them all in one loop.

use tokio::sync::Notify;
use std::sync::Arc;

// Notify is shared via Arc because tasks run on different threads
async fn listen_loop(wakeup: Arc<Notify>) {
    loop {
        tokio::time::sleep(std::time::Duration::from_secs(1)).await;
        wakeup.notify_waiters();        // wake all current waiters
    }
}

async fn relay_loop(wakeup: Arc<Notify>) {
    loop {
        wakeup.notified().await;        // no manual clear needed
        println!("relaying outbox batch...");
    }
}

#[tokio::main]
async fn main() {
    let wakeup = Arc::new(Notify::new());
    let mut set = tokio::task::JoinSet::new();

    set.spawn(listen_loop(wakeup.clone()));    // onto the thread pool
    for _ in 0..4 {
        set.spawn(relay_loop(wakeup.clone())); // Send + 'static required
    }
    while set.join_next().await.is_some() {}   // join the whole set
}

Go (goroutines)

The channel *is* the signal: there is no separate Event/Notify type. Workers range over the channel; a WaitGroup replaces JoinSet/gather. Everything is plain blocking code the scheduler makes concurrent.

package main

import (
	"fmt"
	"sync"
	"time"
)

type OutboxRelayWorker struct {
	workerCount int
}

func (w *OutboxRelayWorker) listenLoop(wakeup chan struct{}) {
	for {
		time.Sleep(1 * time.Second)
		select {
		case wakeup <- struct{}{}: // non-blocking send
		default:                   // drop if no worker is ready
		}
	}
}

func (w *OutboxRelayWorker) relayLoop(wakeup chan struct{}, wg *sync.WaitGroup) {
	defer wg.Done()
	for range wakeup { // ranges until the channel closes
		fmt.Println("relaying outbox batch...")
	}
}

func (w *OutboxRelayWorker) Run() {
	wakeup := make(chan struct{}, 1) // the channel *is* the signal
	var wg sync.WaitGroup

	go w.listenLoop(wakeup) // no async keyword, just "go"

	for i := 0; i < w.workerCount; i++ {
		wg.Add(1)
		go w.relayLoop(wakeup, &wg)
	}
	wg.Wait()
}

func main() {
	(&OutboxRelayWorker{workerCount: 4}).Run()
}

Rust / smol

Same async/await as Tokio, but the primitives come from small crates: event_listener::Event replaces Notify, and you build the multi-thread Executor by hand instead of getting it from a macro.

use event_listener::Event;
use std::sync::Arc;

// event_listener::Event is smol's stand-in for Tokio's Notify
async fn listen_loop(wakeup: Arc<Event>) {
    loop {
        smol::Timer::after(std::time::Duration::from_secs(1)).await;
        wakeup.notify(usize::MAX);      // wake all listeners
    }
}

async fn relay_loop(wakeup: Arc<Event>) {
    loop {
        wakeup.listen().await;          // register + wait for a signal
        println!("relaying outbox batch...");
    }
}

fn main() {
    let wakeup = Arc::new(Event::new());
    // no runtime macro: build a multi-thread executor by hand
    let ex = Arc::new(smol::Executor::new());
    smol::block_on(ex.run(async {
        ex.spawn(listen_loop(wakeup.clone())).detach();
        for _ in 0..4 {
            ex.spawn(relay_loop(wakeup.clone())).detach();
        }
        futures_lite::future::pending::<()>().await;
    }));
}

These snippets are illustrative: each async runtime is an external crate, so it needs the matching dependency in Cargo.toml (tokio, smol, glommio). Only the std::thread version is self-contained; glommio additionally requires Linux io_uring, so it will not run in the standard Rust playground.

// answer them all to reach Poll::Ready

Quiz: Rust futures

A self-check on everything above. The questions run from medium to hard across the async model, the state machine, executors, Pin / Unpin, the Waker contract, and how Rust differs from Go, JavaScript, Python, C# and Kotlin. Nothing is saved, so run it as often as you like. Answer with a click, or with the keyboard (1-4 / A-D to answer, arrows to move).

// warm up, then the sharp edges

Rust futures, medium to hard

A few questions to settle the model, then the parts that actually bite: how the state machine, the executor, and the type system interact, and where Rust diverges from Go, JavaScript, Python, C#, and Kotlin. Hard questions come with the code that triggers them.

impl Future for You {
    type Output = Mastery;
    fn poll(self, cx) -> Poll<Output> {
        // answer them all to reach Poll::Ready
    }
}
7 medium 12 hard Send across awaitCancellation safetyPin / UnpinWaker contractStructured concurrency