Skip to content
DocspackagesDocumentation

@lunora/container

Deploy Docker containers alongside your Lunora app, called from actions over ctx.containers.

PackagesContainer

@lunora/container runs Cloudflare Containers as part of a Lunora app: declare a container with defineContainer, and codegen emits the container-enabled Durable Object class while lunora dev / lunora deploy reconcile the wrangler.jsonc wiring. wrangler deploy builds the image with your local Docker engine and pushes it to Cloudflare's registry. There is no separate container platform to sign up for.

Reach for a container when a job doesn't fit the Workers runtime: an existing binary (ffmpeg, headless Chrome, a Python ML model), a long-running process, or anything that needs a full filesystem.

pnpm add @lunora/container

Containers require the Workers Paid plan. Images must target linux/amd64; on Apple Silicon, build with --platform linux/amd64 (the scaffolded Dockerfile does this for you). See Limits.

Declare a container

Containers live in lunora/containers.ts. Scaffold one (definition plus a starter Dockerfile under containers/<name>/) with:

pnpm vis generate lunora-container --name=transcoder
// lunora/containers.ts
import { defineContainer } from "@lunora/container";

export const transcoder = defineContainer({
    image: "./containers/transcoder", // directory with a Dockerfile, or { registry: "docker.io/acme/transcoder:1.4" }
    defaultPort: 8080,
    instanceType: "standard-1", // lite | basic | standard-1..4 | { vcpu, memoryMib, diskMb }
    maxInstances: 5, // cap concurrent running instances (and the .any() pool)
    sleepAfter: "5m", // idle timeout — instances scale to zero
    secrets: ["TRANSCODER_API_KEY"], // Worker secrets forwarded into the container env
});

image is either a local path (a directory containing a Dockerfile, or a path to the Dockerfile itself) that wrangler deploy builds and pushes, or a pre-built registry reference ({ registry }) from the Cloudflare Registry, Docker Hub, or Amazon ECR.

Re-export the generated classes from your worker entry; wrangler requires every container's class to be exported by the deployed Worker:

// src/server/index.ts
export * from "../../lunora/_generated/containers";

lunora dev reminds you if this is missing, and reconciles the rest of wrangler.jsonc automatically: the containers[] entry, the CONTAINER_* Durable Object binding, the SQLite migration, and observability.enabled (so container logs are captured).

Call a container from an action

Container calls are external I/O, so (like ctx.fetch and ctx.ai) ctx.containers lives on actions, never queries or mutations. Codegen types one handle per declared container.

import { action, v } from "@/lunora/_generated/server";

export const transcode = action.input({ videoId: v.id("videos") }).action(async ({ ctx, args: { videoId } }) => {
    // One instance per entity — same id always routes to the same container.
    const response = await ctx.containers.transcoder.get(videoId).fetch("/transcode", {
        method: "POST",
        body: JSON.stringify({ videoId }),
    });

    return response.json();
});

Routing: .get(name) vs .any()

  • .get(name): one container instance per name. Use it for stateful, per-entity work: a sandbox per user, a room per game, a worker per job id. The same name always reaches the same instance. Names starting with pool- are reserved for the instances .any() / .pool() pick and are rejected, so an entity can never share an instance (disk, lifecycle) with a pool member.
  • .any(count?): a random instance from a fixed pool (defaults to maxInstances, else 3). Use it for stateless, interchangeable work where any instance can serve the request.
  • .pool(options?): like .any(), but resilient. Each fetch picks a random instance and, on a thrown error or a retryable response (5xx by default), retries on a freshly-picked instance with exponential backoff.
// Stateless pool — load-balanced across instances.
const probe = await ctx.containers.transcoder.any().fetch("/healthz");

// Resilient: rides over a single cold/unhealthy instance.
const result = await ctx.containers.transcoder
    .pool({ attempts: 3, backoffMs: 100, retryOn: (response) => response.status >= 500 })
    .fetch("/transcode", { method: "POST", body });

fetch accepts a path string (resolved against the container) or a full Request, and proxies WebSocket upgrades, so a real-time stream to a container works the same as any other fetch.

Cold-start retry

The first request to a sleeping or never-started instance can land while Cloudflare is still provisioning it, surfacing as a 503 "no Container instance available", a 500 "Failed to start container", a 429, or a thrown "not listening" error. .get() and .any() absorb that race: they retry the same instance a few times with exponential backoff (default 3 attempts, 500 ms base, 30 s ceiling) so transient provisioning blips don't reach your handler. Only those platform provisioning signals retry; an honest application 5xx is returned straight through.

// Tune or disable per call (attempts: 1 sends exactly once).
await ctx.containers.transcoder.get(jobId, { attempts: 5, backoffMs: 250 }).fetch("/transcode");

A pre-built Request is sent once and never retried, because its body may be a one-shot stream that can't be replayed. Pass a path string (the common case) to opt into the retry, or set attempts: 1 to send a Request without it. (This is distinct from .pool(), which retries on a fresh instance for stateless load-balancing; cold-start retry sticks to the one instance you named.)

Managing a named instance

The .get(name) handle also exposes lifecycle control, for the per-entity pattern where you manage an instance rather than wait for sleepAfter:

const sandbox = ctx.containers.codeRunner.get(userId);

await sandbox.start({ envVars: { SESSION: userId } }); // explicit start (optional per-instance env)
const state = await sandbox.getState(); // inspect runtime state
await sandbox.stop(); // stop (can restart on next request)
await sandbox.destroy(); // tear down and discard the ephemeral disk

.any() and .pool() return fetch-only handles; lifecycle control is for named instances you own.

Cloudflare has no built-in autoscaling yet: pools are a fixed size and .any()/.pool() pick uniformly, ignoring location. .pool() is the recommended call for stateless work until then; for batch work, drive instances from a scheduler cron or runAfter.

Per-instance images and snapshots

With the default scheduling policy, one image and one instance size in wrangler.jsonc serve every instance, and Cloudflare rolls changes out on deploy. schedulingPolicy: "durable_object" (Cloudflare public beta) moves that choice into each instance's start() instead. It suits sandboxes and agent environments, where each instance may need a different image or size.

// lunora/containers.ts
export const agentComputer = defineContainer({
    schedulingPolicy: "durable_object",
    images: {
        base: "./containers/base", // built and uploaded by wrangler deploy
        gpu: { registry: "registry.cloudflare.com/<account>/gpu@sha256:<digest>" },
    },
    image: "base", // default for a start that names none, including the implicit one a fetch triggers
    instanceType: "lite",
    defaultPort: 8080,
});
const computer = ctx.containers.agentComputer.get(taskId);

await computer.start({ image: "gpu", instanceType: "standard-2" });
const saved = await computer.snapshot({ name: "after-setup" }); // plain data, store it anywhere

await computer.stop();
await ctx.containers.agentComputer.get(otherTaskId).start({ snapshot: saved }); // restore the filesystem
  • image is a key of images or a Cloudflare-managed image such as "cloudflare/debian-trixie". The choice (and instanceType) is persisted like envVars, so the restart after a sleep boots the same image. A start that changes it on a running instance is rejected with CONFLICT.
  • A named { registry } image must be digest-pinned in the Cloudflare registry. Push Docker Hub or ECR images there first.
  • snapshot() saves the writable filesystem, not memory or processes. A restored container runs its entrypoint again. A snapshot is tied to the image it came from and expires 30 days after its last restore. start({ snapshot }) applies to that one start.
  • The policy takes no maxInstances or rollout: running instances count against the account limit, and each keeps its startup image across deploys. The policy is immutable. Switching means a new container application, which replaces every instance.

Needs wrangler 4.131.0 or newer: older releases reject scheduling_policy: "durable_object" and the images map at wrangler deploy. Only Cloudflare supports this. Codegen refuses schedulingPolicy: "durable_object" with a platform_unsupported_feature diagnostic on the Node and celld targets.

Sandbox helpers: files, backups and bucket mounts

Sandbox SDK 1.0 moved container control into the app's own Durable Object, which is what a LunoraContainer already is. Its three helpers, Files, DirectoryBackup and S3Mount, are built in. Opt a container in with sandbox: true. The helpers run Cloudflare's sandbox-shim binary inside the container, which must be at /usr/local/bin/sandbox-shim. Cloudflare publishes it as the image cloudflare/sandbox, which contains nothing but the shim. Copy it into your own image, pinned to the version matching @cloudflare/sandbox:

containers/workspace/Dockerfile
FROM debian:trixie-slim
COPY --from=docker.io/cloudflare/sandbox:1.0.0 /usr/local/bin/sandbox-shim /usr/local/bin/sandbox-shim
lunora/containers.ts
export const workspace = defineContainer({
    image: "./containers/workspace",
    sandbox: true,
    backups: { bucket: "WORKSPACE_BACKUPS", prefix: "workspaces/" },
});

sandbox must be a true/false literal: codegen reads it to decide what the generated container class extends and which Worker entrypoints to export. An app that never opts in never loads @cloudflare/sandbox.

Files

ctx.containers.workspace.get(id).files reads and writes the instance's own disk, the same disk exec and spawn run against. It starts the container if it is stopped.

const box = ctx.containers.workspace.get(userId);

await box.files.mkdir("/workspace/src", { recursive: true });
await box.files.writeFile("/workspace/src/main.ts", source);

const entries = await box.files.readDirectory("/workspace/src");
const text = await (await box.files.readFile("/workspace/out.log")).text();

readFile streams: the response body applies backpressure to the read inside the container, so a large file never sits in Worker memory. writeFile takes a string, bytes or a ReadableStream. A filesystem failure rejects with a LunoraError whose data.errno is the Linux code: ENOENT is NOT_FOUND, EACCES/EPERM are FORBIDDEN, and EEXIST/ENOTEMPTY/EBUSY are CONFLICT.

Backups

backup(directory) saves a directory to the R2 bucket named in backups and returns a plain record. Store it, and pass it to restore later, on this instance or another, even one running a newer image:

const record = await box.backup("/workspace", { exclude: ["node_modules"], gitignore: true });
await ctx.runMutation(internal.workspaces.saveBackup, { userId, record });

// Later, on a fresh instance:
await ctx.containers.workspace.get(userId).restore(record);

The container never holds bucket credentials. Each backup or restore goes through the DirectoryBackupGateway Worker entrypoint, which grants the container access to exactly one object for the length of that one operation, and every restore checks the object's SHA-256. A restore swaps the directory in only after the download is verified, so a rejected restore leaves the target as it was. deleteBackup(record) removes the object.

snapshot() (above) and backup() differ in scope. A snapshot captures the whole filesystem and restores only into the same container. A backup captures one directory into your own bucket and restores across images.

Bucket mounts

mount() attaches an S3-compatible bucket (R2, S3, GCS) at a path inside the container. Credentials are passed by secret name: the Durable Object reads them from the Worker env, and the S3Gateway entrypoint signs every storage request, so the keys never enter the container.

await box.mount({
    path: "/mnt/datasets",
    endpoint: `https://${accountId}.r2.cloudflarestorage.com`,
    region: "auto",
    bucket: "datasets",
    access: "read-only",
    credentials: { accessKeyIdSecret: "R2_ACCESS_KEY_ID", secretAccessKeySecret: "R2_SECRET_ACCESS_KEY" },
});

For R2 this needs an R2 API token's S3 keys; an R2 bucket binding is not used. A mounted path is not a POSIX filesystem, so don't use it for locks or atomic renames. inspectMount(path) reports a mount's state, and unmount(path) removes it.

A bucket mount cannot be combined with an egress policy (allowedHosts, deniedHosts, interceptHttps, outbound handlers or the runtime egress controls) on the same container: the policy's catch-all would take the mount's storage traffic. mount() refuses with BAD_REQUEST in that case. Backups do work with an egress policy, because their route is registered before the policy's catch-all at every start.

Codegen refuses sandbox: true on the Node and celld targets with a platform_unsupported_feature diagnostic (containerSandboxTools).

Running a command: exec

fetch covers containers that serve an HTTP API. When the container is a runner — a build box, a test sandbox, a job worker — what you want is a command and its exit code. exec is that contract:

const result = await ctx.containers.codeRunner.get(userId).exec("pnpm", {
    args: ["install", "--frozen-lockfile"],
    cwd: "/app",
    env: { CI: "1" },
    timeoutMs: 120_000,
});

if (result.code !== 0) {
    throw new Error(`install failed:\n${result.stderr}`);
}

It returns { code, stdout, stderr }, and it is available on every handle — .get(), .any(), .pool() — and through .port(n), inheriting the same cold-start retry and routing as fetch.

A non-zero exit code is not an error. A command that ran and failed is a result, so you get a code to branch on. exec throws only when the command could not be run: the container answered non-2xx (usually "no exec route"), the body was not JSON, or it carried no numeric code. That distinction is the point of the contract — the convention it replaced read the response body back as output, so a container with no exec route handed you its 404 page as though the command had succeeded.

Output is capped, and so is the clock

stdout and stderr come back as one JSON document, which has to be held whole to be parsed. A build box that emits 150MB would take the isolate — and every other request sharing it — down with it, so the response body is capped at 1MB by default and a runner that overruns it fails the call rather than the shard. Raise it per call with maxOutputBytes when a command legitimately produces more, or have the runner cap its own output.

timeoutMs covers the whole call, request and response body. That is not a detail: fetch resolves on headers, so a runner that answers 200 when the command starts and then stalls would otherwise sit past the deadline.

An exec is not cancelled when its deadline fires. exec reaches the container Durable Object over an RPC, and an RPC argument cannot carry an AbortSignal, so the deadline is enforced in your worker: it frees your handler on time, while the call it abandons keeps running and the command keeps going inside the container. The client always sends timeoutMs in the request body — have the runner enforce it, since the runner is the only place that can stop the command.

On .pool(), each attempt re-picks an instance, exactly like a pooled fetch — so a pooled exec must be safe to run more than once. Use .get(name) for anything that isn't idempotent. A pooled exec retries only the platform's cold-start transients (which mean the request never reached the container), not the any-5xx default a pooled fetch uses — a runner that ran the command and then failed must not have it run again.

The container side of the contract

On Cloudflare the container needs nothing: the container Durable Object runs the command through the runtime's native ctx.container.exec(), unshelled, reads stdout and stderr under the same maxOutputBytes cap (killing the process if it overruns, or when timeoutMs fires), and answers the contract itself. The container still has to start the usual way, so exec needs a defaultPort (or a .port(n) handle), as before. The command gets the container's start env (env, secrets, secretsStore values, or a start({ envVars }) override), with the call's own env over it. maxOutputBytes caps stdout and stderr each.

On Cloudflare this path skips an image-side /__lunora/exec runner entirely, so any allowlist, user switch or audit logging that runner applied no longer runs. Put those restrictions in the image itself (a non-root default user, a restricted PATH), and gate model-chosen commands at the caller.

Where the runtime has no native exec, the runner inside the container serves the route instead. Accept a POST at /__lunora/exec taking

{ "command": "pnpm", "args": ["install"], "cwd": "/app", "env": { "CI": "1" }, "maxOutputBytes": 1000000, "timeoutMs": 120000 }

and answer { "code": 0, "stdout": "…", "stderr": "…" } as JSON. cwd, env and timeoutMs are omitted from the body entirely when unset, so a runner never has to special-case an explicit undefined. Absent stdout/stderr in the reply default to ""; an absent code is an error, because "the command ran" and "the command's result is unknown" must not look the same.

Run the command directly — do not concatenate args into a shell string, which would make every argument an injection point.

/__lunora/* is reserved for Lunora's own container routes. handle.fetch refuses any path resolving into it — in any letter case, percent-encoded (once or more), or with ;params on the segment — so serve your application's routes elsewhere and reach exec through exec. The container Durable Object enforces it a second time on every HTTP entry: its fetch and its containerFetch RPC answer 403 for a path touching /__lunora/* (any segment, same spellings), whatever headers the request carries. exec reaches the route over a separate RPC (lunoraExec), which only code holding the container binding can call, so forwarding an inbound request as-is (env.CONTAINER_X.get(id).fetch(request)) cannot open it either.

exec runs whatever it is given. When a model chooses the command — an @lunora/agent containerTool, say — gate it: that tool's default approval policy pauses for a human on exec and on any non-idempotent fetch. A fetch cannot be used to reach the exec route around that gate — /__lunora/* is reserved and handle.fetch refuses it — but restrict what can run inside the container too, rather than relying on the caller.

Streaming a process: spawn

exec buffers: it returns once the command has exited, with at most 1MB of output. spawn is the streaming counterpart for long builds, live logs and interactive programs. It needs the runtime's native exec, and is available on named instances (.get(name)), since a process belongs to one container:

const process = await ctx.containers.codeRunner.get(userId).spawn("pnpm", {
    args: ["test", "--reporter=dot"],
    cwd: "/workspace",
    timeoutMs: 600_000,
});

for await (const chunk of process.stdout!) {
    // Forward to a client, write to R2, …
}

const code = await process.exitCode;

The result carries stdout/stderr streams, exitCode (a promise), kill() and the process pid. The container finishes a process's output, and settles its exit code, only while every output stream is being drained. Lunora drains both for you into a 1MB buffer per stream, so you can read them in any order, or await exitCode before reading, as long as the stream you are not reading stays under that size. For a noisier process, read both concurrently. Pass stdin: true to get a writable stdin, or a ReadableStream to pipe in as the whole input. Pass pty: { cols, rows } to run on a pseudo-terminal: stderr is merged into stdout, and resize() works. An AbortSignal in signal kills the process.

The container is started if it is stopped, and kept awake until the process exits. sleepAfter does not stop it underneath a running process. A spawned process has no default deadline, so set timeoutMs for anything that should not run indefinitely.

Browser terminal

terminal(request) answers a WebSocket upgrade with a shell on a PTY, bridged to the socket. Browser routes have no ctx.containers, so reach the instance with getContainer(env, exportName, name):

lunora/http.ts
import { getContainer } from "@lunora/container";

app.get("/terminal", async (c) => {
    const userId = await requireUser(c); // whoever holds the socket has a shell
    return getContainer(c.env, "workspace", userId).terminal(c.req.raw, { command: "bash", cwd: "/workspace" });
});

The wire protocol is what xterm.js speaks. Binary frames are keystrokes, and a text frame {"cols":120,"rows":40} resizes the terminal. Any other text frame counts as keystrokes too. Output arrives as binary frames, the socket closes with 1000 when the shell exits, and closing the socket kills the shell. The initial size comes from the request's ?cols=&rows= search params, else 80×24.

client
import { FitAddon } from "@xterm/addon-fit";
import { Terminal } from "@xterm/xterm";

const term = new Terminal();
const fit = new FitAddon();
term.loadAddon(fit);
term.open(element);
fit.fit();

const socket = new WebSocket(`wss://${location.host}/terminal?cols=${term.cols}&rows=${term.rows}`);
socket.binaryType = "arraybuffer";
socket.onmessage = (event) => term.write(new Uint8Array(event.data));
term.onData((data) => socket.send(new TextEncoder().encode(data)));
term.onResize(({ cols, rows }) => socket.send(JSON.stringify({ cols, rows })));

The bridge runs where terminal() is called, on top of spawn. No WebSocket crosses into the container Durable Object, which keeps its reserved /__lunora/* routes closed to request content.

Previews

A dev server in the container is already reachable through fetch: route a browser request to its port with .port(n), WebSocket upgrades (HMR) included.

lunora/http.ts
app.all("/previews/:id/*", async (c) => {
    await requireOwner(c, c.req.param("id"));
    const path = c.req.path.replace(`/previews/${c.req.param("id")}`, "") || "/";
    const url = new URL(path + new URL(c.req.url).search, "http://container");

    return getContainer(c.env, "workspace", c.req.param("id")).port(5173).fetch(new Request(url, c.req.raw));
});

The server sees Host: container, and a link that starts with / leaves the /previews/<id> prefix. Use relative links, or serve each preview on its own hostname: a wildcard route (*.preview.example.com/*) where the Worker reads the instance id from the subdomain, so the app runs at / and absolute links work. To share a preview, mint an expiring token, store it, and check it in the route before forwarding. Authenticate whoever mints the token, because a token is access to the dev server.

Multi-port containers

A container that listens on more than one port (an app port plus an admin port, say) declares every one with requiredPorts; start-up waits for all of them to be listening. defaultPort is the target when a request doesn't pick a port; route a single request elsewhere with .port(n), which composes with .get(), .any(), and .pool():

export const app = defineContainer({
    image: "./containers/app",
    defaultPort: 8080,
    requiredPorts: [8080, 9090], // app + admin
});
// in an action:
await ctx.containers.app.get(tenantId).fetch("/work"); // → 8080 (defaultPort)
await ctx.containers.app.get(tenantId).port(9090).fetch("/admin"); // → 9090

Egress firewall

By default a container may open any outbound connection (enableInternet: true), and egress is billed per GB. Constrain it by pairing enableInternet: false with an allowedHosts allow-list, or layer a deniedHosts deny-list that overrides everything, including enableInternet: true and allowedHosts. Glob patterns like *.stripe.com are supported.

export const fetcher = defineContainer({
    image: "./containers/fetcher",
    enableInternet: false,
    allowedHosts: ["*.stripe.com", "api.github.com"],
    deniedHosts: ["*.evil.com"],
    interceptHttps: true, // extend the lists to TLS traffic (image must trust the Cloudflare CA)
});

interceptHttps extends the allow/deny lists to HTTPS connections, not just plain HTTP; it requires the image to trust the Cloudflare CA at /etc/cloudflare/certs/cloudflare-containers-ca.crt. The interception path runs through the ContainerProxy worker entrypoint, which codegen re-exports from the generated container file automatically.

Tighten or relax a single running instance at runtime through its named handle. egress.allow / deny add one host, removeAllowed / removeDenied drop one, and setAllowed / setDenied replace a whole list:

await ctx.containers.fetcher.get(tenantId).egress.allow("hooks.slack.com");
await ctx.containers.fetcher.get(tenantId).egress.setDenied(["*.evil.com", "*.tracking.example"]);

For advanced egress rewriting in worker code (inject auth, route, or mock a container's outbound calls), @lunora/container/do re-exports Cloudflare's custom outbound-handler types (OutboundHandler, OutboundHandlers, outboundParams); wire them onto a hand-authored LunoraContainer subclass.

Readiness gating

The platform health check waits for an open port (defaultPort) and, if pingEndpoint is set, for an HTTP probe on that path on the same container. Neither means the app is ready. readyOn adds application-level probes that gate request proxying: a ctx.containers.<name> fetch holds until every probe responds with its expected status, so callers never hit a container still applying migrations or warming caches.

export const api = defineContainer({
    image: "./containers/api",
    defaultPort: 8080,
    readyOn: [
        { path: "/ready" }, // expect 200 on defaultPort
        { path: "/live", port: 9090, status: 204 }, // own port + expected status
    ],
});

Each probe declares a path (a leading slash is optional), an optional port (defaults to defaultPort), and an optional status (defaults to 200). Probes are declarative data (no handler functions), so codegen and the config layer read them without evaluating code. At start they run in parallel, poll the container's TCP port directly, and fail the start if any probe doesn't go ready within the readiness budget.

Hard timeout

sleepAfter caps idle time; hardTimeout caps total lifetime: a runaway-cost backstop that fires regardless of activity, measured from start. It uses the same grammar as sleepAfter ("30s", "5m", "1h", or a plain number of seconds):

export const job = defineContainer({
    image: "./containers/job",
    sleepAfter: "5m", // sleep after 5 min idle
    hardTimeout: "1h", // …stopped after an hour, busy or not
});

When it elapses, the generated class's onHardTimeoutExpired hook runs; the default action is stop(), which sends SIGTERM and does not escalate. A container that traps or ignores SIGTERM therefore keeps running past its cap — if yours might, override onHardTimeoutExpired and follow the stop with a destroy() after whatever grace period your workload needs. The timer is armed through the container's own scheduler and stamped with a run generation, so a stale timer left over from a previous (slept or crashed) run never kills a fresh one. Advanced apps that hand-author their container class can override onHardTimeoutExpired (the base class lives in @lunora/container/do) to drain or checkpoint before stopping.

Calling Lunora from inside a container

Container code calls back into your app's functions with the bridge client (any JS runtime: Node, Bun, Deno) over the Worker's HTTP RPC endpoint, so a container reads/writes app state through the same queries/mutations the browser uses, instead of reaching into a database directly.

import { createContainerBridge } from "@lunora/container/bridge";

const lunora = createContainerBridge({ baseUrl: process.env.LUNORA_URL, token: process.env.LUNORA_TOKEN });

const pending = await lunora.query("jobs:listPending", { limit: 10 });
await lunora.mutation("jobs:markDone", { id: pending[0].id });

For full type-safety, pass a generated api reference to run(); args and result are inferred from it:

import { api } from "../lunora/_generated/api";

const job = await lunora.run(api.jobs.next, { queue: "transcode" }); // typed args + result

The token is a bearer your Worker's resolveIdentity recognizes. Forward it as a container secret, never bake it into the image. Non-JS containers can POST /_lunora/rpc with { functionPath, args } directly (same contract).

Securing the bridge

The bridge sends its token as Authorization: Bearer <token>. Your Worker's resolveIdentity is what validates it and maps it to the identity the called functions run as. Without that check, anyone who reaches /_lunora/rpc runs as whatever identity you return. Validate the bearer against a Worker secret and return null for anything that doesn't match (an unrecognised request runs anonymously and fails your functions' own authorization checks):

// in your worker options (createWorker / withLunora)
resolveIdentity: (request, env) => {
    const header = request.headers.get("authorization");
    const token = header?.startsWith("Bearer ") ? header.slice("Bearer ".length) : undefined;

    // Compare against a Worker secret you also forward to the container.
    if (!token || token !== env.LUNORA_CONTAINER_TOKEN) {
        return null; // anonymous — function-level authorization still applies
    }

    return { userId: "container:transcoder" }; // identity the bridge calls run as
},

Generate the token once, set it with wrangler secret put LUNORA_CONTAINER_TOKEN for production (and in .dev.vars locally), and pass the same value to the container as a secret so it can send it back. Never bake it into the image.

Building with Railpack (no Dockerfile)

Point image at a { build } source directory and Lunora builds the image with Railpack, Railway's Dockerfile-less builder, instead of you writing a Dockerfile:

export const transcoder = defineContainer({
    image: { build: "./services/transcoder" }, // Railpack detects the stack and builds it
    defaultPort: 8080,
});

On lunora deploy, each { build } container is built with Railpack and pushed to the Cloudflare Registry (under a deterministic lunora-<name>:build tag) before the Worker deploys. This needs the railpack CLI and a running BuildKit instance reachable via BUILDKIT_HOST, e.g.:

docker run --rm --privileged -d --name buildkit moby/buildkit
export BUILDKIT_HOST=docker-container://buildkit

lunora deploy preflights both and stops with setup directions if either is missing. The Dockerfile path remains the zero-extra-deps default; Railpack is opt-in.

Lifecycle logs

The generated container classes emit a structured event on start, stop, and error, tagged with the container name and per-instance id (the container's CLOUDFLARE_DURABLE_OBJECT_ID). lunora dev surfaces them inline with your ctx.log and RPC lines:

[lunora] container:transcoder#a1b2c3d4  start
[lunora] container:transcoder#a1b2c3d4  error  exited unexpectedly

The same events also appear in the Studio Logs panel: each lifecycle transition is best-effort pushed into the root shard's log buffer under the container:<name> source, so a crash-looping container shows up next to your function logs without leaving the dashboard. The push is best-effort by design: if it can't reach the shard, the lunora dev terminal above remains the source of truth.

Advisor lints

Two advisor lints flag container cost/security footguns in the Studio Advisors table and at codegen time:

  • container_oversized_instance: a large standard-3/standard-4 (or large custom) instance, whose provisioned memory + disk are billed while any instance runs.
  • container_public_internet: enableInternet left at the default (enabled); set it explicitly (false to close egress, true to opt in).

Configuration

FieldPurpose
imageLocal Dockerfile path/directory, { registry } for a pre-built image, or { build } (Railpack). Required.
defaultPortPort the container listens on; fetch targets it. Must be EXPOSEd for local dev.
requiredPortsEvery port start-up must wait for (multi-port); route a request to one with .port(n).
instanceTypelite | basic | standard-1..4, or a custom { vcpu, memoryMib, diskMb }.
maxInstancesCap on concurrently running instances (and the default .any() pool size).
sleepAfterIdle timeout before an instance sleeps, e.g. "5m", "30s", or seconds. Default "10m".
hardTimeoutHard cap on total lifetime from start, ignoring activity. Same grammar as sleepAfter.
readyOnApplication-level readiness probes that gate request proxying until the app reports ready.
pingEndpointHTTP path the platform polls to decide an instance is healthy. Defaults to upstream's "ping".
envStatic environment variables passed to the container on every start (runtime). For secrets use secrets.
secretsNames of Worker secrets forwarded into the container env at start.
buildArgsBuild-time args for a built image (wrangler image_vars, like --build-arg); ignored for { registry }.
entrypointOverride the image's ENTRYPOINT/CMD for every start.
enableInternetWhether the container may open outbound connections. Default true; egress is billed.
allowedHostsEgress allow-list (globs): hosts reachable even with enableInternet: false.
deniedHostsEgress deny-list (globs): overrides everything, including allowedHosts.
interceptHttpsExtend the egress lists to HTTPS too (image must trust the Cloudflare CA). Default false.
labelsKey-value metadata attached to every instance for metrics/observability.
nameOverride the wrangler containers[].name identifier.
rolloutRolling-deploy tuning: { stepPercentage, gracePeriodSeconds }.
schedulingPolicy"durable_object" picks image and size per start(); see per-instance images.
imagesNamed images under schedulingPolicy: "durable_object"; image then names the default among them.

Secrets and environment

env holds static values; secrets names Worker secrets, managed exactly like every other Lunora secret. List them in .dev.vars for local dev and wrangler secret put for production (see Deployment). A declared secret that isn't set fails fast at container start with a directed error, rather than launching the container without it.

export const worker = defineContainer({
    image: "./containers/worker",
    env: { LOG_LEVEL: "info" },
    secrets: ["STRIPE_API_KEY"], // resolved from env.STRIPE_API_KEY at start
});

To pull from a Cloudflare Secrets Store binding instead of a plain Worker secret, map container env-var name → Secrets Store binding name with secretsStore. Each binding is resolved with its async .get() the first time the instance starts (memoised thereafter) and injected as that env var:

export const worker = defineContainer({
    image: "./containers/worker",
    secretsStore: { STRIPE_KEY: "STRIPE_SECRET" }, // env.STRIPE_SECRET.get() → STRIPE_KEY inside the container
});

A name that collides with env/secrets, or a binding missing from the Worker env, is rejected (the collision at authoring time, the missing binding at start): the same fail-closed stance as secrets.

A per-instance start({ envVars }) replaces the env set wholesale (the Secrets Store is not read) and is persisted for that instance: every later start uses it — an explicit start(), and the implicit restart a fetch or exec triggers after the container slept, crashed or hit hardTimeout — until you destroy() the instance. That makes start({ envVars: {} }) (or with only what the job needs) the way to run an untrusted or model-driven exec sandbox without the credentials the definition declares, and keep it that way across restarts. The override cannot change while the container is running or still starting: a start({ envVars }) that differs from its env is rejected with CONFLICT in either case, rather than silently ignored — stop() the instance (and let any start in flight finish) first. Pass every variable the container needs.

Both env and secrets are runtime values, available when the container starts. For build-time values (docker build --build-arg, exposed to the Dockerfile as ARG, wrangler's image_vars), use buildArgs. They only apply to an image Lunora builds (a Dockerfile or Railpack { build } source) and are ignored for a pre-built { registry } image:

export const worker = defineContainer({
    image: "./containers/worker",
    buildArgs: { NODE_VERSION: "22", BUILD_TARGET: "production" }, // → docker --build-arg
});

Dockerfile

The scaffolded Dockerfile encodes the platform's requirements:

# linux/amd64 is required — keep it explicit so Apple Silicon builds don't
# produce an arm64 image that fails at deploy.
FROM --platform=linux/amd64 node:22-slim

WORKDIR /app
COPY . .

# Local dev needs the listening port EXPOSEd (production exposes all ports).
EXPOSE 8080

# Exec form (not shell form) so SIGTERM reaches the process — Cloudflare sends
# SIGTERM on rollouts and gives 15 minutes before SIGKILL.
ENTRYPOINT ["node", "server.mjs"]

Disk is ephemeral: every (re)start gives a fresh filesystem. For persistence, write to @lunora/storage (R2).

Pushing to a Cloudflare Artifacts repo

A container's disk is ephemeral. When what it produces is a tree of files you want history and diffs for (an agent's workspace, a generated site, a build output), push it to a Cloudflare Artifacts repo. The Artifacts binding can't write files, so the push comes from git inside the container, authenticated with short-lived tokens the action mints.

The build runs code the container fetched (pnpm install, pnpm run build, and every dependency's scripts). The recipe keeps tokens out of the build's environment and out of the checkout, which stops a careless build from picking one up. It does not isolate the push from a build that is trying to: see the rules below before you run code you don't trust.

import { internalAction } from "@/lunora/_generated/server";
import { v } from "@lunora/values";

// Single-quoted: `exec` runs the command unshelled, so `sh -c` expands the
// variables inside the container. `core.hooksPath=/dev/null` and `--no-verify`
// keep hooks the build may have written from running.
const CHECKOUT =
    'rm -rf /work && if [ "$EMPTY" = 1 ]; then git init -b "$BRANCH" /work; else git -c core.hooksPath=/dev/null clone --branch "$BRANCH" "$REMOTE" /work; fi';
const BUILD = "pnpm install --frozen-lockfile && pnpm run build";
// Prints `unchanged` instead of committing when the build changed nothing.
const COMMIT =
    'git -C /work add -A && if git -C /work diff --cached --quiet; then echo unchanged; else git -C /work -c core.hooksPath=/dev/null -c user.name=lunora -c user.email=bot@example.com commit --no-verify -q -m "$MESSAGE"; fi';
const PUSH = 'git -C /work -c core.hooksPath=/dev/null push "$REMOTE" "HEAD:$BRANCH"';

/**
 * Hand a token to one `git` process as an `http.<remote>.extraHeader` through
 * git's `GIT_CONFIG_*` environment: it is in neither the command line nor
 * `/work/.git/config`, and it is sent only to `remote`, so a `url.*.insteadOf`
 * rewrite can't redirect it to another host. The `?expires=` suffix is not part
 * of the secret.
 */
const gitAuthEnv = (remote: string, plaintext: string): Record<string, string> => {
    return {
        GIT_CONFIG_COUNT: "1",
        GIT_CONFIG_KEY_0: `http.${remote}.extraHeader`,
        GIT_CONFIG_VALUE_0: `Authorization: Basic ${btoa(`x:${plaintext.split("?expires=")[0] ?? ""}`)}`,
    };
};

// Internal: only your own code decides who may push to which repo.
export const buildAndPush = internalAction.input({ repo: v.string(), sessionId: v.string() }).action(async ({ ctx, args }) => {
    const { defaultBranch, remote } = await ctx.artifacts.info(args.repo);
    // A freshly created repo has no commit, so no branch to clone yet.
    const empty = (await ctx.artifacts.withRepo(args.repo, (repo) => repo.log({ limit: 1 }))).length === 0;
    const box = ctx.containers.workspace.get(args.sessionId);
    const base = { BRANCH: defaultBranch, MESSAGE: `build ${args.sessionId}`, REMOTE: remote };
    const live = new Set<string>();

    const mint = async (scope: "read" | "write") => {
        const token = await ctx.artifacts.withRepo(args.repo, (repo) => repo.createToken(scope, 600));

        live.add(token.id);

        return token;
    };
    // A failed revoke must not mask the result; the token still expires on its own.
    const revoke = async (id: string): Promise<void> => {
        live.delete(id);

        try {
            await ctx.artifacts.withRepo(args.repo, (repo) => repo.revokeToken(id));
        } catch (error: unknown) {
            ctx.log.warn("artifacts token revoke failed", { error, tokenId: id });
        }
    };

    try {
        const readToken = empty ? undefined : await mint("read");
        const checkout = await box.exec("sh", {
            args: ["-c", CHECKOUT],
            env: { ...base, EMPTY: empty ? "1" : "0", ...(readToken === undefined ? {} : gitAuthEnv(remote, readToken.plaintext)) },
            timeoutMs: 120_000,
        });

        // The read token's job is done before any of the repo's code runs.
        if (readToken !== undefined) {
            await revoke(readToken.id);
        }

        if (checkout.code !== 0) {
            throw new Error(`checkout failed (exit ${String(checkout.code)})`);
        }

        // No token in this env. Install and build logs routinely pass exec's 1MB
        // output default, so raise the cap (or have the build write less).
        const build = await box.exec("sh", { args: ["-c", BUILD], cwd: "/work", maxOutputBytes: 16 * 1024 * 1024, timeoutMs: 600_000 });

        if (build.code !== 0) {
            throw new Error(`build failed (exit ${String(build.code)})`);
        }

        const commit = await box.exec("sh", { args: ["-c", COMMIT], env: base, timeoutMs: 60_000 });

        if (commit.code !== 0) {
            throw new Error(`commit failed (exit ${String(commit.code)})`);
        }

        if (commit.stdout.trim() === "unchanged") {
            return { changed: false, pushed: false };
        }

        // The write token exists only now, and only in the push's environment.
        const writeToken = await mint("write");
        const push = await box.exec("sh", { args: ["-c", PUSH], env: { ...base, ...gitAuthEnv(remote, writeToken.plaintext) }, timeoutMs: 120_000 });

        return { changed: true, pushed: push.code === 0 };
    } finally {
        // Revoke by id whatever happened — the write token, and the read token if the checkout threw.
        for (const id of [...live]) {
            await revoke(id);
        }
    }
});

After the push, ctx.artifacts.withRepo(repo, (r) => r.readFile({ ref: defaultBranch, path })) reads the pushed files back from any action, with no container running. The checkout starts from rm -rf /work, so calling the action again on the same session starts clean.

Keeping the token contained

  • Mint late, and only for the step that needs it. The clone gets a read token that is revoked as soon as the checkout ends. The write token is minted after the build and only exists in the push's environment. The install, build and commit run with neither. Hooks are disabled on every git step, so a .git/hooks or core.hooksPath the build wrote doesn't run with a token in its environment.

  • Pass the token through GIT_CONFIG_*, never through args, a URL or a log line. It never lands in /work/.git/config or in a process's arguments, and the header is scoped to the repo's remote. Don't print it, don't return it from the action, and don't leave set -x on in a script that uses it.

  • If the build runs code you don't trust, push from a separate container session. Everything above shares one container with the build, and a build that wants the token can still reach the push:

    • a process it leaves running reads the push's environment from /proc/<pid>/environ;
    • it writes ~/.gitconfig or /etc/gitconfig (a proxy, a credential helper, its own http.*.extraHeader), which the push's -c flags don't override;
    • it rewrites /work/.git/config, for example a core.fsmonitor command;
    • it puts its own git earlier on PATH.

    So for untrusted code, build in one session and move only the output out: with sandbox: true, backup() the output directory (never its .git) and restore() it into a fresh session the build never ran in (see Backups), or copy it through R2 with @lunora/storage. Clone, commit and push from that fresh session. The hardening above is for builds you trust; it is not a sandbox.

  • The image needs git. node:22-slim doesn't include it, so add RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates to the Dockerfile.

This is a different trust model from saving files to R2 with @lunora/storage, where the Worker holds the credentials and the container never sees them. Here a repo-scoped write credential lives inside the container for the length of the push. Choose Artifacts when you want Git history, diffs and forks of the output, or want other Git clients to clone it. Choose R2 when you want an opaque snapshot restored as is.

For repos too large to clone on every start, mount them with ArtifactFS inside the image. That is an image-level choice Lunora doesn't wrap.

Local development

lunora dev (and vite dev) build and run containers locally through the Cloudflare Vite plugin, so a Docker-compatible engine (Docker Desktop, Colima) must be running; Lunora warns up front if it isn't. A few things differ from production:

  • Container code is not hot-reloaded. Edit the Worker and it reloads; press r to rebuild the container.
  • Ports must be EXPOSEd in the Dockerfile (production exposes all ports).
  • vite dev can't pull from the Cloudflare Registry. Use a local Dockerfile, FROM-ing a registry image if you need one.

Testing

ctx.containers has a Docker-free test double, so action tests stay fast and deterministic. createContainerTestContext maps each container to a fetch handler that plays the container:

import { createContainerTestContext } from "@lunora/container";

const containers = createContainerTestContext({
    transcoder: async (request, { name }) => new Response(JSON.stringify({ ok: true, name })),
});

// Inject `containers` as ctx.containers when exercising the action.

A named instance from the double also has files, backed by an in-memory disk per instance with Linux-shaped errors (ENOENT, EEXIST, ENOTEMPTY), so a handler that writes and reads container files can be tested without Docker. spawn, backup/restore and mount reject with NOT_IMPLEMENTED: there is no process or bucket behind the double, so stub those handle methods in the test.

Run the real container behind lunora dev for integration tests; keep those in a Docker-enabled CI job, separate from the default unit-test matrix.

CLI

Manage images and instances with lunora containers, thin wrappers over wrangler containers with a Docker preflight on the build/push paths:

lunora containers build ./containers/transcoder --tag transcoder:v1 --push
lunora containers push transcoder:v1
lunora containers images list
lunora containers images delete transcoder:v1

Splitting build/push from lunora deploy lets CI build the image in one job and deploy in another. lunora deploy itself runs a Docker preflight and stops with an actionable message when a Dockerfile-built container is declared but no engine is available.

Stability and how a surface would be removed

The package is Experimental: it sits outside the Lunora 1.0 SemVer promise, and the graduation bar is what moves it out. Its public surface is snapshotted in api-snapshots/container.api.md and gated by pnpm run api:check, so every addition, removal and signature change shows up as a reviewed diff — no export carries @experimental individually any more, which is what makes the signatures tracked rather than skipped.

What that commits us to once the package graduates:

  • Names are the contract. The binding name (CONTAINER_<NAME>), the generated class name (<Name>Container), the exec route (/__lunora/exec) and the exec body keys (command, args, cwd, env, timeoutMs → code, stdout, stderr) are the names we intend to keep. A container image built against them keeps working.
  • Removal is deprecate-then-remove. A surface being withdrawn is first marked deprecated in the snapshot and docs, and keeps working for a full minor cycle with a runtime warning naming its replacement; only then is it removed, in a major. Pre-1.0 (alpha/next/beta) the old path is deleted in the same change as the new one, per the repo's branch policy — which is exactly what happened when exec replaced the ad-hoc POST-to-/exec convention.
  • Config keys outlive their implementation. defineContainer options are additive; one that stops being meaningful is accepted and ignored with a warning rather than becoming a hard error mid-cycle.
  • Errors are identified by code, not prose. Every failure raises a LunoraError with a stable code; message text may be reworded to be clearer without that counting as a break, so match on the code.

Known platform limitations

Some constraints live in Cloudflare Containers itself, not in this package, and are worth knowing before you design around them. They track open issues on cloudflare/containers. Lunora papers over what it can (cold-start retry, WebSocket keep-alive); the rest are listed here.

  • No autoscaling or location-aware routing. Pools are a fixed size and .any()/.pool() pick uniformly at random, ignoring instance region. Drive batch work from a scheduler cron and prefer .pool() for stateless requests. (cloudflare/containers#226)
  • Disk is ephemeral. Every (re)start gives a fresh filesystem; persist to @lunora/storage (R2). FUSE mounts, a writable tmpfs, and some low-level node:net socket modes are not available in the container sandbox. (cloudflare/containers#112, #160, #67)
  • Egress interception is HTTP-first. The allowedHosts/deniedHosts firewall gates HTTP egress out of the box; HTTPS needs interceptHttps: true plus the Cloudflare CA trusted in the image, and raw gRPC interception isn't supported yet. (cloudflare/containers#195)
  • Long jobs can be terminated on rollout. A new deploy moves instances after rollout.gracePeriodSeconds; a job longer than that may be sent SIGTERM (15 min before SIGKILL). Use hardTimeout as a cost backstop and make long work resumable. (cloudflare/containers#138)
  • Local dev can't pull from the Cloudflare Registry. lunora dev builds from a local Dockerfile; FROM a registry image inside it if you need one. (cloudflare/containers#155)