@lunora/container runs Cloudflare Containers
as part of a Lunora app: declare a container with defineContainer, and codegen
emits the container-enabled Durable Object class while lunora dev / lunora deploy
reconcile the wrangler.jsonc wiring. wrangler deploy builds the image with
your local Docker engine and pushes it to Cloudflare's registry. There is no
separate container platform to sign up for.
Reach for a container when a job doesn't fit the Workers runtime: an existing binary (ffmpeg, headless Chrome, a Python ML model), a long-running process, or anything that needs a full filesystem.
pnpm add @lunora/containerContainers require the Workers Paid plan. Images must target linux/amd64; on Apple Silicon, build with --platform linux/amd64 (the scaffolded
Dockerfile does this for you). See Limits.
Declare a container
Containers live in lunora/containers.ts. Scaffold one (definition plus a
starter Dockerfile under containers/<name>/) with:
pnpm vis generate lunora-container --name=transcoder// lunora/containers.ts
import { defineContainer } from "@lunora/container";
export const transcoder = defineContainer({
image: "./containers/transcoder", // directory with a Dockerfile, or { registry: "docker.io/acme/transcoder:1.4" }
defaultPort: 8080,
instanceType: "standard-1", // lite | basic | standard-1..4 | { vcpu, memoryMib, diskMb }
maxInstances: 5, // cap concurrent running instances (and the .any() pool)
sleepAfter: "5m", // idle timeout — instances scale to zero
secrets: ["TRANSCODER_API_KEY"], // Worker secrets forwarded into the container env
});image is either a local path (a directory containing a Dockerfile, or a
path to the Dockerfile itself) that wrangler deploy builds and pushes, or a
pre-built registry reference ({ registry }) from the Cloudflare Registry,
Docker Hub, or Amazon ECR.
Re-export the generated classes from your worker entry; wrangler requires every container's class to be exported by the deployed Worker:
// src/server/index.ts
export * from "../../lunora/_generated/containers";lunora dev reminds you if this is missing, and reconciles the rest of
wrangler.jsonc automatically: the containers[] entry, the CONTAINER_*
Durable Object binding, the SQLite migration, and observability.enabled (so
container logs are captured).
Call a container from an action
Container calls are external I/O, so (like ctx.fetch and ctx.ai)
ctx.containers lives on actions, never queries or mutations. Codegen types
one handle per declared container.
import { action, v } from "@/lunora/_generated/server";
export const transcode = action.input({ videoId: v.id("videos") }).action(async ({ ctx, args: { videoId } }) => {
// One instance per entity — same id always routes to the same container.
const response = await ctx.containers.transcoder.get(videoId).fetch("/transcode", {
method: "POST",
body: JSON.stringify({ videoId }),
});
return response.json();
});Routing: .get(name) vs .any()
.get(name): one container instance per name. Use it for stateful, per-entity work: a sandbox per user, a room per game, a worker per job id. The same name always reaches the same instance. Names starting withpool-are reserved for the instances.any()/.pool()pick and are rejected, so an entity can never share an instance (disk, lifecycle) with a pool member..any(count?): a random instance from a fixed pool (defaults tomaxInstances, else 3). Use it for stateless, interchangeable work where any instance can serve the request..pool(options?): like.any(), but resilient. Eachfetchpicks a random instance and, on a thrown error or a retryable response (5xxby default), retries on a freshly-picked instance with exponential backoff.
// Stateless pool — load-balanced across instances.
const probe = await ctx.containers.transcoder.any().fetch("/healthz");
// Resilient: rides over a single cold/unhealthy instance.
const result = await ctx.containers.transcoder
.pool({ attempts: 3, backoffMs: 100, retryOn: (response) => response.status >= 500 })
.fetch("/transcode", { method: "POST", body });fetch accepts a path string (resolved against the container) or a full
Request, and proxies WebSocket upgrades, so a real-time stream to a container
works the same as any other fetch.
Cold-start retry
The first request to a sleeping or never-started instance can land while
Cloudflare is still provisioning it, surfacing as a 503 "no Container
instance available", a 500 "Failed to start container", a 429, or a thrown
"not listening" error. .get() and .any() absorb that race: they retry the
same instance a few times with exponential backoff (default 3 attempts,
500 ms base, 30 s ceiling) so transient provisioning blips don't reach your
handler. Only those platform provisioning signals retry; an honest application
5xx is returned straight through.
// Tune or disable per call (attempts: 1 sends exactly once).
await ctx.containers.transcoder.get(jobId, { attempts: 5, backoffMs: 250 }).fetch("/transcode");A pre-built Request is sent once and never retried, because its body may
be a one-shot stream that can't be replayed. Pass a path string (the common case) to
opt into the retry, or set attempts: 1 to send a Request without it. (This
is distinct from .pool(), which retries on a fresh instance for stateless
load-balancing; cold-start retry sticks to the one instance you named.)
Managing a named instance
The .get(name) handle also exposes lifecycle control, for the per-entity
pattern where you manage an instance rather than wait for sleepAfter:
const sandbox = ctx.containers.codeRunner.get(userId);
await sandbox.start({ envVars: { SESSION: userId } }); // explicit start (optional per-instance env)
const state = await sandbox.getState(); // inspect runtime state
await sandbox.stop(); // stop (can restart on next request)
await sandbox.destroy(); // tear down and discard the ephemeral disk.any() and .pool() return fetch-only handles; lifecycle control is for
named instances you own.
Cloudflare has no built-in autoscaling yet: pools are a fixed size and .any()/.pool() pick uniformly, ignoring location. .pool() is the recommended
call for stateless work until then; for batch work, drive instances from a scheduler cron or runAfter.
Per-instance images and snapshots
With the default scheduling policy, one image and one instance size in
wrangler.jsonc serve every instance, and Cloudflare rolls changes out on
deploy. schedulingPolicy: "durable_object" (Cloudflare public beta) moves
that choice into each instance's start() instead. It suits sandboxes and
agent environments, where each instance may need a different image or size.
// lunora/containers.ts
export const agentComputer = defineContainer({
schedulingPolicy: "durable_object",
images: {
base: "./containers/base", // built and uploaded by wrangler deploy
gpu: { registry: "registry.cloudflare.com/<account>/gpu@sha256:<digest>" },
},
image: "base", // default for a start that names none, including the implicit one a fetch triggers
instanceType: "lite",
defaultPort: 8080,
});const computer = ctx.containers.agentComputer.get(taskId);
await computer.start({ image: "gpu", instanceType: "standard-2" });
const saved = await computer.snapshot({ name: "after-setup" }); // plain data, store it anywhere
await computer.stop();
await ctx.containers.agentComputer.get(otherTaskId).start({ snapshot: saved }); // restore the filesystemimageis a key ofimagesor a Cloudflare-managed image such as"cloudflare/debian-trixie". The choice (andinstanceType) is persisted likeenvVars, so the restart after a sleep boots the same image. A start that changes it on a running instance is rejected withCONFLICT.- A named
{ registry }image must be digest-pinned in the Cloudflare registry. Push Docker Hub or ECR images there first. snapshot()saves the writable filesystem, not memory or processes. A restored container runs its entrypoint again. A snapshot is tied to the image it came from and expires 30 days after its last restore.start({ snapshot })applies to that one start.- The policy takes no
maxInstancesorrollout: running instances count against the account limit, and each keeps its startup image across deploys. The policy is immutable. Switching means a new container application, which replaces every instance.
Needs wrangler 4.131.0 or newer: older releases reject scheduling_policy: "durable_object" and the images map at wrangler deploy. Only Cloudflare
supports this. Codegen refuses schedulingPolicy: "durable_object" with a platform_unsupported_feature diagnostic on the Node and celld targets.
Sandbox helpers: files, backups and bucket mounts
Sandbox SDK 1.0 moved container
control into the app's own Durable Object, which is what a LunoraContainer
already is. Its three helpers, Files, DirectoryBackup and S3Mount, are
built in. Opt a container in with sandbox: true. The helpers run Cloudflare's
sandbox-shim binary inside the container, which must be at
/usr/local/bin/sandbox-shim. Cloudflare publishes it as the image
cloudflare/sandbox, which contains nothing but the shim. Copy it into your own
image, pinned to the version matching @cloudflare/sandbox:
FROM debian:trixie-slim
COPY --from=docker.io/cloudflare/sandbox:1.0.0 /usr/local/bin/sandbox-shim /usr/local/bin/sandbox-shimexport const workspace = defineContainer({
image: "./containers/workspace",
sandbox: true,
backups: { bucket: "WORKSPACE_BACKUPS", prefix: "workspaces/" },
});sandbox must be a true/false literal: codegen reads it to decide what the
generated container class extends and which Worker entrypoints to export. An app
that never opts in never loads @cloudflare/sandbox.
Files
ctx.containers.workspace.get(id).files reads and writes the instance's own
disk, the same disk exec and spawn run against. It starts the container if
it is stopped.
const box = ctx.containers.workspace.get(userId);
await box.files.mkdir("/workspace/src", { recursive: true });
await box.files.writeFile("/workspace/src/main.ts", source);
const entries = await box.files.readDirectory("/workspace/src");
const text = await (await box.files.readFile("/workspace/out.log")).text();readFile streams: the response body applies backpressure to the read inside
the container, so a large file never sits in Worker memory. writeFile takes a
string, bytes or a ReadableStream. A filesystem failure rejects with a
LunoraError whose data.errno is the Linux code: ENOENT is NOT_FOUND,
EACCES/EPERM are FORBIDDEN, and EEXIST/ENOTEMPTY/EBUSY are
CONFLICT.
Backups
backup(directory) saves a directory to the R2 bucket named in backups and
returns a plain record. Store it, and pass it to restore later, on this
instance or another, even one running a newer image:
const record = await box.backup("/workspace", { exclude: ["node_modules"], gitignore: true });
await ctx.runMutation(internal.workspaces.saveBackup, { userId, record });
// Later, on a fresh instance:
await ctx.containers.workspace.get(userId).restore(record);The container never holds bucket credentials. Each backup or restore goes
through the DirectoryBackupGateway Worker entrypoint, which grants the
container access to exactly one object for the length of that one operation, and
every restore checks the object's SHA-256. A restore swaps the directory in only
after the download is verified, so a rejected restore leaves the target as it
was. deleteBackup(record) removes the object.
snapshot() (above) and backup() differ in scope. A snapshot captures the
whole filesystem and restores only into the same container. A backup captures
one directory into your own bucket and restores across images.
Bucket mounts
mount() attaches an S3-compatible bucket (R2, S3, GCS) at a path inside the
container. Credentials are passed by secret name: the Durable Object reads
them from the Worker env, and the S3Gateway entrypoint signs every storage
request, so the keys never enter the container.
await box.mount({
path: "/mnt/datasets",
endpoint: `https://${accountId}.r2.cloudflarestorage.com`,
region: "auto",
bucket: "datasets",
access: "read-only",
credentials: { accessKeyIdSecret: "R2_ACCESS_KEY_ID", secretAccessKeySecret: "R2_SECRET_ACCESS_KEY" },
});For R2 this needs an R2 API token's S3 keys; an R2 bucket binding is not used.
A mounted path is not a POSIX filesystem, so don't use it for locks or atomic
renames. inspectMount(path) reports a mount's state, and unmount(path)
removes it.
A bucket mount cannot be combined with an egress policy (allowedHosts, deniedHosts, interceptHttps, outbound handlers or the runtime egress
controls) on the same container: the policy's catch-all would take the mount's storage traffic. mount() refuses with BAD_REQUEST in that case. Backups
do work with an egress policy, because their route is registered before the policy's catch-all at every start.
Codegen refuses sandbox: true on the Node and celld targets with a
platform_unsupported_feature diagnostic (containerSandboxTools).
Running a command: exec
fetch covers containers that serve an HTTP API. When the container is a
runner — a build box, a test sandbox, a job worker — what you want is a
command and its exit code. exec is that contract:
const result = await ctx.containers.codeRunner.get(userId).exec("pnpm", {
args: ["install", "--frozen-lockfile"],
cwd: "/app",
env: { CI: "1" },
timeoutMs: 120_000,
});
if (result.code !== 0) {
throw new Error(`install failed:\n${result.stderr}`);
}It returns { code, stdout, stderr }, and it is available on every handle —
.get(), .any(), .pool() — and through .port(n), inheriting the same
cold-start retry and routing as fetch.
A non-zero exit code is not an error. A command that ran and failed is a
result, so you get a code to branch on. exec throws only when the command
could not be run: the container answered non-2xx (usually "no exec route"),
the body was not JSON, or it carried no numeric code. That distinction is the
point of the contract — the convention it replaced read the response body back
as output, so a container with no exec route handed you its 404 page as though
the command had succeeded.
Output is capped, and so is the clock
stdout and stderr come back as one JSON document, which has to be held whole
to be parsed. A build box that emits 150MB would take the isolate — and every
other request sharing it — down with it, so the response body is capped at
1MB by default and a runner that overruns it fails the call rather than the
shard. Raise it per call with maxOutputBytes when a command legitimately
produces more, or have the runner cap its own output.
timeoutMs covers the whole call, request and response body. That is not a
detail: fetch resolves on headers, so a runner that answers 200 when the
command starts and then stalls would otherwise sit past the deadline.
An exec is not cancelled when its deadline fires. exec reaches the container Durable Object over an RPC, and an RPC argument cannot carry an
AbortSignal, so the deadline is enforced in your worker: it frees your handler on time, while the call it abandons keeps running and the command keeps
going inside the container. The client always sends timeoutMs in the request body — have the runner enforce it, since the runner is the only place that
can stop the command.
On .pool(), each attempt re-picks an instance, exactly like a pooled fetch — so a pooled exec must be safe to run more than once. Use .get(name) for
anything that isn't idempotent. A pooled exec retries only the platform's cold-start transients (which mean the request never reached the container), not
the any-5xx default a pooled fetch uses — a runner that ran the command and then failed must not have it run again.
The container side of the contract
On Cloudflare the container needs nothing: the container Durable Object
runs the command through the runtime's native ctx.container.exec(), unshelled,
reads stdout and stderr under the same maxOutputBytes cap (killing the process
if it overruns, or when timeoutMs fires), and answers the contract itself. The
container still has to start the usual way, so exec needs a defaultPort (or a
.port(n) handle), as before. The command gets the container's start env
(env, secrets, secretsStore values, or a start({ envVars }) override),
with the call's own env over it. maxOutputBytes caps stdout and stderr each.
On Cloudflare this path skips an image-side /__lunora/exec runner entirely, so any allowlist, user switch or audit logging that runner applied no longer
runs. Put those restrictions in the image itself (a non-root default user, a restricted PATH), and gate model-chosen commands at the caller.
Where the runtime has no native exec, the runner inside the container serves the
route instead. Accept a POST at /__lunora/exec taking
{ "command": "pnpm", "args": ["install"], "cwd": "/app", "env": { "CI": "1" }, "maxOutputBytes": 1000000, "timeoutMs": 120000 }and answer { "code": 0, "stdout": "…", "stderr": "…" } as JSON. cwd, env
and timeoutMs are omitted from the body entirely when unset, so a runner never
has to special-case an explicit undefined. Absent stdout/stderr in the
reply default to ""; an absent code is an error, because "the command ran"
and "the command's result is unknown" must not look the same.
Run the command directly — do not concatenate args into a shell string, which
would make every argument an injection point.
/__lunora/* is reserved for Lunora's own container routes. handle.fetch
refuses any path resolving into it — in any letter case, percent-encoded (once
or more), or with ;params on the segment — so serve your application's routes
elsewhere and reach exec through exec. The container Durable Object enforces
it a second time on every HTTP entry: its fetch and its containerFetch RPC
answer 403 for a path touching /__lunora/* (any segment, same spellings),
whatever headers the request carries. exec reaches the route over a separate
RPC (lunoraExec), which only code holding the container binding can call, so
forwarding an inbound request as-is (env.CONTAINER_X.get(id).fetch(request))
cannot open it either.
exec runs whatever it is given. When a model chooses the command — an @lunora/agent containerTool, say — gate it: that tool's default approval policy
pauses for a human on exec and on any non-idempotent fetch. A fetch cannot be used to reach the exec route around that gate — /__lunora/* is
reserved and handle.fetch refuses it — but restrict what can run inside the container too, rather than relying on the caller.
Streaming a process: spawn
exec buffers: it returns once the command has exited, with at most 1MB of
output. spawn is the streaming counterpart for long builds, live logs and
interactive programs. It needs the runtime's native exec, and is available on
named instances (.get(name)), since a process belongs to one container:
const process = await ctx.containers.codeRunner.get(userId).spawn("pnpm", {
args: ["test", "--reporter=dot"],
cwd: "/workspace",
timeoutMs: 600_000,
});
for await (const chunk of process.stdout!) {
// Forward to a client, write to R2, …
}
const code = await process.exitCode;The result carries stdout/stderr streams, exitCode (a promise), kill()
and the process pid. The container finishes a process's output, and settles
its exit code, only while every output stream is being drained. Lunora drains
both for you into a 1MB buffer per stream, so you can read them in any order,
or await exitCode before reading, as long as the stream you are not reading
stays under that size. For a noisier process, read both concurrently. Pass stdin: true to get a writable stdin, or a
ReadableStream to pipe in as the whole input. Pass pty: { cols, rows } to
run on a pseudo-terminal: stderr is merged into stdout, and resize() works.
An AbortSignal in signal kills the process.
The container is started if it is stopped, and kept awake until the process
exits. sleepAfter does not stop it underneath a running process. A spawned
process has no default deadline, so set timeoutMs for anything that should not
run indefinitely.
Browser terminal
terminal(request) answers a WebSocket upgrade with a shell on a PTY, bridged to
the socket. Browser routes have no ctx.containers, so reach the instance with
getContainer(env, exportName, name):
import { getContainer } from "@lunora/container";
app.get("/terminal", async (c) => {
const userId = await requireUser(c); // whoever holds the socket has a shell
return getContainer(c.env, "workspace", userId).terminal(c.req.raw, { command: "bash", cwd: "/workspace" });
});The wire protocol is what xterm.js speaks. Binary frames are keystrokes, and
a text frame {"cols":120,"rows":40} resizes the terminal. Any other text
frame counts as keystrokes too. Output arrives as
binary frames, the socket closes with 1000 when the shell exits, and closing
the socket kills the shell. The initial size comes from the request's
?cols=&rows= search params, else 80×24.
import { FitAddon } from "@xterm/addon-fit";
import { Terminal } from "@xterm/xterm";
const term = new Terminal();
const fit = new FitAddon();
term.loadAddon(fit);
term.open(element);
fit.fit();
const socket = new WebSocket(`wss://${location.host}/terminal?cols=${term.cols}&rows=${term.rows}`);
socket.binaryType = "arraybuffer";
socket.onmessage = (event) => term.write(new Uint8Array(event.data));
term.onData((data) => socket.send(new TextEncoder().encode(data)));
term.onResize(({ cols, rows }) => socket.send(JSON.stringify({ cols, rows })));The bridge runs where terminal() is called, on top of spawn. No WebSocket
crosses into the container Durable Object, which keeps its reserved
/__lunora/* routes closed to request content.
Previews
A dev server in the container is already reachable through fetch: route a
browser request to its port with .port(n), WebSocket upgrades (HMR) included.
app.all("/previews/:id/*", async (c) => {
await requireOwner(c, c.req.param("id"));
const path = c.req.path.replace(`/previews/${c.req.param("id")}`, "") || "/";
const url = new URL(path + new URL(c.req.url).search, "http://container");
return getContainer(c.env, "workspace", c.req.param("id")).port(5173).fetch(new Request(url, c.req.raw));
});The server sees Host: container, and a link that starts with / leaves the
/previews/<id> prefix. Use relative links, or serve each preview on its own
hostname: a wildcard route (*.preview.example.com/*) where the Worker reads
the instance id from the subdomain, so the app runs at / and absolute links
work. To share a preview, mint an expiring token, store it, and check it in the
route before forwarding. Authenticate whoever mints the token, because a token
is access to the dev server.
Multi-port containers
A container that listens on more than one port (an app port plus an admin port,
say) declares every one with requiredPorts; start-up waits for all of them to
be listening. defaultPort is the target when a request doesn't pick a port;
route a single request elsewhere with .port(n), which composes with .get(),
.any(), and .pool():
export const app = defineContainer({
image: "./containers/app",
defaultPort: 8080,
requiredPorts: [8080, 9090], // app + admin
});// in an action:
await ctx.containers.app.get(tenantId).fetch("/work"); // → 8080 (defaultPort)
await ctx.containers.app.get(tenantId).port(9090).fetch("/admin"); // → 9090Egress firewall
By default a container may open any outbound connection (enableInternet: true),
and egress is billed per GB. Constrain it by pairing enableInternet: false
with an allowedHosts allow-list, or layer a deniedHosts deny-list that
overrides everything, including enableInternet: true and allowedHosts. Glob
patterns like *.stripe.com are supported.
export const fetcher = defineContainer({
image: "./containers/fetcher",
enableInternet: false,
allowedHosts: ["*.stripe.com", "api.github.com"],
deniedHosts: ["*.evil.com"],
interceptHttps: true, // extend the lists to TLS traffic (image must trust the Cloudflare CA)
});interceptHttps extends the allow/deny lists to HTTPS connections, not just
plain HTTP; it requires the image to trust the Cloudflare CA at
/etc/cloudflare/certs/cloudflare-containers-ca.crt. The interception path runs
through the ContainerProxy worker entrypoint, which codegen re-exports from the
generated container file automatically.
Tighten or relax a single running instance at runtime through its named handle.
egress.allow / deny add one host, removeAllowed / removeDenied drop one,
and setAllowed / setDenied replace a whole list:
await ctx.containers.fetcher.get(tenantId).egress.allow("hooks.slack.com");
await ctx.containers.fetcher.get(tenantId).egress.setDenied(["*.evil.com", "*.tracking.example"]);For advanced egress rewriting in worker code (inject auth, route, or mock a
container's outbound calls), @lunora/container/do re-exports Cloudflare's
custom outbound-handler types (OutboundHandler, OutboundHandlers,
outboundParams); wire them onto a hand-authored LunoraContainer subclass.
Readiness gating
The platform health check waits for an open port (defaultPort) and, if
pingEndpoint is set, for an HTTP probe on that path on the same container.
Neither means the app is ready. readyOn adds
application-level probes that gate request proxying: a ctx.containers.<name>
fetch holds until every probe responds with its expected status, so callers
never hit a container still applying migrations or warming caches.
export const api = defineContainer({
image: "./containers/api",
defaultPort: 8080,
readyOn: [
{ path: "/ready" }, // expect 200 on defaultPort
{ path: "/live", port: 9090, status: 204 }, // own port + expected status
],
});Each probe declares a path (a leading slash is optional), an optional port
(defaults to defaultPort), and an optional status (defaults to 200).
Probes are declarative data (no handler functions), so codegen and the
config layer read them without evaluating code. At start they run in parallel,
poll the container's TCP port directly, and fail the start if any probe doesn't
go ready within the readiness budget.
Hard timeout
sleepAfter caps idle time; hardTimeout caps total lifetime: a
runaway-cost backstop that fires regardless of activity, measured from start. It
uses the same grammar as sleepAfter ("30s", "5m", "1h", or a plain
number of seconds):
export const job = defineContainer({
image: "./containers/job",
sleepAfter: "5m", // sleep after 5 min idle
hardTimeout: "1h", // …stopped after an hour, busy or not
});When it elapses, the generated class's onHardTimeoutExpired hook runs; the
default action is stop(), which sends SIGTERM and does not escalate. A
container that traps or ignores SIGTERM therefore keeps running past its cap —
if yours might, override onHardTimeoutExpired and follow the stop with a
destroy() after whatever grace period your workload needs. The timer is armed through the container's own
scheduler and stamped with a run generation, so a stale timer left over from a
previous (slept or crashed) run never kills a fresh one. Advanced apps that
hand-author their container class can override onHardTimeoutExpired (the base
class lives in @lunora/container/do) to drain or checkpoint before stopping.
Calling Lunora from inside a container
Container code calls back into your app's functions with the bridge client (any JS runtime: Node, Bun, Deno) over the Worker's HTTP RPC endpoint, so a container reads/writes app state through the same queries/mutations the browser uses, instead of reaching into a database directly.
import { createContainerBridge } from "@lunora/container/bridge";
const lunora = createContainerBridge({ baseUrl: process.env.LUNORA_URL, token: process.env.LUNORA_TOKEN });
const pending = await lunora.query("jobs:listPending", { limit: 10 });
await lunora.mutation("jobs:markDone", { id: pending[0].id });For full type-safety, pass a generated api reference to run(); args and
result are inferred from it:
import { api } from "../lunora/_generated/api";
const job = await lunora.run(api.jobs.next, { queue: "transcode" }); // typed args + resultThe token is a bearer your Worker's resolveIdentity recognizes. Forward it
as a container secret, never bake it into the image. Non-JS containers can
POST /_lunora/rpc with { functionPath, args } directly (same contract).
Securing the bridge
The bridge sends its token as Authorization: Bearer <token>. Your Worker's
resolveIdentity is what validates it and maps it to the identity the called
functions run as. Without that check, anyone who reaches /_lunora/rpc runs as
whatever identity you return. Validate the bearer against a Worker secret and
return null for anything that doesn't match (an unrecognised request runs
anonymously and fails your functions' own authorization checks):
// in your worker options (createWorker / withLunora)
resolveIdentity: (request, env) => {
const header = request.headers.get("authorization");
const token = header?.startsWith("Bearer ") ? header.slice("Bearer ".length) : undefined;
// Compare against a Worker secret you also forward to the container.
if (!token || token !== env.LUNORA_CONTAINER_TOKEN) {
return null; // anonymous — function-level authorization still applies
}
return { userId: "container:transcoder" }; // identity the bridge calls run as
},Generate the token once, set it with wrangler secret put LUNORA_CONTAINER_TOKEN
for production (and in .dev.vars locally), and pass the same value to the
container as a secret so it can send it back. Never bake it into the image.
Building with Railpack (no Dockerfile)
Point image at a { build } source directory and Lunora builds the image
with Railpack, Railway's Dockerfile-less builder,
instead of you writing a Dockerfile:
export const transcoder = defineContainer({
image: { build: "./services/transcoder" }, // Railpack detects the stack and builds it
defaultPort: 8080,
});On lunora deploy, each { build } container is built with Railpack and
pushed to the Cloudflare Registry (under a deterministic lunora-<name>:build
tag) before the Worker deploys. This needs the railpack CLI and a running
BuildKit instance reachable via BUILDKIT_HOST, e.g.:
docker run --rm --privileged -d --name buildkit moby/buildkit
export BUILDKIT_HOST=docker-container://buildkitlunora deploy preflights both and stops with setup directions if either is
missing. The Dockerfile path remains the zero-extra-deps default; Railpack is
opt-in.
Lifecycle logs
The generated container classes emit a structured event on start, stop,
and error, tagged with the container name and per-instance id (the
container's CLOUDFLARE_DURABLE_OBJECT_ID). lunora dev surfaces them inline
with your ctx.log and RPC lines:
[lunora] container:transcoder#a1b2c3d4 start
[lunora] container:transcoder#a1b2c3d4 error exited unexpectedlyThe same events also appear in the Studio Logs panel: each lifecycle
transition is best-effort pushed into the root shard's log buffer under the
container:<name> source, so a crash-looping container shows up next to your
function logs without leaving the dashboard. The push is best-effort by design:
if it can't reach the shard, the lunora dev terminal above remains the source
of truth.
Advisor lints
Two advisor lints flag container cost/security footguns in the Studio Advisors table and at codegen time:
container_oversized_instance: a largestandard-3/standard-4(or large custom) instance, whose provisioned memory + disk are billed while any instance runs.container_public_internet:enableInternetleft at the default (enabled); set it explicitly (falseto close egress,trueto opt in).
Configuration
| Field | Purpose |
|---|---|
image | Local Dockerfile path/directory, { registry } for a pre-built image, or { build } (Railpack). Required. |
defaultPort | Port the container listens on; fetch targets it. Must be EXPOSEd for local dev. |
requiredPorts | Every port start-up must wait for (multi-port); route a request to one with .port(n). |
instanceType | lite | basic | standard-1..4, or a custom { vcpu, memoryMib, diskMb }. |
maxInstances | Cap on concurrently running instances (and the default .any() pool size). |
sleepAfter | Idle timeout before an instance sleeps, e.g. "5m", "30s", or seconds. Default "10m". |
hardTimeout | Hard cap on total lifetime from start, ignoring activity. Same grammar as sleepAfter. |
readyOn | Application-level readiness probes that gate request proxying until the app reports ready. |
pingEndpoint | HTTP path the platform polls to decide an instance is healthy. Defaults to upstream's "ping". |
env | Static environment variables passed to the container on every start (runtime). For secrets use secrets. |
secrets | Names of Worker secrets forwarded into the container env at start. |
buildArgs | Build-time args for a built image (wrangler image_vars, like --build-arg); ignored for { registry }. |
entrypoint | Override the image's ENTRYPOINT/CMD for every start. |
enableInternet | Whether the container may open outbound connections. Default true; egress is billed. |
allowedHosts | Egress allow-list (globs): hosts reachable even with enableInternet: false. |
deniedHosts | Egress deny-list (globs): overrides everything, including allowedHosts. |
interceptHttps | Extend the egress lists to HTTPS too (image must trust the Cloudflare CA). Default false. |
labels | Key-value metadata attached to every instance for metrics/observability. |
name | Override the wrangler containers[].name identifier. |
rollout | Rolling-deploy tuning: { stepPercentage, gracePeriodSeconds }. |
schedulingPolicy | "durable_object" picks image and size per start(); see per-instance images. |
images | Named images under schedulingPolicy: "durable_object"; image then names the default among them. |
Secrets and environment
env holds static values; secrets names Worker secrets, managed exactly like
every other Lunora secret. List them in .dev.vars for local dev and
wrangler secret put for production (see Deployment).
A declared secret that isn't set fails fast at container start with a directed
error, rather than launching the container without it.
export const worker = defineContainer({
image: "./containers/worker",
env: { LOG_LEVEL: "info" },
secrets: ["STRIPE_API_KEY"], // resolved from env.STRIPE_API_KEY at start
});To pull from a Cloudflare Secrets Store
binding instead of a plain Worker secret, map container env-var name → Secrets
Store binding name with secretsStore. Each binding is resolved with its async
.get() the first time the instance starts (memoised thereafter) and injected as
that env var:
export const worker = defineContainer({
image: "./containers/worker",
secretsStore: { STRIPE_KEY: "STRIPE_SECRET" }, // env.STRIPE_SECRET.get() → STRIPE_KEY inside the container
});A name that collides with env/secrets, or a binding missing from the Worker
env, is rejected (the collision at authoring time, the missing binding at start):
the same fail-closed stance as secrets.
A per-instance start({ envVars }) replaces the env set wholesale (the
Secrets Store is not read) and is persisted for that instance: every later
start uses it — an explicit start(), and the implicit restart a fetch or
exec triggers after the container slept, crashed or hit hardTimeout — until
you destroy() the instance. That makes start({ envVars: {} }) (or with only
what the job needs) the way to run an untrusted or model-driven exec sandbox
without the credentials the definition declares, and keep it that way across
restarts. The override cannot change while the container is running or still
starting: a start({ envVars }) that differs from its env is rejected with
CONFLICT in either case, rather than silently ignored — stop() the instance
(and let any start in flight finish) first. Pass every variable the container needs.
Both env and secrets are runtime values, available when the container
starts. For build-time values (docker build --build-arg, exposed to the
Dockerfile as ARG, wrangler's image_vars), use buildArgs. They only apply
to an image Lunora builds (a Dockerfile or Railpack { build } source) and are
ignored for a pre-built { registry } image:
export const worker = defineContainer({
image: "./containers/worker",
buildArgs: { NODE_VERSION: "22", BUILD_TARGET: "production" }, // → docker --build-arg
});Dockerfile
The scaffolded Dockerfile encodes the platform's requirements:
# linux/amd64 is required — keep it explicit so Apple Silicon builds don't
# produce an arm64 image that fails at deploy.
FROM --platform=linux/amd64 node:22-slim
WORKDIR /app
COPY . .
# Local dev needs the listening port EXPOSEd (production exposes all ports).
EXPOSE 8080
# Exec form (not shell form) so SIGTERM reaches the process — Cloudflare sends
# SIGTERM on rollouts and gives 15 minutes before SIGKILL.
ENTRYPOINT ["node", "server.mjs"]Disk is ephemeral: every (re)start gives a fresh filesystem. For
persistence, write to @lunora/storage (R2).
Pushing to a Cloudflare Artifacts repo
A container's disk is ephemeral. When what it produces is a tree of files you
want history and diffs for (an agent's workspace, a generated site, a build
output), push it to a Cloudflare Artifacts
repo. The Artifacts binding can't write files, so the push comes from git inside
the container, authenticated with short-lived tokens the action mints.
The build runs code the container fetched (pnpm install, pnpm run build, and
every dependency's scripts). The recipe keeps tokens out of the build's
environment and out of the checkout, which stops a careless build from picking
one up. It does not isolate the push from a build that is trying to: see
the rules below before you run code you don't trust.
import { internalAction } from "@/lunora/_generated/server";
import { v } from "@lunora/values";
// Single-quoted: `exec` runs the command unshelled, so `sh -c` expands the
// variables inside the container. `core.hooksPath=/dev/null` and `--no-verify`
// keep hooks the build may have written from running.
const CHECKOUT =
'rm -rf /work && if [ "$EMPTY" = 1 ]; then git init -b "$BRANCH" /work; else git -c core.hooksPath=/dev/null clone --branch "$BRANCH" "$REMOTE" /work; fi';
const BUILD = "pnpm install --frozen-lockfile && pnpm run build";
// Prints `unchanged` instead of committing when the build changed nothing.
const COMMIT =
'git -C /work add -A && if git -C /work diff --cached --quiet; then echo unchanged; else git -C /work -c core.hooksPath=/dev/null -c user.name=lunora -c user.email=bot@example.com commit --no-verify -q -m "$MESSAGE"; fi';
const PUSH = 'git -C /work -c core.hooksPath=/dev/null push "$REMOTE" "HEAD:$BRANCH"';
/**
* Hand a token to one `git` process as an `http.<remote>.extraHeader` through
* git's `GIT_CONFIG_*` environment: it is in neither the command line nor
* `/work/.git/config`, and it is sent only to `remote`, so a `url.*.insteadOf`
* rewrite can't redirect it to another host. The `?expires=` suffix is not part
* of the secret.
*/
const gitAuthEnv = (remote: string, plaintext: string): Record<string, string> => {
return {
GIT_CONFIG_COUNT: "1",
GIT_CONFIG_KEY_0: `http.${remote}.extraHeader`,
GIT_CONFIG_VALUE_0: `Authorization: Basic ${btoa(`x:${plaintext.split("?expires=")[0] ?? ""}`)}`,
};
};
// Internal: only your own code decides who may push to which repo.
export const buildAndPush = internalAction.input({ repo: v.string(), sessionId: v.string() }).action(async ({ ctx, args }) => {
const { defaultBranch, remote } = await ctx.artifacts.info(args.repo);
// A freshly created repo has no commit, so no branch to clone yet.
const empty = (await ctx.artifacts.withRepo(args.repo, (repo) => repo.log({ limit: 1 }))).length === 0;
const box = ctx.containers.workspace.get(args.sessionId);
const base = { BRANCH: defaultBranch, MESSAGE: `build ${args.sessionId}`, REMOTE: remote };
const live = new Set<string>();
const mint = async (scope: "read" | "write") => {
const token = await ctx.artifacts.withRepo(args.repo, (repo) => repo.createToken(scope, 600));
live.add(token.id);
return token;
};
// A failed revoke must not mask the result; the token still expires on its own.
const revoke = async (id: string): Promise<void> => {
live.delete(id);
try {
await ctx.artifacts.withRepo(args.repo, (repo) => repo.revokeToken(id));
} catch (error: unknown) {
ctx.log.warn("artifacts token revoke failed", { error, tokenId: id });
}
};
try {
const readToken = empty ? undefined : await mint("read");
const checkout = await box.exec("sh", {
args: ["-c", CHECKOUT],
env: { ...base, EMPTY: empty ? "1" : "0", ...(readToken === undefined ? {} : gitAuthEnv(remote, readToken.plaintext)) },
timeoutMs: 120_000,
});
// The read token's job is done before any of the repo's code runs.
if (readToken !== undefined) {
await revoke(readToken.id);
}
if (checkout.code !== 0) {
throw new Error(`checkout failed (exit ${String(checkout.code)})`);
}
// No token in this env. Install and build logs routinely pass exec's 1MB
// output default, so raise the cap (or have the build write less).
const build = await box.exec("sh", { args: ["-c", BUILD], cwd: "/work", maxOutputBytes: 16 * 1024 * 1024, timeoutMs: 600_000 });
if (build.code !== 0) {
throw new Error(`build failed (exit ${String(build.code)})`);
}
const commit = await box.exec("sh", { args: ["-c", COMMIT], env: base, timeoutMs: 60_000 });
if (commit.code !== 0) {
throw new Error(`commit failed (exit ${String(commit.code)})`);
}
if (commit.stdout.trim() === "unchanged") {
return { changed: false, pushed: false };
}
// The write token exists only now, and only in the push's environment.
const writeToken = await mint("write");
const push = await box.exec("sh", { args: ["-c", PUSH], env: { ...base, ...gitAuthEnv(remote, writeToken.plaintext) }, timeoutMs: 120_000 });
return { changed: true, pushed: push.code === 0 };
} finally {
// Revoke by id whatever happened — the write token, and the read token if the checkout threw.
for (const id of [...live]) {
await revoke(id);
}
}
});After the push, ctx.artifacts.withRepo(repo, (r) => r.readFile({ ref: defaultBranch, path }))
reads the pushed files back from any action, with no container running. The
checkout starts from rm -rf /work, so calling the action again on the same
session starts clean.
Keeping the token contained
-
Mint late, and only for the step that needs it. The clone gets a read token that is revoked as soon as the checkout ends. The write token is minted after the build and only exists in the push's environment. The install, build and commit run with neither. Hooks are disabled on every
gitstep, so a.git/hooksorcore.hooksPaththe build wrote doesn't run with a token in its environment. -
Pass the token through
GIT_CONFIG_*, never throughargs, a URL or a log line. It never lands in/work/.git/configor in a process's arguments, and the header is scoped to the repo's remote. Don't print it, don't return it from the action, and don't leaveset -xon in a script that uses it. -
If the build runs code you don't trust, push from a separate container session. Everything above shares one container with the build, and a build that wants the token can still reach the push:
- a process it leaves running reads the push's environment from
/proc/<pid>/environ; - it writes
~/.gitconfigor/etc/gitconfig(a proxy, a credential helper, its ownhttp.*.extraHeader), which the push's-cflags don't override; - it rewrites
/work/.git/config, for example acore.fsmonitorcommand; - it puts its own
gitearlier onPATH.
So for untrusted code, build in one session and move only the output out: with
sandbox: true,backup()the output directory (never its.git) andrestore()it into a fresh session the build never ran in (see Backups), or copy it through R2 with@lunora/storage. Clone, commit and push from that fresh session. The hardening above is for builds you trust; it is not a sandbox. - a process it leaves running reads the push's environment from
-
The image needs
git.node:22-slimdoesn't include it, so addRUN apt-get update && apt-get install -y --no-install-recommends git ca-certificatesto the Dockerfile.
This is a different trust model from saving files to R2 with
@lunora/storage, where the Worker holds the
credentials and the container never sees them. Here a repo-scoped write
credential lives inside the container for the length of the push. Choose
Artifacts when you want Git history, diffs and forks of the output, or want
other Git clients to clone it. Choose R2 when you want an opaque snapshot
restored as is.
For repos too large to clone on every start, mount them with ArtifactFS inside the image. That is an image-level choice Lunora doesn't wrap.
Local development
lunora dev (and vite dev) build and run containers locally through the
Cloudflare Vite plugin, so a Docker-compatible engine (Docker Desktop, Colima)
must be running; Lunora warns up front if it isn't. A few things differ from
production:
- Container code is not hot-reloaded. Edit the Worker and it reloads; press
rto rebuild the container. - Ports must be
EXPOSEd in the Dockerfile (production exposes all ports). vite devcan't pull from the Cloudflare Registry. Use a local Dockerfile,FROM-ing a registry image if you need one.
Testing
ctx.containers has a Docker-free test double, so action tests stay fast and
deterministic. createContainerTestContext maps each container to a fetch
handler that plays the container:
import { createContainerTestContext } from "@lunora/container";
const containers = createContainerTestContext({
transcoder: async (request, { name }) => new Response(JSON.stringify({ ok: true, name })),
});
// Inject `containers` as ctx.containers when exercising the action.A named instance from the double also has files, backed by an in-memory disk
per instance with Linux-shaped errors (ENOENT, EEXIST, ENOTEMPTY), so a
handler that writes and reads container files can be tested without Docker.
spawn, backup/restore and mount reject with NOT_IMPLEMENTED: there is
no process or bucket behind the double, so stub those handle methods in the test.
Run the real container behind lunora dev for integration tests; keep those in
a Docker-enabled CI job, separate from the default unit-test matrix.
CLI
Manage images and instances with lunora containers, thin wrappers over
wrangler containers with a Docker preflight on the build/push paths:
lunora containers build ./containers/transcoder --tag transcoder:v1 --push
lunora containers push transcoder:v1
lunora containers images list
lunora containers images delete transcoder:v1Splitting build/push from lunora deploy lets CI build the image in one job and
deploy in another. lunora deploy itself runs a Docker preflight and stops with
an actionable message when a Dockerfile-built container is declared but no engine
is available.
Stability and how a surface would be removed
The package is Experimental: it sits outside the Lunora 1.0 SemVer promise,
and the graduation bar
is what moves it out. Its public surface is snapshotted in
api-snapshots/container.api.md and gated by pnpm run api:check, so every
addition, removal and signature change shows up as a reviewed diff — no
export carries @experimental individually any more, which is what makes the
signatures tracked rather than skipped.
What that commits us to once the package graduates:
- Names are the contract. The binding name (
CONTAINER_<NAME>), the generated class name (<Name>Container), the exec route (/__lunora/exec) and the exec body keys (command,args,cwd,env,timeoutMs→code,stdout,stderr) are the names we intend to keep. A container image built against them keeps working. - Removal is deprecate-then-remove. A surface being withdrawn is first
marked deprecated in the snapshot and docs, and keeps working for a full minor
cycle with a runtime warning naming its replacement; only then is it removed,
in a major. Pre-1.0 (
alpha/next/beta) the old path is deleted in the same change as the new one, per the repo's branch policy — which is exactly what happened whenexecreplaced the ad-hoc POST-to-/execconvention. - Config keys outlive their implementation.
defineContaineroptions are additive; one that stops being meaningful is accepted and ignored with a warning rather than becoming a hard error mid-cycle. - Errors are identified by code, not prose. Every failure raises a
LunoraErrorwith a stable code; message text may be reworded to be clearer without that counting as a break, so match on the code.
Known platform limitations
Some constraints live in Cloudflare Containers itself, not in this package, and
are worth knowing before you design around them. They track open issues on
cloudflare/containers. Lunora
papers over what it can (cold-start retry, WebSocket keep-alive); the rest are
listed here.
- No autoscaling or location-aware routing. Pools are a fixed size and
.any()/.pool()pick uniformly at random, ignoring instance region. Drive batch work from a scheduler cron and prefer.pool()for stateless requests. (cloudflare/containers#226) - Disk is ephemeral. Every (re)start gives a fresh filesystem; persist to
@lunora/storage(R2). FUSE mounts, a writabletmpfs, and some low-levelnode:netsocket modes are not available in the container sandbox. (cloudflare/containers#112, #160, #67) - Egress interception is HTTP-first. The
allowedHosts/deniedHostsfirewall gates HTTP egress out of the box; HTTPS needsinterceptHttps: trueplus the Cloudflare CA trusted in the image, and raw gRPC interception isn't supported yet. (cloudflare/containers#195) - Long jobs can be terminated on rollout. A new deploy moves instances after
rollout.gracePeriodSeconds; a job longer than that may be sentSIGTERM(15 min beforeSIGKILL). UsehardTimeoutas a cost backstop and make long work resumable. (cloudflare/containers#138) - Local dev can't pull from the Cloudflare Registry.
lunora devbuilds from a local Dockerfile;FROMa registry image inside it if you need one. (cloudflare/containers#155)