Skip to content
DocspackagesDocumentation

@lunora/mcp

Model Context Protocol servers for Lunora — one exposing a deployment to AI agents, one exposing the framework's documentation.

PackagesMcp

@lunora/mcp ships two Model Context Protocol surfaces.

The deployment server (this package's main entry) exposes a deployed Lunora app to AI agents. It registers tools for introspecting a deployment (lunora_list_functions, lunora_list_tables, lunora_get_function_schema) and invoking its functions (lunora_run_query, lunora_run_mutation, lunora_run_action), each backed by @lunora/client over HTTP RPC. It needs an admin token.

The documentation server (@lunora/mcp/docs) exposes the framework's docs, so an agent writing Lunora code can look up the real API instead of inventing one. It reads published documentation only, with no credentials and no writes, which is why it can be hosted unauthenticated; Lunora runs it at https://lunora.sh/mcp.

Most users never install this package. lunora mcp install wires both servers into whichever editor they use. See AI coding agents.

pnpm add @lunora/mcp

The server is transport-agnostic. The shipped lunora-mcp binary speaks JSON-RPC over stdio (the transport MCP clients use when they spawn a process), or you can build a server with createLunoraMcpServer and connect any transport yourself.

The lunora-mcp binary

MCP clients spawn the lunora-mcp binary and talk to it over stdio. Configuration comes from the environment, so the spawn config stays a plain { command, env }:

  • LUNORA_URL (required): base URL of the deployed Worker.
  • LUNORA_ADMIN_TOKEN (required): the deployment's admin bearer, sent on every RPC. It cannot be scoped down — lunora_list_functions, lunora_list_tables and the allowlist precheck that runs before every lunora_run_* call all hit admin-gated /_lunora/admin/* routes, so a least-privilege token returns ADMIN_FORBIDDEN on the first tool call. The binary refuses to start without it rather than advertising tools that can only 403.
  • LUNORA_MCP_ALLOW_WRITES (optional): set to 1/true/yes/on to expose the mutation/action run tools. Default: read-only (writes disabled).
  • LUNORA_MCP_ALLOW_OBSERVABILITY (optional): set to 1/true/yes/on to expose the five lunora_get_* observability tools. Default: disabled. They are read-only, but what they return — production log lines, request metadata, grouped error messages — is user data that lands in the model's context and therefore at its provider, so holding the admin bearer (which every tool needs) is not by itself consent to ship it.
  • LUNORA_MCP_ALLOW_DATA_READS (optional): set to 1/true/yes/on to expose lunora_find_related. Default: disabled. It returns raw table rows read through the deployment's admin writer, so RLS policies and column masks do not apply to what it hands the model — and it returns everything reachable within depth hops of the start row, not just that row. Separate from the observability gate on purpose: log lines and table rows are different data classes, and enabling one must not silently enable the other.
  • LUNORA_MCP_ALLOW_AGENTS (optional): set to 1/true/yes/on to expose the agent tools. Default: agents disabled.
  • LUNORA_MCP_AGENTS (optional): a ;-separated list of name:description pairs selecting which agents to expose (see Expose an agent).
{
    "mcpServers": {
        "lunora": {
            "command": "lunora-mcp",
            "env": {
                "LUNORA_URL": "https://app.example.workers.dev",
                "LUNORA_ADMIN_TOKEN": "...",
            },
        },
    },
}

The binary exits non-zero if LUNORA_URL is missing or the transport fails to connect, so the spawning client surfaces a startup failure immediately.

Exposed tools

Each tool maps onto a method LunoraClient already provides. toolDefinitions(allowWrites, allowObservability, allowDataReads) returns the advertised tool surface for a given set of gates (the tiers are also exported on their own as READ_ONLY_TOOL_DEFINITIONS, OBSERVABILITY_TOOL_DEFINITIONS, ROW_READ_TOOL_DEFINITIONS and WRITE_TOOL_DEFINITIONS); callTool dispatches a call against a client.

ToolInputWhat it does
lunora_list_functionsnoneLists the deployment's public functions (queries, mutations, actions) with their kinds.
lunora_list_tablesnoneLists the deployment's .global() tables with their row counts.
lunora_get_function_schemafunctionPathReturns one function's argument descriptors and kind, so a caller can build valid args.
lunora_run_queryfunctionPath, args?, shardKey?Runs a query and returns its result. Read-only.
lunora_explain_errorcode?, message?Explains a Lunora error from the static catalog: status, title, hint, any matched solution, and a link to the reference. No deployment needed.
lunora_run_mutationfunctionPath, args?, shardKey?, confirmed?, actionDigest?, idempotencyKey?Proposes, then (on a second confirming call) runs a mutation. Writes data. Requires writes enabled.
lunora_run_actionfunctionPath, args?, shardKey?, confirmed?, actionDigest?, idempotencyKey?Proposes, then (on a second confirming call) runs an action. May call external services. Requires writes enabled.
agent_<name>prompt, threadKey?, title?Starts a durable @lunora/agent run and awaits its answer. One tool per exposed agent. Requires agents enabled.
lunora_agent_statusthreadKeyPolls a running agent by threadKey and returns its answer once finished. Requires agents enabled.
lunora_get_logslimit?, level?, shardKey?Recent ctx.log.* lines from the deployment. Requires observability enabled.
lunora_get_issueslimit?, status?, shardKey?Grouped error Issues with their fingerprints and messages. Requires observability enabled.
lunora_get_advisorieslimit?, shardKey?Advisor findings for the running deployment. Requires observability enabled.
lunora_get_query_insightslimit?, range?, shardKey?Slow/frequent query insights over a recent window. Requires observability enabled.
lunora_get_migration_statusshardKey?Applied and pending migrations for a shard. Requires observability enabled.
lunora_find_relatedtable, id, depth?, direction?, edges?, limit?, cursor?, shardKey?Follows the schema's foreign keys out of one row and returns what it is connected to. Reads through the ADMIN writer, so RLS policies and column masks do NOT apply. Requires data reads enabled.

The default surface is read-only and non-privileged: only the introspection tools, lunora_run_query and lunora_explain_error are exposed. Three separate gates hold the rest back, and each of them both hides the tools from the advertised list and refuses them at dispatch. lunora_run_mutation and lunora_run_action need LUNORA_MCP_ALLOW_WRITES (or allowWrites: true). The five lunora_get_* tools need LUNORA_MCP_ALLOW_OBSERVABILITY (or allowObservability: true) — they are read-only, but they return production log lines, request metadata and grouped error messages, all of which land in the model's context and therefore at its provider. lunora_find_related has its own gate, LUNORA_MCP_ALLOW_DATA_READS (or allowDataReads: true), deliberately NOT folded into the observability one: it returns raw table rows read through the deployment's admin writer with RLS policies and column masks bypassed, and enabling log reading for debugging must not silently also hand over every row. Those gates decide what the server may reach at all: the token must be the deployment's admin bearer (every tool reads admin-gated routes), so it cannot be scoped down to enforce anything from the credential side. Past the write gate, each individual write additionally needs its own confirmation. Guard who can reach the server process too.

The three run-tools share one base input schema:

  • functionPath (required, string): a function reference, e.g. "messages:send".
  • args (object): the arguments object passed to the function. Omitted (or null) means an empty bag, so an all-optional function runs with its defaults, and a JSON-stringified object (which models commonly emit) is parsed. Anything else — an array, a number, a boolean, a string that isn't a JSON object — is rejected with a BAD_REQUEST isError result naming the actual mistake, not coerced to {}.
  • shardKey (string): an optional shard key when the function is .shardBy()-partitioned.

The two write tools take three more fields — confirmed, actionDigest and idempotencyKey — which drive the confirmation handshake below.

Results are returned as MCP content text: the function's return value JSON, or the JSON null literal for a void-returning mutation/action. Unknown tools and thrown errors come back as isError tool results (not rejections), so the calling model sees the failure as tool output.

Write confirmation

allowWrites decides whether this server may write at all. It says nothing about whether a particular write was reviewed — and that is the gap that matters most for lunora_run_action, the tool that can send mail, charge a card, or call a third-party API. The destructiveHint annotation is a UI hint, not a gate.

So past that gate, lunora_run_mutation and lunora_run_action each take two calls. The first executes nothing and returns the proposal:

{
    "status": "action_required",
    "actionDigest": "1789129912052.0ZR2…",
    "expiresAt": "2026-09-11T14:41:52.052Z",
    "proposedAction": {
        "tool": "lunora_run_mutation",
        "kind": "mutation",
        "functionPath": "messages:send",
        "args": { "roomId": "r1", "text": "hi" },
    },
    "nextStep": "Show proposedAction to a human. To execute, call …",
}

The client renders proposedAction for a human, then calls the same tool again — before expiresAt — with the identical functionPath, args, shardKey and idempotencyKey, plus confirmed: true and that actionDigest. Only that second call writes.

The confirmation is bound to the exact action a human saw, for ten minutes. actionDigest is <expiresAt>.<signature>, the signature an HMAC over a canonical, sorted-key encoding of the tool name, function path, arguments, shard key, idempotency key and that deadline, keyed by the deployment's own identity — so re-serializing args in a different key order still verifies, while changing the target, any argument, or the shard key produces a different digest and the confirmation is refused with nothing written. A digest minted against another deployment never verifies, a lunora_run_mutation digest can never confirm a lunora_run_action, and an expired one is refused rather than silently re-proposed.

The digest carries its own proof rather than naming a stored record, because there is no store to name: createMcpFetchHandler serves statelessly — a fresh MCP server per HTTP request, no session, no cross-request state — and the confirming request need not even reach the instance that issued the proposal. Any instance holding the same deployment URL and admin bearer recomputes and verifies the same digest; nothing without that bearer can mint one. The deadline rides inside the digest for the same reason: there is nowhere else to keep it.

What the handshake does not do

It binds intent, not human presence.

A verified digest proves the call about to run is exactly the call that was proposed, on this deployment, inside its window. It does not prove a human saw it, and nothing server-side can: an MCP server has no channel to a person, and MCP puts the human-in-the-loop at the host — the client is what renders a tool call for approval. A client that asks nobody can send the digest it was just handed straight back with confirmed: true and the write runs.

That is why the write surface is off by default and refused at dispatch as well as omitted from ListTools. Setting allowWrites is the operator's statement that the client on the other end does the asking; the handshake is a client-UI affordance and an audit record of what was proposed, not a gate against the model.

Its scope is also deployment-wide, not principal-bound: the signing key is the deployment URL plus the admin bearer, with nothing identifying a user in it. On an OAuth-fronted server every principal shares that bearer, so inside the ten-minute window any principal holding write scope can confirm another's identical proposal. Binding a confirmation to a person would mean folding the verified sub claim into the key, which this package does not do today.

What idempotencyKey does and does not guarantee

idempotencyKey is an optional caller-chosen token folded into the digest.

It guarantees: a client that timed out can resubmit the confirmation it already holds, for as long as that digest is inside its window, instead of asking for a second human review — and a deliberately-repeated identical write sent under a new key gets its own digest, so it cannot ride the first review.

It does not guarantee deduplication. This server keeps no state between requests, so it cannot remember that a call already ran and cannot replay an earlier result; it also never forwards the key to your function, which never declared it as an argument. A resubmitted confirmed call therefore executes again. If the underlying write must happen at most once, make the function itself idempotent — for example by storing a caller-supplied key in a row with a unique index and short-circuiting on a repeat.

An agent discovers the surface before it calls anything:

  1. lunora_list_functions: discover the available paths and their kinds.
  2. lunora_get_function_schema: fetch the argument descriptors for one path.
  3. lunora_run_query / lunora_run_mutation / lunora_run_action: call the function with a well-formed args object. For the two write tools this is two calls — propose, have a human review proposedAction, then resubmit with confirmed: true and the actionDigest.

lunora_get_function_schema returns { path, kind, args }, where kind is "query", "mutation", or "action" and args is an array of argument descriptors (name, kind, optional, and optionally element or table). It returns an isError result if the path doesn't exist.

Building a server programmatically

createLunoraMcpServer returns a transport-agnostic MCP Server. Pass a url (and optional token), and the tools dispatch against a LunoraClient built from them:

import { createLunoraMcpServer } from "@lunora/mcp";

const server = createLunoraMcpServer({ url: "https://app.example.workers.dev", token: "..." });

await server.connect(myTransport);

For the common stdio case, connectStdio builds the server and connects it over a StdioServerTransport in one step:

import { connectStdio } from "@lunora/mcp";

await connectStdio({ url: process.env.LUNORA_URL, token: process.env.LUNORA_ADMIN_TOKEN });

LunoraMcpServerOptions accepts either a url with a token (plus an optional fetch) or a pre-built client; the latter is the injection point for tests. You must supply one or the other; passing neither — or a url with no token — throws.

Auth and admin gating

The token (sourced from LUNORA_ADMIN_TOKEN for the binary) is set as the client's auth token and sent as a bearer token on every RPC the tools make. It has to be the admin bearer: introspection (lunora_list_functions, lunora_list_tables) and the allowlist precheck in front of every run tool go through admin-gated /_lunora/admin/* routes.

So the read-only guarantee does not come from the token's scope. It comes from allowWrites (LUNORA_MCP_ALLOW_WRITES) defaulting off, which omits the write tools from tools/list and refuses them at dispatch. With writes on, the confirmation handshake is the second layer: a mutation or action executes only on a call carrying a digest issued for those exact arguments, so "this server may write" and "this write was reviewed" stay separate questions. Beyond that, gating is enforced by your deployment: the tools call ordinary queries, mutations, and actions, so whatever auth those functions require applies unchanged. Treat the MCP server process itself as the trust boundary and control who can reach it — createAuthedMcpFetchHandler is the supported way to expose it beyond a local stdio process. Run-tools open no WebSocket (every call is plain HTTP RPC), so the server is safe to run as a short-lived stdio process.

Serving MCP over OAuth

createMcpFetchHandler serves anyone who can reach the URL. That is the right trade for a stdio binary on your laptop, and the wrong one for a public endpoint: the tools carry the deployment's admin bearer, so the network path is the authorization.

createAuthedMcpFetchHandler mounts the same server behind an OAuth 2.1 gate — the flow MCP clients already know how to walk, discovered through the RFC 9728 protected-resource metadata that better-auth's mcp() plugin serves. The authorization server is a @lunora/auth instance running that plugin; the resource server is the handler:

// lunora/auth.ts
import { createAuth } from "@lunora/auth";
import { jwt, mcp } from "@lunora/auth/plugins";

// The canonical protected-resource id (RFC 8707 / RFC 9728). Issued tokens are
// audience-bound to it. HTTPS, no query or fragment.
export const mcpResource = "https://app.example.workers.dev/mcp";

export const auth = createAuth({
    database: env.DB,
    secret: env.AUTH_SECRET,
    plugins: [
        jwt(),
        mcp({
            consentPage: "/consent",
            loginPage: "/login",
            resource: mcpResource,
            // Declare the scopes the MCP route checks. Left out, `mcp()` falls back
            // to the OIDC defaults and no client can be granted `lunora:read`.
            scopes: ["lunora:read", "lunora:write"],
        }),
    ],
});
// The MCP route.
import { requireMcpAuth } from "@lunora/auth/plugins";
import { createAuthedMcpFetchHandler, mcpTokenScopes } from "@lunora/mcp";

import { auth, mcpResource } from "./auth";

export const handleMcp = createAuthedMcpFetchHandler({
    // The same `resource` as `mcp()`: tokens are audience-bound to it. Lunora's
    // `requireMcpAuth` requires it and throws at startup when it is missing or not
    // an absolute URL. (better-auth's own version silently falls back to the auth
    // `baseURL`, an audience no `mcp()` token carries, and refuses every request.)
    protect: (handler) => requireMcpAuth(auth, handler, { requiredScopes: ["lunora:read"], resource: mcpResource }),
    server: (claims) => ({
        // Writes need a scope a read-only token does not carry.
        allowWrites: mcpTokenScopes(claims).has("lunora:write"),
        token: env.LUNORA_ADMIN_TOKEN,
        url: env.LUNORA_URL,
    }),
});

An unauthenticated request never reaches the MCP server at all — it gets a 401 carrying the WWW-Authenticate challenge that starts the client's authorization flow, and no LunoraClient holding the admin bearer is ever constructed for it.

The discovery documents

The challenge sends the client to https://app.example.workers.dev/.well-known/oauth-protected-resource/mcp, and from there to the authorization server's metadata at /.well-known/oauth-authorization-server/api/auth. better-auth serves both, outside the /api/auth/* base path, and a Lunora worker built with .auth(...) serves them for you: it derives exactly those two paths from the auth options (the resource of mcp(), the issuer's base path) and answers them only when an mcp() or oauthProvider() plugin is configured. Use mcp from @lunora/auth/plugins rather than @better-auth/mcp: Lunora's records its resource, which is how the worker knows the first path. Without that record, the worker falls back to the provider's only resource, and createAuth throws AUTH_MCP_RESOURCE_AMBIGUOUS when there are several to choose from.

The documents are the worker's last matcher, reached only by a GET/HEAD that the explicit routes and your httpRouter left unanswered (a 404). An app that already routes either path itself keeps serving its own answer. Nothing else under /.well-known/ (security.txt, app links) is touched.

A worker you compose by hand, outside the generated app builder, routes them itself. mcpDiscoveryPaths(resource) derives both paths from the resource URL (pass the auth basePath as a second argument if you changed it):

import { mcpDiscoveryPaths } from "@lunora/auth/plugins";

import { auth, mcpResource } from "./auth";

const discoveryPaths = new Set(mcpDiscoveryPaths(mcpResource));

export const handleRequest = async (request: Request): Promise<Response> => {
    const { pathname } = new URL(request.url);

    if (discoveryPaths.has(pathname) && (request.method === "GET" || request.method === "HEAD")) {
        return await auth.handler(request);
    }

    return await handleMcp(request);
};

Scoping tools to the token

server takes the verified token claims, not just a fixed options object. A gate that only answers yes/no gives every authorized agent the same capabilities, which throws away the scopes the token was issued with. Resolving allowWrites per request is what lets one endpoint serve a read-only agent and a read-write one: for the former the write tools are omitted from tools/list and refused at dispatch.

mcpTokenScopes(claims) parses the space-delimited scope claim (RFC 6749 §3.3) into a Set. A malformed or absent claim yields an empty set rather than throwing, so a scope check denies instead of surfacing a 500 a client might retry.

Step-up for writes

Hiding the side-effecting tools suits a client that can never get the write scope (a client_credentials agent). An interactive client can instead re-authorize for more scope when it needs it, if the server tells it which. Set stepUp, and a read-only token still sees the write tools (lunora_run_mutation, lunora_run_action) and the agent tools (agent_<name> or its toolName, and lunora_agent_status: starting a durable run is a side effect too). A tools/call to one is answered with HTTP 403 and an RFC 6750 challenge, WWW-Authenticate: Bearer error="insufficient_scope", scope="lunora:write", resource_metadata="…", which the client's auth layer turns into a new authorization for that scope:

import { createInsufficientScopeError, requireMcpAuth } from "@lunora/auth/plugins";
import { createAuthedMcpFetchHandler } from "@lunora/mcp";

import { auth, mcpResource } from "./auth";

export const handleMcp = createAuthedMcpFetchHandler({
    protect: (handler) => requireMcpAuth(auth, handler, { requiredScopes: ["lunora:read"], resource: mcpResource }),
    // The scope gate decides writes now, so the server always offers them.
    server: { allowWrites: true, token: env.LUNORA_ADMIN_TOKEN, url: env.LUNORA_URL },
    stepUp: {
        // better-auth only turns an error this factory created into the challenge.
        challenge: createInsufficientScopeError,
        scope: "lunora:write",
    },
});

stepUp needs both fields; an empty scope or a missing challenge throws at construction, because a scope with no way to raise its challenge would let every write through. The handler reads the JSON-RPC message through the same bounded parse the transport uses before it decides, so a body too large or not JSON is refused, and a batch carrying a side-effecting call is challenged like a single one. Without stepUp, nothing changes: tools follow allowWrites and allowAgents alone.

protect is a lambda rather than an auth instance because @lunora/mcp does not depend on better-auth — and the same seam takes either better-auth entry point. Use requireMcpAuth(auth, handler, opts) when the resource server shares a deployment with the authorization server, or createMcpProtectedRequestHandler(verifyOptions, handler) when it does not (a separate resource server holds verification config, not an auth instance).

Letting MCP clients register themselves (CIMD)

mcp() registers no clients on its own: dynamic client registration is off unless you enable it. MCP 2026-07-28 pins Client ID Metadata Documents instead. The client's client_id is an HTTPS URL, and the authorization server fetches the client's metadata from it. That is how Claude, ChatGPT and other MCP hosts connect to a server they have never seen. Add the cimd() plugin next to mcp():

import { createAuth } from "@lunora/auth";
import workersCimdFetch from "@lunora/auth/cimd/workers";
import { cimd, jwt, mcp } from "@lunora/auth/plugins";

export const auth = createAuth({
    database: env.DB,
    secret: env.AUTH_SECRET,
    plugins: [
        jwt(),
        mcp({ consentPage: "/consent", loginPage: "/login", resource: mcpResource, scopes: ["lunora:read", "lunora:write"] }),
        cimd({
            fetchClientMetadataResource: workersCimdFetch(),
            // MCP's profile: `client_name` and `redirect_uris` become mandatory.
            metadataProfile: "mcp-2026-07-28",
            // Recommended in production: only fetch metadata from hosts you trust.
            isMetadataDocumentUrlAllowed: (url) => ["https://claude.ai", "https://chatgpt.com"].includes(new URL(url).origin),
            // The allowlist gates only the client_id fetch. Bind jwks_uri to that
            // origin too, or a document can point the key fetch at any public host.
            originBoundFields: ["post_logout_redirect_uris", "client_uri", "jwks_uri"],
        }),
    ],
});

The authorization-server metadata then advertises client_id_metadata_document_supported: true, and a client that presents a URL client_id is created from its document on first use.

That fetch goes to a URL the client chose, so the transport must keep it on the public internet. On Workers, workersCimdFetch() relies on the global_fetch_strictly_public compatibility flag, which stops fetch from reaching private and special-use addresses. Add it to wrangler.jsonc:

{
    "compatibility_flags": ["nodejs_compat", "global_fetch_strictly_public"],
}

workersCimdFetch() throws at startup when the flag is missing, and lunora doctor reports it (cimd-fetch-not-strictly-public). The transport also refuses anything but an https: GET/HEAD and never follows a redirect. On Node, pass the default export of @lunora/auth/cimd/node instead (import fetchClientMetadataResource from "@lunora/auth/cimd/node"); it pins the connection to the address it resolved.

The flag is a weaker guarantee than cimd asks for: a Worker cannot pin the connection to the address it resolved, so a host that rebinds its DNS still reaches some public address. That is the same trade Cloudflare's own workers-oauth-provider makes. An isMetadataDocumentUrlAllowed origin allowlist narrows it to hosts you trust, but does not close it, and it only covers the client_id fetch: oauth-provider fetches a document's jwks_uri through the same transport without asking it. Add jwks_uri to originBoundFields, as above, so the key fetch stays on the allowlisted origin. Metadata is cached per isolate, so each new isolate fetches a client's document again on its first authorization.

Connecting an MCP client

Any MCP client that can spawn a stdio server works. With the mcpServers config above, an agent like Claude can call lunora_list_functions to discover the deployment's surface, then lunora_run_query / lunora_run_mutation / lunora_run_action to invoke it, passing functionPath (e.g. "messages:list"), an args object, and an optional shardKey. The two write tools answer the first call with an action_required proposal; the client shows it to the user and calls again with confirmed: true and the returned actionDigest — see write confirmation.

Resources and annotations

Both servers implement MCP tools. The documentation server additionally exposes every page as an MCP resource (lunora-docs:/docs/…, text/markdown), so a client can list and attach a page on the user's behalf, before the model knows what to search for.

Every tool carries annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint, title), so a client can badge the read-only surface and prompt before a write. These are hints for presentation; the guarantee itself is still enforced at dispatch — by the allowWrites / allowObservability gates, and for the two write tools by the confirmation handshake.

The documentation server

@lunora/mcp/docs is an independent surface with three tools:

ToolDescription
lunora_search_docsSearch the docs; returns matching pages and sections with their URLs.
lunora_get_docReturn one page in full, as Markdown.
lunora_list_docsList every page with its title and description.

It touches no user data and needs no token, so it is safe to expose publicly. Nothing in the subpath imports @lunora/client or a Node built-in, so it runs unchanged on Workers, Netlify/Vercel functions, Deno, and Bun.

The tools read a DocsIndex, and two implementations satisfy that contract. A docs site wires up its own in-process search index and mounts the server as a route:

import { createDocsMcpFetchHandler } from "@lunora/mcp/docs";

const handle = createDocsMcpFetchHandler({ index: myDocsIndex });

Anything else (the CLI's lunora mcp serve, a script) reads a published site over its /api/search, /llms.mdx/*, and /llms.txt endpoints:

import { createDocsMcpServer, createRemoteDocsIndex } from "@lunora/mcp/docs";

const server = createDocsMcpServer({ index: createRemoteDocsIndex({ baseUrl: "https://lunora.sh" }) });

Because both backends map results through the same helper, a model sees identical hits whichever one answered.

Hosting it safely

createDocsMcpFetchHandler screens each request before the transport sees it, because this surface is meant to be public and unauthenticated:

  • Bodies are capped: 128 KiB by default, overridable with maxRequestBytes.
  • JSON-RPC batches are refused. The stateless transport buffers a whole batch's replies into one response body, so a single small request carrying thousands of tools/call messages would amplify into hundreds of megabytes out, with no initialize and no session to rate-limit against. A documentation client gains nothing from batching.
  • lunora_search_docs bounds its query, and lunora_list_docs caps how many pages it serialises in one call.

Composing a local server

createLocalMcpServer / connectLocalStdio assemble the docs tools, the deployment tools, and any extra tools a host supplies into one stdio server. This is what lunora mcp serve runs, and it is why the CLI depends only on this package rather than on the protocol SDK.

import { connectLocalStdio } from "@lunora/mcp";

await connectLocalStdio({
    deployment: () => readMyDevServer(),
    docs: { baseUrl: "https://lunora.sh" },
    extraTools: myLocalTools,
});

deployment accepts a resolver, consulted on every tool call. An editor spawns its MCP servers when the project opens, usually before a dev server is running, and keeps them alive across every restart, so a URL captured once at startup would be absent for the whole first session and stale after the first restart. The deployment tools are advertised either way (clients cache the tool list); calling one with nothing running returns an actionable error.

The observability tools are the one exception. Their gate is a snapshot taken when the tool list is built, so a session that started before lunora dev was running does not advertise them — and because clients cache the list, they stay absent for that session even after the dev server comes up. Restart the MCP server (or the editor) once the dev server is running to get them.

Expose an agent

A deployment's durable @lunora/agent agents can be fronted as MCP tools. Like writes, this is opt-in and fail-closed: starting an agent run is a side effect, so the agent tools are hidden from the advertised list and refused at dispatch unless you enable them. @lunora/mcp takes no dependency on @lunora/agent; it calls the agent's public agents:agentRun mutation over HTTP RPC like any other function.

Two opt-ins are needed, and they sit on opposite sides of the RPC:

On the agent, mark it runnable from outside the deployment:

export const support = defineAgent({ name: "support", publicRun: true /* … */ });

agents:agentRun throws FORBIDDEN for any agent whose handle is not publicRun: true — starting a durable run is a side effect, so the agent's author allows it, not the MCP server. The env vars below cannot grant it: an agent exposed here without the flag is advertised in tools/list and fails on its first call.

On this server, the env vars (or the matching createLunoraMcpServer options):

  • LUNORA_MCP_ALLOW_AGENTS: 1/true/yes/on to expose the agent tools.
  • LUNORA_MCP_AGENTS: a ;-separated list of name:description pairs selecting which agents to expose, e.g. "support:Support questions;billing:Billing help".
  • LUNORA_MCP_AGENT_TIMEOUT_MS (optional): wall-clock budget a single agent_<name> call awaits before returning a pending result to poll.
{
    "mcpServers": {
        "lunora": {
            "command": "lunora-mcp",
            "env": {
                "LUNORA_URL": "https://app.example.workers.dev",
                "LUNORA_ADMIN_TOKEN": "...",
                "LUNORA_MCP_ALLOW_AGENTS": "1",
                "LUNORA_MCP_AGENTS": "support:Support questions;billing:Billing help",
            },
        },
    },
}

Each exposed agent gets an agent_<name> tool taking:

  • prompt (required, string): the task or message for the agent.
  • threadKey (string): reuse to continue a conversation; omit to start a new thread. The tool returns the threadKey it used either way.
  • title (string): an optional thread title, applied on the first run only.

The tool starts a durable run and awaits it up to the timeout budget. If the run outlasts the budget, the tool returns a pending result ({ status: "running", threadKey, runId, hint }) instead of hanging. Feed that threadKey to the generic lunora_agent_status tool to poll for the final answer once the run finishes.

Agent runs are owner-scoped to the identity the configured token resolves to — the deployment's admin identity, since that is the only token every tool works with. Every thread this server starts therefore shares one owner; run a separate MCP server per deployment to keep threads isolated. If the token resolves to no identity, threads are not owner-isolated at all.