Skip to content
DocspackagesDocumentation

@lunora/ai

Workers AI inference from your functions, provider-agnostic, Workers AI by default.

PackagesAi

@lunora/ai is a small helper over the Vercel AI SDK and Cloudflare's workers-ai-provider. Call generateText / streamText / generateObject / embed / tool from any function. Workers AI is the zero-config default, but every call is provider-agnostic: pass a Workers AI model id (a string) or any AI SDK model object (@ai-sdk/openai, @ai-sdk/anthropic, OpenRouter, …).

pnpm add @lunora/ai

When a function uses AI, the dev server / lunora prepare reconciles the ai binding into wrangler.jsonc for you ({ "ai": { "binding": "AI" } }), and codegen wires a typed ctx.ai onto your action contexts.

ctx.ai in an action

Inference is an external, non-deterministic call, so like ctx.fetch, ctx.ai lives on actions, not queries or mutations.

import { action, v } from "@/lunora/_generated/server";
import { generateText } from "@lunora/ai";

export const summarize = action.input({ text: v.string().max(20_000) }).action(async ({ ctx, args: { text } }) => {
    const { text: summary } = await generateText({
        // Both ends of the token bill are bounded: the input by `.max()` above,
        // the completion here. An `action` is public RPC and inference is
        // metered, so an unbounded completion is a denial-of-wallet vector — a
        // short prompt can ask for an arbitrarily long answer, and output
        // tokens are the expensive half.
        maxOutputTokens: 300,
        model: ctx.ai.model("@cf/meta/llama-3.3-70b-instruct-fp8-fast"),
        prompt: `Summarize:\n\n${text}`,
    });

    return summary;
});

Bound every generation reachable from a public procedure, and rate-limit the entry point — the ai_unbounded_generation_public advisor lint flags the ones that are not. See the ai registry item (lunora add ai) for the full auth + rate-limit + bounds shape.

ctx.ai.model(id) resolves a model from the binding — a Workers AI id, or any provider's model through AI Gateway (next section). Pass the resolved model to the AI SDK functions re-exported from @lunora/ai (generateText, streamText, generateObject, streamObject, embed, embedMany, tool).

Any provider, same call

Switch provider by changing the model string — nothing else:

const model = ctx.ai.model("anthropic/claude-sonnet-5"); // or "openai/gpt-5", "google/gemini-2.5-pro", …

A "<provider>/<model>" id is routed through Cloudflare AI Gateway over the same AI binding, so the app holds no provider API key. Which way the provider is paid depends on the provider, and is set up once in the Cloudflare dashboard:

  • Unified Billing — for the unified-catalog providers (OpenAI, Anthropic, Google, xAI, Groq, DeepSeek). Their calls take the binding's run path, where Cloudflare holds the provider credentials and bills against prepaid credits; a key stored on the gateway is not used for them.
  • Keys stored on the gateway — for the other OpenAI-compatible providers (Mistral, Perplexity, Cerebras, OpenRouter, Fireworks, …), which only exist on the gateway path. The key lives in the gateway's Secrets Store, not in the app.

Calls go to the gateway named by LUNORA_AI_GATEWAY_ID, or to the account's auto-created default gateway when it is unset. lunora ai gateway creates a gateway for the app and writes that var for you. Each call carries its function path and trace id as gateway metadata, so the gateway's logs attribute spend to the function that made it.

A dynamic/<route> id calls an AI Gateway dynamic route — the place to configure fallback models, per-model budgets and spend limits without a code change.

@cf/… ids stay on Workers AI, exactly as before.

Your own proxy (celld, or no Cloudflare)

Set LUNORA_AI_PROXY_URL to any OpenAI-compatible endpoint (LiteLLM, OpenRouter, a self-hosted proxy) and <provider>/<model> slugs go there instead of AI Gateway. The slug is sent unchanged as the request's model, and LUNORA_AI_PROXY_TOKEN (a secret) is sent as the bearer token. No AI binding is needed, so this is how ctx.ai works on celld. @cf/… ids and ctx.ai.run still need the binding.

{ "vars": { "LUNORA_AI_PROXY_URL": "https://ai-proxy.internal/v1" } }

Bring your own provider

An AI SDK model object passes straight through, for a provider or feature the gateway catalog does not cover:

import { streamText } from "@lunora/ai";
import { createOpenAI } from "@ai-sdk/openai";

const openai = createOpenAI({ apiKey: env.OPENAI_API_KEY });
const result = streamText({ model: ctx.ai.model(openai("gpt-5")), messages });

Wrap it in ctx.ai.model(...) as above so its usage is recorded like any other call (see Usage and spend per function).

Structured output

import { generateObject } from "@lunora/ai";
import { z } from "zod";

const { object } = await generateObject({
    model: ctx.ai.model("@cf/meta/llama-3.3-70b-instruct-fp8-fast"),
    schema: z.object({ sentiment: z.enum(["positive", "neutral", "negative"]) }),
    prompt: review,
});

Failing fast when Workers AI is busy

By default a synchronous Workers AI call waits in a capacity queue when no capacity is free. Pass rejectIfBusy: true to fail at once instead. It applies to Workers AI model ids (@cf/…) only:

import { generateText } from "@lunora/ai";

const { text } = await generateText({
    model: ctx.ai.model("@cf/meta/llama-3.3-70b-instruct-fp8-fast", { rejectIfBusy: true }),
    // The rejection is an AI SDK `APICallError` with `statusCode: 429`, which the
    // AI SDK retries by default. Turn retries off to actually fail fast.
    maxRetries: 0,
    prompt,
});

// Raw binding: a rejection throws a `LunoraError` with code `RATE_LIMITED`.
await ctx.ai.run("@cf/meta/m2m100-1.2b", inputs, { rejectIfBusy: true });

ctx.ai.websearch(query, options?) calls the Cloudflare Web Search API (beta) over the same AI binding, so a response can be grounded in live results. The providers are "ceramic" (the default), "exa" and "linkup", all under Zero Data Retention. The search goes through the gateway inference already routes through (LUNORA_AI_GATEWAY_ID), else the account's default gateway; pass gatewayId to pick another. It bills a provider key stored on that gateway under the default alias (or the one byokAlias names) when there is one, and AI Gateway credits at list price otherwise.

const { items } = await ctx.ai.websearch("cloudflare durable objects pricing", { limit: 5, provider: "exa" });
// items: [{ url, title, description?, lastModifiedDate?, imageUrl?, faviconUrl? }, …]

limit defaults to 10 and accepts at most 10. Results describe pages; reading a page is a separate fetch of its url.

A rejected request throws a LunoraError with code BAD_REQUEST. Gateway or credential problems keep their own codes (UNAUTHORIZED, FORBIDDEN, NOT_FOUND), a throttled search is RATE_LIMITED, and a provider or gateway failure SERVICE_UNAVAILABLE. An unreadable response is INTERNAL. For agents, webSearchTool() from @lunora/agent wraps this as a tool.

RAG — defineRag (@lunora/ai/rag)

defineRag composes ctx.ai (embeddings) with ctx.vectors (Vectorize) into a declared index → retrieve pipeline: chunk → embed → upsert on the write side, embed → query → assemble on the read side. It's a thin library over the two facades every action already has, with no new binding and no codegen.

// lunora/rag.ts
import { defineRag } from "@lunora/ai/rag";

export const docs = defineRag({
    embeddingModel: "@cf/baai/bge-base-en-v1.5", // declared once → index + retrieve embed identically
    index: "docs", // a ctx.vectors index binding key
});

embeddingModel is optional only if a default is configured: the later examples on this page omit it, and they resolve through ctx.ai.embeddingModel(), which reads LUNORA_AI_DEFAULT_EMBEDDING_MODEL from the Worker env (wrangler vars or .dev.vars). With neither, the first index/retrieve throws. The same holds for ctx.ai.model() with no argument and LUNORA_AI_DEFAULT_MODEL.

// inside an action:
import { docs } from "@/lunora/rag";

await docs(ctx).index({ id: doc._id, metadata: { title: doc.title }, namespace: ctx.shardKey, text: doc.body });

const { chunks, context, sources } = await docs(ctx).retrieve(question, { namespace: ctx.shardKey, topK: 5 });
// `context` is prompt-ready; `chunks` are ranked; `sources` are deduped refs.

What you get beyond the manual loop:

  • Deterministic chunk ids (${sourceId}#${n}): re-indexing a source replaces its chunks; shrinking documents have stale trailing chunks deleted automatically.
  • Content-hash short-circuit: re-indexing an unchanged source skips chunking, embedding, and every write ({ unchanged: true }), so periodic re-syncs are free. The hash covers text, metadata and importance together, so changing a document's tenant or ACL tags over an unchanged body still re-indexes.
  • Tenant isolation: thread namespace (your shard/tenant key) through both sides; a namespace-less call gets a one-time dev warning (Vectorize indexes are account-global). Multi-tenant apps should set requireNamespace: true to turn the warning into a hard error; single-tenant apps suppress it with allowSharedNamespace.
  • Ranking controls: minScore threshold, per-source importance weighting (0 to 1, multiplied into scores), and chunkContext: { before, after } to stitch neighbouring chunks around each match ("embed small, retrieve big").
  • asTool(): expose retrieval as an AI SDK tool so a model can decide to search the index itself: tools: { searchDocs: docs(ctx).asTool() }.
  • Traced: when the bound context carries ctx.trace, each embedding call is a generation span (gen_ai.operation.name: "embeddings", gen_ai.request.model), so RAG shows up on the trace waterfall. A hand-built context without ctx.trace embeds untraced.

Chunk text lives in vector metadata by default (returnMetadata: "all", topK capped at 50, 10 KiB of metadata per vector). For long documents or deeper retrieval, supply a textStore ({ put, getMany, remove? }: a DO table, KV, …): text moves out of metadata and the topK ceiling lifts to 100. The default chunker is a fixed 1000-char window with 200 overlap; pass chunk for the sentence, Markdown, or token strategies below.

The 10 KiB is Vectorize's, and it covers the whole metadata object: chunk text, Lunora's bookkeeping keys, and any metadata you attach. defineRag measures each chunk's metadata as it is assembled and refuses one that would not fit, rather than letting the upsert fail at Vectorize with nothing naming the cause. It also rejects an unworkable chunkSize up front, though that earlier check can only compare characters against a byte ceiling; multibyte text costs up to three bytes each, so the index-time measurement is the one that holds.

The topK 50 is Vectorize V2's own ceiling for full-metadata queries. Legacy V1 indexes cap at 20 and will reject a larger topK remotely; a binding handle does not expose its index version, so the check cannot branch on it.

Chunking strategies

chunkSize/chunkOverlap drive the built-in fixed character window. For real documents, pass chunk one of the structure-aware chunkers instead — a chunk that starts mid-clause embeds to a worse vector than the same prose split where the author ended a thought:

import { defineRag, markdownChunker, sentenceChunker, tokenChunker } from "@lunora/ai/rag";

export const docs = defineRag({
    chunk: markdownChunker({ overlap: 200, size: 1000 }),
    index: "docs",
});
  • sentenceChunker({ size, overlap }) — packs whole sentences into size-bounded windows. The default choice for prose.
  • markdownChunker({ size, overlap }) — splits at ATX headings (code fences are tracked, so a # comment inside a fence never starts a section), then packs each section's sentences. Every chunk is prefixed with its heading trail (# Guide > ## Auth > ### OAuth), so a chunk from deep inside a long document still carries what it is about — which is what makes it retrievable by a query naming its section rather than its prose.
  • tokenChunker({ countTokens, maxTokens, overlapTokens }) — bounds chunks by a real token count, for when the embedding model's context window is the binding constraint (anything longer is silently truncated, so the tail is embedded as if it were never written). countTokens is required and injected: @lunora/ai will not add a tokenizer dependency or pretend a characters-per-token constant is a token count. Pass js-tiktoken, gpt-tokenizer, or your provider's counter.

None of them can exceed size: an atom too large to fit falls back to a hard character split rather than emitting a chunk the store would reject.

Embedding batching and caching

Indexing embeds a whole document in one embedMany call rather than one call per chunk: the per-chunk embed callbacks ctx.vectors.upsert invokes resolve from that batch. A 200-chunk document collapses from 200 round-trips to a handful. It is best-effort — a provider that rejects the batch falls back to per-chunk embeds rather than failing the index — and identical chunks dedupe to a single embed.

Retrieval re-embeds the query on every call. Set cacheEmbeddings to retain embeddings across calls on a bound context:

export const docs = defineRag({ cacheEmbeddings: 32, index: "docs" });

Sized in entries but budget in bytes: one 1536-dimension embedding is ~12 KB, so 100 entries is over a megabyte held in the isolate. Keep it small. The cache is scoped to the bound context, never module level, so it cannot outlive the request that built it.

Embedding dimensions

Vectorize stores at most 1536 dimensions per vector, which rules out most current large embedding models (text-embedding-3-large and Gemini embedding at 3072, Qwen3-Embedding at 4096). defineRag measures the first embedding each bound context produces and refuses a wider one, naming the ceiling and both escapes — rather than letting the upsert fail at Vectorize with nothing saying why:

import { openai } from "@ai-sdk/openai";

export const docs = defineRag({
    // Matryoshka truncation keeps a large model under the ceiling:
    embeddingModel: openai.textEmbeddingModel("text-embedding-3-large", { dimensions: 1536 }),
    index: "docs",
});

Set maxEmbeddingDimensions: false to disable the check when the index is not Vectorize-backed, or give it a number to enforce a different store's ceiling. The check runs once per bound context, not once per chunk.

Bring your own vector store

defineRag used to hard-code Vectorize's limits — its topK ceilings, its 10 KiB metadata budget, its 1536-dimension cap — which made every RAG index a Cloudflare index. A store now declares its own capabilities and defineRag reads them, so a backend without those constraints is not held to them:

import { defineRag } from "@lunora/ai/rag";
import type { RagVectorStore } from "@lunora/ai/rag";

const pgStore = (ctx: { sql: SqlClient }): RagVectorStore => ({
    capabilities: {
        maxDimensions: false, // pgvector has no 1536 ceiling
        maxMetadataBytes: false, // nor a 10 KiB metadata budget
        maxTopK: 1000,
        maxTopKWithMetadata: 1000,
    },
    deleteByIds: (ids, namespace) => ctx.sql`delete from chunks where id = any(${ids}) ...`,
    getByIds: (ids, namespace) => ctx.sql`select ... `,
    query: (input) => ctx.sql`select ... order by embedding <=> ${vector} limit ${input.topK}`,
    upsert: (input) => ctx.sql`insert into chunks ... on conflict (id) do update ...`,
});

export const docs = defineRag({ index: "docs", store: (ctx) => pgStore(ctx) });

store is called once per bound context with that context, so a store needing per-request state — a Hyperdrive connection off ctx.sql, a shard's own SQLite — builds itself from it.

Omit store and the bound ctx.vectors is wrapped by vectorizeStore, which declares exactly the limits listed above. That is the default and its behaviour is unchanged.

An explicit maxEmbeddingDimensions still wins over the store's own maxDimensions, so you can tighten a permissive backend; leave it unset and the store speaks for itself.

Shipped: a SQLite-backed store

sqliteVectorStore runs a RAG index on any SQL engine reachable through an injected executor — a Durable Object's SQLite, D1, or node:sqlite. No Vectorize, no extra Cloudflare product.

import type { RagSqlExec } from "@lunora/ai/rag";
import { defineRag, sqliteVectorStore } from "@lunora/ai/rag";

// Whatever your engine exposes. On D1:
const exec: RagSqlExec = async (sql, parameters) =>
    (
        await env.DB.prepare(sql)
            .bind(...parameters)
            .all()
    ).results;
// Inside a Durable Object it is `(sql, parameters) => state.storage.sql.exec(sql, ...parameters).toArray()`.

export const docs = defineRag({
    index: "docs",
    store: () => sqliteVectorStore({ exec }),
});

ctx.vectors is not needed here — with a store configured it is never read, and RagContext.vectors is optional, so an app with no Vectorize index type- checks.

When the executor is a shard's own SQLite, the shard is the tenant boundary — so the account-global-namespace hazard that Vectorize forces you to guard against does not exist here, rather than being filtered away.

Search is brute force by default. Every vector in the namespace is read and scored in JS; there is no ANN index by default, because SQLite has no vector type and sqlite-vec is not loadable inside workerd. That is linear in namespace size, so this suits many small per-tenant indexes — the shape most sharded apps already have — and not one large shared corpus. A namespace past maxScan (default 50,000) throws a named error rather than letting a Worker get killed on CPU with nothing explaining why.

Where sqlite-vec is loaded, add an index. A celld Durable Object whose worker sets the sqlite_vec compatibility flag, or a node:sqlite database with the extension loaded, can pass ann:

sqliteVectorStore({ ann: { dimensions: 768 }, exec });

The store then keeps a vec0 index (cosine distance, one partition per namespace) in <table>_ann, backfilled from the table the first time, and an unfiltered query is a KNN lookup instead of a scan. A query with a filter still scans: vec0 filters after choosing the k nearest, which would return a short page. Without the extension the first operation throws naming the missing vec0 module. Workerd cannot load sqlite-vec, so this does not apply on Cloudflare.

Limits of the index:

  • It needs code that runs inside the Durable Object. celld loads sqlite-vec only into Durable Object SQLite, not into D1, and Lunora actions run in the Worker rather than inside a shard, so a Lunora function has no exec over a shard's SQLite to hand the store. Use it from a Durable Object class your app writes (its state.storage.sql), or from node:sqlite.
  • Cosine only, over one fixed width: ann.dimensions is the vec0 column width, so an embedding of another width — or reopening the table with a different ann.dimensions — fails with RAG_DIMENSION_MISMATCH. Changing the model's width means a new table.
  • An interrupted build is rebuilt, not trusted. The index counts as built only once its backfill finishes, so a build cut off part-way is dropped and redone on the next call.
  • Filtered queries still scan the namespace, with the same maxScan bound as without the index.
  • Writes are two statements, not one transaction. An interrupted upsert or delete can leave an index entry the table no longer holds; a query drops it, so it costs a topK slot rather than returning a stale record.

Bring your own embeddings — no env.AI binding

embeddingModel takes a Workers AI model id (a string, resolved through ctx.ai, so it needs the env.AI binding) or a ready-made AI SDK EmbeddingModel object. Pass an object and the helper embeds through it directly, never touching ctx.ai, so a RAG index over OpenAI (or any provider) needs no Workers AI binding at all. ctx.vectors (Vectorize) is still required: it's the store.

import { openai } from "@ai-sdk/openai";
import { defineRag } from "@lunora/ai/rag";

// A model *object*, not an id → embeds without ctx.ai / env.AI.
export const docs = defineRag({
    embeddingModel: openai.textEmbeddingModel("text-embedding-3-small"),
    index: "docs",
});

Because the object path skips ctx.ai, you can even bind a hand-built context carrying only vectors (e.g. in a test or a non-action caller): docs({ vectors: ctx.vectors }).retrieve(question). A model-id string with no ctx.ai present throws a directed error telling you to pass a model object or wire ctx.ai.

Embedding-model versioning

Embeddings from different models are not comparable: swap the model and every old vector becomes noise a nearest-neighbour query still happily returns. Set an opt-in embeddingModelVersion discriminator (^[A-Za-z0-9._-]{1,40}$) and it is folded into the Vectorize namespace, so bumping the tag re-partitions the index: new writes and reads share a fresh partition and the old vectors become unreachable to new queries (an empty result beats a wrong one). Chunk #0 also stamps the tag in metadata for auditability.

export const docs = defineRag({
    embeddingModel: "@cf/baai/bge-large-en-v1.5",
    embeddingModelVersion: "bge-large-v1.5", // bump when you change embeddingModel
    index: "docs",
});

Leaving embeddingModelVersion unset is byte-identical to before; existing indexes are untouched. The discriminator must live in the namespace (not just the id prefix): a Vectorize id prefix does not partition nearest-neighbour results.

Hybrid search (vector + lexical)

Dense retrieval misses exact keywords, rare tokens, and identifiers. Supply a lexicalStore and retrieve runs a keyword leg alongside the vector leg and fuses the two rankings with Reciprocal Rank Fusion. That combines semantic recall with lexical precision and needs no reranker call.

import { bm25LexicalStore, defineRag } from "@lunora/ai/rag";

export const docs = defineRag({
    embeddingModel: "@cf/baai/bge-base-en-v1.5",
    index: "docs",
    lexicalStore: bm25LexicalStore(), // Okapi BM25 keyword leg
    lexicalTopK: 20, // fanout of the lexical leg (defaults to the query topK)
});

Both legs fetch a deeper candidate pool than topK (topK × 4 by default, tunable with candidates) and the fused list is trimmed to topK afterwards. That widening is what makes the lexical leg useful at all: its job is to surface a chunk the vector leg ranked below topK, which it cannot do if it was never asked for more than topK.

Two lexical stores ship:

  • bm25LexicalStore() — in-memory. It lives in the worker isolate, so it is not durable and not shared across isolates: restart the isolate and the keyword leg silently returns nothing until every source is re-indexed. Use it for tests, local dev, and single-isolate workloads.
  • sqlLexicalStore({ exec }) — a durable inverted index over the same injected SQL executor sqliteVectorStore takes, so hybrid search survives a deploy.
import { defineRag, sqlLexicalStore } from "@lunora/ai/rag";

export const docs = defineRag({
    index: "docs",
    lexicalStore: sqlLexicalStore({ exec }), // the same `exec` as above
});

Attaching a lexical store to a corpus that is already indexed needs one pass with index({ …, reindex: true }): an ordinary re-sync short-circuits on the content hash before it reaches the new store, so the keyword leg would stay empty.

Both score through the same BM25 kernel, so swapping one for the other does not move the ranking — asserted directly in the test suite. Indexing mirrors chunks into the lexical store automatically, and removals fan out to it too.

Reranking

Vector search embeds a passage without ever seeing the query, so it cannot tell which of two topically-similar chunks actually answers this question. A cross-encoder sees both at once. rerank runs over the candidate pool after fusion and before the trim to topK, and its output order is final:

import { defineRag, scoreReranker } from "@lunora/ai/rag";

export const docs = defineRag({
    candidates: 40, // pool the reranker gets to work with (default: topK × 4)
    index: "docs",
    rerank: scoreReranker({
        minScore: 0.2, // also *reject* weak matches, not just reorder
        score: async (query, text) => {
            const result = await ctx.ai.run("@cf/baai/bge-reranker-base", { contexts: [{ text }], query });
            return result.response[0].score;
        },
    }),
});

The scorer is injected — @lunora/ai takes no provider dependency to make a model call. scoreReranker calls it once per candidate with bounded concurrency; batchReranker hands the whole pool to a batch endpoint in one call (and refuses a score list that does not line up with the passages, rather than zipping scores onto the wrong ones). Pass { rerank: false } on a single retrieve() to skip the round-trip on a latency-sensitive path.

Query transformation

The raw user query is often the worst possible search string. A conversational follow-up ("what about the other one?") carries its meaning in the preceding turns, and a short question shares few terms with the long passage answering it.

transformQuery rewrites it — or expands it into several, each searched independently and fused with RRF, recovering passages any single phrasing would miss:

export const docs = defineRag({
    index: "docs",
    // Return a string to rewrite, or an array for multi-query retrieval.
    transformQuery: async (query, { conversationId }) => {
        const { text } = await generateText({
            // A rewritten query is a phrase. Retrieval is reachable from public
            // RPC, so leave this uncapped and every search is an uncapped
            // generation — the same denial-of-wallet shape as an unbounded
            // answer, one hop further back.
            maxOutputTokens: 100,
            model: ctx.ai.model("@cf/meta/llama-3.3-70b-instruct-fp8-fast"),
            prompt: `Rewrite this search query to be self-contained: ${query}`,
        });
        return text;
    },
});

It receives conversationId from the bound context, so a follow-up can be rewritten against its thread. Also injected, for the same reason as rerank: every useful strategy (HyDE, multi-query expansion, follow-up rewriting) needs a language model, and this package does not pick one for you. Returning the query unchanged — or nothing usable — falls back to the original rather than searching for an empty string. Pass { transformQuery: false } to skip it per call.

RLS-filtered retrieval

rlsFilter derives a metadata filter from the retrieval identity so per-request row-level security applies without every call site remembering to pass it. It receives ctx.auth (an action's ctx satisfies RagContext.auth structurally) and returns a Vectorize metadata filter, which is merged over any explicit filter with the RLS keys winning, so a caller can never widen past the tenant/RBAC scope.

export const docs = defineRag({
    embeddingModel: "@cf/baai/bge-base-en-v1.5",
    index: "docs",
    rlsFilter: (auth) => ({ orgId: (auth as { orgId: string }).orgId }),
});

// the retrieval is transparently scoped to the caller's org:
const { chunks } = await docs(ctx).retrieve(question, { topK: 5 });

The filter applies to both the vector and lexical legs, and only to retrieval; indexing stays a trusted server path.

Each chunk's source metadata is mirrored into the lexical store at index time (StoredRagChunk.metadata), so bm25LexicalStore evaluates the same predicate the vector leg is given — hybrid search and metadata-based RLS compose. A lexical hit the filter excludes never reaches fusion, which matters because the filter carries the tenant/RBAC scope: a lexical leg that ignored it would leak excluded chunk text no matter what the vector leg returned.

matchesMetadataFilter is exported if you are writing your own store. It covers implicit equality, $eq / $ne / $lt / $lte / $gt / $gte, $in / $nin, and dot-notation paths into nested objects — and fails closed on anything it does not recognise, including a range predicate over an incomparable value and a chunk indexed with no metadata at all. Guessing at an unknown operator risks admitting a row RLS meant to exclude, and the failure mode there is a cross-tenant leak, not a missing result.

Bulk ingestion — defineRagSource

rag.index() takes one document's text, which leaves the whole crawl — list the objects, fetch each, extract text, index it, notice the ones that were deleted — as something every app writes for itself. defineRagSource is that loop. If you would rather not run it at all, AI Search ingests an R2 bucket or a website on its own schedule, at the cost of the controls on this page.

import { defineRagSource } from "@lunora/ai/rag";

const ingest = defineRagSource(docs(ctx), { namespace: ctx.shardKey });

const report = await ingest.sync(
    {
        list: async function* () {
            for await (const object of bucket.list()) {
                yield { key: object.key, contentType: object.httpMetadata?.contentType, metadata: { url: object.key } };
            }
        },
        get: async (object) => (await bucket.get(object.key))?.text(),
    },
    // Optional. What you already indexed — anything here that this pass does
    // not list is deleted from the index.
    { knownKeys: await loadIndexedKeys() },
);
// → { indexed: [...], unchanged: [...], skipped: [...], pruned: [...] }

The object source is injected, so this runs over an R2 bucket, S3, a filesystem, or a database table without @lunora/ai depending on any of them. list may be an async generator, so the caller decides how to page a large bucket. One pass collects the keys it yields — it needs the full current set to work out what disappeared — while object bodies are fetched concurrency at a time (concurrency, default 4), so peak memory is that many bodies, not one. Lower it when the objects are large.

Re-syncing is free. rag.index short-circuits on a content hash — over the object's body, metadata and importance together — so an unchanged object costs one get and no embedding, which makes running this on a cron the normal way to use it. An object whose metadata changed is not unchanged: it re-indexes, so a tenant move never leaves the old scope on the vectors.

Pruning keeps the index a mirror, and it is the caller's set. Pass knownKeys — the keys you believe the index holds — and anything missing from this pass's list() is removed, because a document deleted upstream but left indexed keeps being retrieved and cited. Omit it and nothing is pruned.

It is explicit because nothing else can hold that set honestly: the shape above builds the source per request, from a per-request docs(ctx), so a set remembered inside the instance would be empty every time — a prune that defaulted on would in practice never run and never say so. Persist the keys (a table, a KV entry, the bucket listing itself) and hand them in.

Extractors are injected too, keyed by content type with "*" as a fallback. Parsing PDF is a large dependency, and pulling one in for everybody to serve the users who need it is the wrong trade. Plain-text types (text/plain, text/markdown, text/csv, application/json) need no extractor; anything else without one is skipped, never indexed as raw bytes — a PDF's binary or an HTML file's markup embeds to something that matches nothing.

Cloudflare AI Search (formerly AutoRAG) runs the whole pipeline on Cloudflare's side: ingestion from an R2 bucket, a website or uploaded files, chunking, embedding, a hybrid vector and BM25 index, reranking and a query cache. An action reaches it as ctx.aiSearch: the raw ai_search_namespaces binding, passed through unwrapped, so every method is Cloudflare's own AiSearchNamespace. It is typed as @lunora/bindings/ai-search's AiSearch, a structural mirror of that class, so it type-checks without @cloudflare/workers-types in your tsconfig.

// lunora/search.ts
import { action } from "./_generated/server";
import { v } from "@lunora/values";

export const askDocs = action.input({ question: v.string() }).action(async ({ args, ctx }) => {
    const { chunks } = await ctx.aiSearch.get("docs").search({
        ai_search_options: { retrieval: { max_num_results: 5 } },
        query: args.question,
    });

    return chunks.map((chunk) => ({ key: chunk.item.key, score: chunk.score, text: chunk.text }));
});

Reading ctx.aiSearch is the whole setup. lunora dev and lunora deploy add { "binding": "AI_SEARCH", "namespace": "default" } to wrangler.jsonc the first time a lunora/ source reads it, and leave any ai_search_namespaces entry you wrote alone. If AI_SEARCH is already the name of another binding (an ai_search instance, a KV namespace, a var), nothing is written and the reconcile warning says which; add your own entry under another name. default exists on every account and already holds the instances you created before AI Search went GA. To bind another namespace, keep your own entry and point the app at it with defineApp().aiSearch((env) => env.MY_NAMESPACE). The instances themselves (source, embedding model, index method, OCR) are configured in the dashboard, with wrangler, or through ctx.aiSearch.create(...).

It is action-only. Every call is billed network I/O over a corpus that re-indexes underneath you, so a query or mutation running it would re-bill on every subscription re-run and return a different answer each time. There is no local simulator either: plain lunora dev proxies the binding to the deployed service (wrangler warns that it may incur charges), so you need to be logged in with wrangler login. On the celld and node targets the capability matrix rates it unsupported, and codegen omits ctx.aiSearch with a platform_unsupported_feature diagnostic.

AI Search or defineRag?

Choose AI Search when…Choose defineRag when…
the corpus is an R2 bucket or a website you want indexed with no codethe corpus is your app's own rows, indexed as they are written
you need OCR for scanned PDFs or the multimodal embedding modelsyou need your own chunker, embedder, lexical or graph leg, reranker or query transform
one tenant per instance is fine, or the corpus is sharedretrieval must be filtered by row-level security across many tenants in one index
Cloudflare is the only targetthe app also runs on celld or node

Two limits decide the tenancy row. An instance has at most 5 custom metadata fields, and a filterable string is indexed only to its first 64 bytes, so a metadata filter on a tenant id longer than that stops filtering correctly without an error. That is a cross-tenant leak, not a missing result. For per-tenant isolation, give each tenant its own instance:

import { internalAction } from "./_generated/server";
import { v } from "@lunora/values";

export const provisionTenant = internalAction.input({ tenantId: v.string() }).action(async ({ args, ctx }) => {
    // Built-in storage: omit `type` and `source`, then upload into it.
    const instance = await ctx.aiSearch.create({ id: `tenant-${args.tenantId}` });

    await instance.items.upload("welcome.md", "# Welcome\n…");
});

Two more things to know before you put it on a hot path:

  • Billing. AI Search is usage-billed from 2026-11-01: past the monthly free allowance, semantic queries cost $0.75 and full-text queries $0.10 per 1,000, and ingestion $0.75 per million tokens. Workers AI embedding and reranking calls are folded into that bill. ctx.ai usage counters do not see these calls.
  • Text only over the binding. A message's content is a string. Image and file content parts, one of the GA headlines, are accepted only by the REST API and the public endpoint, not by ctx.aiSearch.

Streaming an answer

chatCompletions retrieves and generates in one call, and with stream: true it resolves to a ReadableStream of server-sent events. An HTTP route can hand that stream straight back. ctx.aiSearch is not on the HTTP action context, so read the binding off the route's env:

// lunora/http.ts
import type { AiSearch } from "@lunora/bindings/ai-search";
import { httpRouter } from "lunorash/server";

export const app = httpRouter();

app.post("/ask", async (c) => {
    const { question } = await c.req.json<{ question: string }>();
    const search = c.env.AI_SEARCH as AiSearch;
    const stream = await search.get("docs").chatCompletions({ messages: [{ content: question, role: "user" }], stream: true });

    return new Response(stream, { headers: { "content-type": "text/event-stream" } });
});

Across several instances, call chatCompletions (or search) on the namespace itself with ai_search_options: { instance_ids: ["docs", "blog"] } (1 to 10 ids). Each returned chunk then carries its instance_id.

Feeding agent memory

Agent memory consumes defineRag's RetrieveResult. AI Search returns its own chunks, so map them across by hand. There is no adapter to keep in sync:

import type { RetrieveResult } from "@lunora/ai/rag";
import type { AiSearchInstance } from "@lunora/bindings/ai-search";

const toRetrieveResult = (chunks: Awaited<ReturnType<AiSearchInstance["search"]>>["chunks"]): RetrieveResult => ({
    chunks: chunks.map((chunk, index) => ({
        chunkIndex: index,
        id: chunk.id,
        importance: 1,
        metadata: chunk.item.metadata,
        score: chunk.score,
        sourceId: chunk.item.key,
        text: chunk.text,
    })),
    context: chunks.map((chunk) => `[source:${chunk.item.key}]\n${chunk.text}`).join("\n\n"),
    sources: [...new Set(chunks.map((chunk) => chunk.item.key))].map((id) => ({ id })),
});

AiSearchInstance is what ctx.aiSearch.get(name) returns, exported next to AiSearch. Both are types only; the subpath ships no runtime code.

Tagging gateway calls for cost attribution

AI Gateway can only break cost down by keys that were sent with the call, so a dimension you did not tag is one you cannot recover afterwards. Two scopes, because they are fixed at different times:

  • Deployment-scoped — which app, which environment. Set LUNORA_AI_GATEWAY_TAGS to a flat JSON object of string values and every gateway call carries them:

    // wrangler.jsonc
    "vars": { "LUNORA_AI_GATEWAY_TAGS": "{\"app\":\"checkout\",\"env\":\"prod\"}" }
  • Per-call — which feature, which (hashed) user. Pass them at the call site as the tags field of AiGatewayMetadata.

Both land in the same place: the cf-aig-metadata header for bring-your-own providers, and the Workers AI binding's native gateway.metadata option. The built-in functionPath / traceId correlation fields always win over a tag of the same name.

AI Gateway rejects the whole metadata object past AI_GATEWAY_METADATA_MAX_KEYS (5) keys, so the builder trims to that rather than sending something the gateway will throw away. A malformed LUNORA_AI_GATEWAY_TAGS is ignored with a warn-once rather than failing the call — telemetry configuration must not take inference down — and non-string values are dropped rather than coerced, since a number silently becoming "1" is worse than a visibly absent tag.

Cost telemetry without an AI Gateway

Per-request dollar cost used to reach a span only when a Cloudflare AI Gateway put it in providerMetadata. Call the same model through @ai-sdk/openai directly, or run on a non-Cloudflare host, and spend visibility disappeared.

estimateModelCost derives it from token usage and a price table instead. defineRag's embed span and @lunora/agent's generation spans both fall back to it automatically, so a chat turn carries a cost with or without a gateway:

import { estimateModelCost } from "@lunora/ai";

estimateModelCost("text-embedding-3-small", { inputTokens: 1_000_000 }); // → 0.02

An estimate is never presented as a measurement. A provider-reported cost always wins, and the span records which it got:

AttributeValue
gen_ai.usage.costthe cost in USD
lunora.usage.cost.source"provider" or "estimated"

The shipped table is indicative, not authoritative — a hand-maintained snapshot, and providers change prices without warning. It is deliberately small, because a table trying to cover every model is a table that is wrong about most of them; anything it does not cover returns undefined rather than a guess, and an unpriced model yields no attribute rather than a 0 that would quietly sum into a total. Pass your own prices for anything you are invoicing against.

Usage and spend per function

Every model resolved through ctx.ai.model(...) is instrumented: each call gets an ai.generate (or ai.stream) span, and its usage is counted into the function's durable metrics. The counters are grouped by the model id you asked for (anthropic/claude-sonnet-5, not the provider's internal id) and carry the same cost-source rule as above:

CounterValueAttributes
gen_ai.usage.input_tokensinput tokens of one callgen_ai.request.model
gen_ai.usage.output_tokensoutput tokens of one callgen_ai.request.model
gen_ai.usage.costUSD cost of one callgen_ai.request.model, lunora.usage.cost.source

Studio's AI usage page reads them: total spend and tokens, a breakdown by function and by model, and each recent call linked to its trace. A stream's span stays open until the stream finishes, so its duration is the whole generation.

Outside an action

ctx.ai is only wired onto action contexts. In the worker entry, a Durable Object, or a queue / scheduled handler, build the helper directly from the binding:

import { createAi } from "@lunora/ai";

const ai = createAi({ binding: env.AI });

The raw binding escape hatch, ctx.ai.run(model, inputs) (or ai.run(...)), covers Workers-AI-only model families (image, ASR, translation) that aren't surfaced through the AI SDK provider.