
Chrome Built-in AI in Practice: Running Gemini Nano On-Device with the Prompt and Summarizer APIs
By Ghazi Khan | Sep 28, 2026 - 10 min read
Most "AI feature" tutorials start with the same line: call an LLM API from your server. That works, but it has real costs. Every request is billed per token, user text leaves the device, latency includes a network round trip, and the feature stops working offline. Chrome now offers another option. It ships a small language model with the browser and exposes it through JavaScript APIs, so your page can summarize, translate and prompt a model that runs on the user's own machine.
This post explains how those APIs work, what the requirements really are, and how to build a feature around them without breaking for the large share of users whose device cannot run the model. If you can explain this architecture in an interview, you show that you understand AI features as a frontend systems problem, not only as an API call.
What "built-in AI" actually means
Built-in AI means the browser downloads a foundation model once, stores it on disk, and shares it across every origin. Your site never bundles model weights. You ask the browser for a capability, and the browser decides whether it can serve it.
According to Chrome's built-in AI documentation, the status of the APIs is not uniform. At the time of writing, the Translator, Language Detector and Summarizer APIs are stable from Chrome 138. The Prompt API is stable for web pages in Chrome 148 and for extensions from Chrome 138. The Writer and Rewriter APIs are in developer trial, and the Proofreader API is in an origin trial. Treat the second group as experiments, not production dependencies.
| API | Purpose | Status (per Chrome docs) |
|---|---|---|
| Summarizer | Condense long text | Stable, Chrome 138+ |
| Translator | Translate between languages | Stable, Chrome 138+ |
| Language Detector | Identify input language | Stable, Chrome 138+ |
| Prompt | Free-form prompting of the model | Stable on web in Chrome 148 |
| Writer / Rewriter | Generate or refine text | Developer trial |
| Proofreader | Interactive proofreading | Origin trial |
The hardware reality
Chrome documents concrete requirements for the on-device model (see the Summarizer API docs): desktop Windows 10 or 11, macOS 13 or newer, Linux, or Chromebook Plus. It needs at least 22 GB of free space on the volume holding the Chrome profile, and either a GPU with strictly more than 4 GB of VRAM or a CPU path with 16 GB of RAM and 4 or more cores. Mobile is not supported yet. Firefox and Safari do not implement these APIs.
This single paragraph drives your whole design. A large share of your users, especially on mid-range Indian phones and older laptops, will never have the model. So built-in AI is a progressive enhancement. It is never the only path.
Diagram100%flowchart TD A[Page calls availability] --> B{Result} B -->|unavailable| C[Fallback: server API or hide feature] B -->|downloadable| D[Ask user, then create with monitor] B -->|downloading| E[Show progress, wait] B -->|available| F[Create session and run] D --> E E --> Fvisualized by
The four availability states
Every built-in API follows the same lifecycle. You call availability() first, and it resolves to one of four strings.
unavailable: the device or browser cannot run the model, or the requested options (for example a language) are not supported.downloadable: the device can run it, but the model is not on disk yet.downloading: a download is in progress.available: ready to use immediately.
You never guess. You branch on this value. Here is a small reusable gate that returns a status and keeps the rest of the app simple.
type AiStatus = 'unavailable' | 'downloadable' | 'downloading' | 'available';
export async function summarizerStatus(): Promise<AiStatus> {
if (!('Summarizer' in self)) return 'unavailable';
return Summarizer.availability();
}
The 'Summarizer' in self check matters. In Safari and Firefox the global does not exist, and calling it would throw a ReferenceError. Feature detection always comes before availability.
Summarizer API step by step
The Summarizer API is the simplest entry point because it is purpose-built. You do not write a prompt. You describe the kind of summary you want and the browser handles the instructions.
The create() options documented by Chrome include type (key-points, tldr, teaser or headline), format (markdown or plain-text), length (short, medium or long), and sharedContext, a string that tells the model what kind of content it is reading.
export async function createSummarizer(onProgress: (pct: number) => void) {
return Summarizer.create({
type: 'key-points',
format: 'markdown',
length: 'short',
sharedContext: 'Technical blog posts about frontend engineering',
monitor(m) {
m.addEventListener('downloadprogress', (e) => onProgress(e.loaded));
},
});
}
Two details deserve attention. First, the monitor callback gives you a downloadprogress event, which is how you show a progress bar during the first-time model download. That download can be large, so telling users what is happening is both courteous and, per Chrome's guidance, recommended. Second, create() that triggers a download requires user activation, meaning it must run inside a click or key handler. Calling it from a useEffect on page load will fail when the model still has to be fetched.
Streaming the result
A summary can take seconds on-device. summarizeStreaming() returns a ReadableStream, and you can consume it with for await, exactly like the streaming chat pattern used with server LLMs.
export async function streamSummary(
summarizer: Summarizer,
text: string,
onChunk: (soFar: string) => void,
signal?: AbortSignal,
) {
let result = '';
const stream = summarizer.summarizeStreaming(text, { signal });
for await (const chunk of stream) {
result += chunk;
onChunk(result);
}
return result;
}
Passing an AbortSignal lets you cancel when the user navigates away. Without it, the model keeps computing on the user's GPU for a result nobody will see.
The Prompt API: sessions, context and structured output
The Prompt API gives you the raw model. It is more flexible and easier to misuse, so understanding its session model is important.
A session holds conversation state. Each prompt() call appends the user message and the model reply to that session's context, so a follow-up question can refer to the previous answer. The context window is finite, which means long conversations eventually lose their oldest turns. A session is also a real resource that holds memory, so you should destroy sessions you no longer need.
const controller = new AbortController();
const availability = await LanguageModel.availability({
expectedInputs: [{ type: 'text', languages: ['en'] }],
expectedOutputs: [{ type: 'text', languages: ['en'] }],
});
if (availability === 'unavailable') throw new Error('No on-device model');
const session = await LanguageModel.create({ signal: controller.signal });
const answer = await session.prompt(
'Explain event delegation in two sentences.',
);
Declaring expectedInputs and expectedOutputs with languages is not decoration. Availability can differ per language, so you ask about the exact combination you will use.
Cloning a session for branching work
session.clone() creates a new session with the same accumulated context. This is useful when you want one expensive system setup and many independent questions. You set up context once, then clone per task so tasks do not pollute each other.
Structured output with a JSON Schema
Small models are unreliable at producing parseable text when you merely ask nicely. The Prompt API accepts a responseConstraint, a JSON Schema that constrains the model's output.
const schema = {
type: 'object',
properties: {
sentiment: { enum: ['positive', 'neutral', 'negative'] },
confidence: { type: 'number' },
},
required: ['sentiment', 'confidence'],
};
const raw = await session.prompt(
'Classify the sentiment of: "The new build is fast but the docs are thin."',
{ responseConstraint: schema },
);
const result = JSON.parse(raw) as {
sentiment: 'positive' | 'neutral' | 'negative';
confidence: number;
};
Constrained decoding means the model can only emit tokens that keep the output valid against the schema, so JSON.parse is far safer than parsing free text. You should still wrap it in try/catch and validate with a runtime schema library, because a type assertion proves nothing at runtime. The Prompt API also supports image and audio inputs, with audio requiring GPU acceleration, but text covers most frontend use cases.
Diagram100%sequenceDiagram participant UI as React UI participant API as LanguageModel participant Model as Gemini Nano on disk UI->>API: availability(options) API-->>UI: available UI->>API: create(signal) API->>Model: load into memory UI->>API: prompt(text, responseConstraint) API->>Model: constrained decoding Model-->>API: schema-valid JSON API-->>UI: string result UI->>API: session.destroy()visualized by
A production pattern: one interface, two engines
Because the model is optional, define your feature against an interface and plug in either engine. This is the Strategy pattern, and it keeps components unaware of where the intelligence runs.
export interface Summarize {
run(
text: string,
onChunk: (s: string) => void,
signal?: AbortSignal,
): Promise<string>;
}
class OnDeviceSummarizer implements Summarize {
constructor(private inner: Summarizer) {}
run(text: string, onChunk: (s: string) => void, signal?: AbortSignal) {
return streamSummary(this.inner, text, onChunk, signal);
}
}
class ServerSummarizer implements Summarize {
async run(text: string, onChunk: (s: string) => void, signal?: AbortSignal) {
const res = await fetch('/api/summarize', {
method: 'POST',
body: JSON.stringify({ text }),
signal,
});
const data = await res.json();
onChunk(data.summary);
return data.summary as string;
}
}
export async function pickSummarizer(): Promise<Summarize | null> {
if ((await summarizerStatus()) === 'available') {
return new OnDeviceSummarizer(await createSummarizer(() => {}));
}
return new ServerSummarizer();
}
A component calls pickSummarizer() once and uses the result. If you have no server fallback, return null and hide the button entirely. A missing button is better than a button that errors.
Diagram100%flowchart LR C[Component] --> P[pickSummarizer] P -->|model available| O[OnDeviceSummarizer] P -->|otherwise| S[ServerSummarizer] O --> R[Streamed text] S --> Rvisualized by
Trade-offs you should be able to defend
| Concern | On-device (built-in) | Server LLM |
|---|---|---|
| Privacy | Text stays on the device | Text sent to a provider |
| Cost per request | None for you | Billed per token |
| Offline | Works after model download | Needs network |
| Reach | Narrow, hardware gated, Chrome and Edge | Every browser |
| Quality | Small model, weaker on hard tasks | Larger, stronger models |
| First use | Large one-time download | None |
Notice what this table implies. On-device models suit small, private, high-volume tasks: summarizing a page, classifying a comment, translating a UI string, cleaning up a draft. They do not suit complex reasoning or code generation. Picking the right tool for the task is the engineering judgment interviewers look for.
Also remember access rules. The Summarizer docs state it works in top-level windows and same-origin iframes, cross-origin iframes need the allow="summarizer" permission policy, and Web Workers are not currently supported. If you planned to run inference off the main thread, check current support first.
Practical takeaway
When you add a built-in AI feature at work, follow this order. Detect the global with in self. Call availability() with the exact options you will use. Only trigger create() from a user gesture, and show download progress through monitor. Pass an AbortSignal everywhere. Destroy sessions on unmount. Put both engines behind one interface, and hide the feature when neither exists.
In an interview, if asked "how would you add summarization to a content site while keeping user data private and costs low?", a strong answer names on-device models as the first tier, explains the hardware gating and the four availability states, and describes a server fallback behind a shared interface. That answer shows system thinking, not tool familiarity.
Conclusion
Chrome's built-in AI turns a language model into a browser capability, with the cost and privacy profile of local code and the reach limits of a hardware-gated feature. The APIs are small, but the engineering lives in the edges: availability states, user activation, streaming, cancellation and graceful fallback. Build the fallback first, add the on-device path as an enhancement, and you get an AI feature that is cheap, private when it can be, and never broken when it cannot.
Sources and further reading
- Built-in AI APIs overview, Chrome for Developers. API status table and general guidance.
- Summarize with built-in AI, Chrome for Developers. Hardware requirements,
create()options and streaming. - Prompt API, Chrome for Developers. Sessions,
responseConstraintand multimodal input. - Translation with built-in AI, Chrome for Developers. The Translator API, if you want to extend the pattern in this post.
- AI APIs in stable and origin trials, Chrome blog. Background on how the APIs moved through trials.
API status and requirements change often, so check these pages before shipping.
Advertisement
Ready to practice?
Test your skills with our interactive UI challenges and build your portfolio.
Start Coding Challenge