Search IOCombats

Search challenges, guides, questions and articles

Chrome Built-in AI in Practice: Running Gemini Nano On-Device with the Prompt and Summarizer APIs
Chrome Built-in AIPrompt APISummarizer APIOn-device AIProgressive EnhancementFrontend Engineering

Chrome Built-in AI in Practice: Running Gemini Nano On-Device with the Prompt and Summarizer APIs

By Ghazi Khan | Sep 28, 2026 - 10 min read

Most "AI feature" tutorials start with the same line: call an LLM API from your server. That works, but it has real costs. Every request is billed per token, user text leaves the device, latency includes a network round trip, and the feature stops working offline. Chrome now offers another option. It ships a small language model with the browser and exposes it through JavaScript APIs, so your page can summarize, translate and prompt a model that runs on the user's own machine.

This post explains how those APIs work, what the requirements really are, and how to build a feature around them without breaking for the large share of users whose device cannot run the model. If you can explain this architecture in an interview, you show that you understand AI features as a frontend systems problem, not only as an API call.

What "built-in AI" actually means

Built-in AI means the browser downloads a foundation model once, stores it on disk, and shares it across every origin. Your site never bundles model weights. You ask the browser for a capability, and the browser decides whether it can serve it.

According to Chrome's built-in AI documentation, the status of the APIs is not uniform. At the time of writing, the Translator, Language Detector and Summarizer APIs are stable from Chrome 138. The Prompt API is stable for web pages in Chrome 148 and for extensions from Chrome 138. The Writer and Rewriter APIs are in developer trial, and the Proofreader API is in an origin trial. Treat the second group as experiments, not production dependencies.

APIPurposeStatus (per Chrome docs)
SummarizerCondense long textStable, Chrome 138+
TranslatorTranslate between languagesStable, Chrome 138+
Language DetectorIdentify input languageStable, Chrome 138+
PromptFree-form prompting of the modelStable on web in Chrome 148
Writer / RewriterGenerate or refine textDeveloper trial
ProofreaderInteractive proofreadingOrigin trial

The hardware reality

Chrome documents concrete requirements for the on-device model (see the Summarizer API docs): desktop Windows 10 or 11, macOS 13 or newer, Linux, or Chromebook Plus. It needs at least 22 GB of free space on the volume holding the Chrome profile, and either a GPU with strictly more than 4 GB of VRAM or a CPU path with 16 GB of RAM and 4 or more cores. Mobile is not supported yet. Firefox and Safari do not implement these APIs.

This single paragraph drives your whole design. A large share of your users, especially on mid-range Indian phones and older laptops, will never have the model. So built-in AI is a progressive enhancement. It is never the only path.

Diagram
100%
flowchart TD A[Page calls availability] --> B{Result} B -->|unavailable| C[Fallback: server API or hide feature] B -->|downloadable| D[Ask user, then create with monitor] B -->|downloading| E[Show progress, wait] B -->|available| F[Create session and run] D --> E E --> F
visualized byIOCombats

The four availability states

Every built-in API follows the same lifecycle. You call availability() first, and it resolves to one of four strings.

  • unavailable: the device or browser cannot run the model, or the requested options (for example a language) are not supported.
  • downloadable: the device can run it, but the model is not on disk yet.
  • downloading: a download is in progress.
  • available: ready to use immediately.

You never guess. You branch on this value. Here is a small reusable gate that returns a status and keeps the rest of the app simple.

type AiStatus = 'unavailable' | 'downloadable' | 'downloading' | 'available';

export async function summarizerStatus(): Promise<AiStatus> {
  if (!('Summarizer' in self)) return 'unavailable';
  return Summarizer.availability();
}

The 'Summarizer' in self check matters. In Safari and Firefox the global does not exist, and calling it would throw a ReferenceError. Feature detection always comes before availability.

Summarizer API step by step

The Summarizer API is the simplest entry point because it is purpose-built. You do not write a prompt. You describe the kind of summary you want and the browser handles the instructions.

The create() options documented by Chrome include type (key-points, tldr, teaser or headline), format (markdown or plain-text), length (short, medium or long), and sharedContext, a string that tells the model what kind of content it is reading.

export async function createSummarizer(onProgress: (pct: number) => void) {
  return Summarizer.create({
    type: 'key-points',
    format: 'markdown',
    length: 'short',
    sharedContext: 'Technical blog posts about frontend engineering',
    monitor(m) {
      m.addEventListener('downloadprogress', (e) => onProgress(e.loaded));
    },
  });
}

Two details deserve attention. First, the monitor callback gives you a downloadprogress event, which is how you show a progress bar during the first-time model download. That download can be large, so telling users what is happening is both courteous and, per Chrome's guidance, recommended. Second, create() that triggers a download requires user activation, meaning it must run inside a click or key handler. Calling it from a useEffect on page load will fail when the model still has to be fetched.

Streaming the result

A summary can take seconds on-device. summarizeStreaming() returns a ReadableStream, and you can consume it with for await, exactly like the streaming chat pattern used with server LLMs.

export async function streamSummary(
  summarizer: Summarizer,
  text: string,
  onChunk: (soFar: string) => void,
  signal?: AbortSignal,
) {
  let result = '';
  const stream = summarizer.summarizeStreaming(text, { signal });
  for await (const chunk of stream) {
    result += chunk;
    onChunk(result);
  }
  return result;
}

Passing an AbortSignal lets you cancel when the user navigates away. Without it, the model keeps computing on the user's GPU for a result nobody will see.

The Prompt API: sessions, context and structured output

The Prompt API gives you the raw model. It is more flexible and easier to misuse, so understanding its session model is important.

A session holds conversation state. Each prompt() call appends the user message and the model reply to that session's context, so a follow-up question can refer to the previous answer. The context window is finite, which means long conversations eventually lose their oldest turns. A session is also a real resource that holds memory, so you should destroy sessions you no longer need.

const controller = new AbortController();

const availability = await LanguageModel.availability({
  expectedInputs: [{ type: 'text', languages: ['en'] }],
  expectedOutputs: [{ type: 'text', languages: ['en'] }],
});

if (availability === 'unavailable') throw new Error('No on-device model');

const session = await LanguageModel.create({ signal: controller.signal });
const answer = await session.prompt(
  'Explain event delegation in two sentences.',
);

Declaring expectedInputs and expectedOutputs with languages is not decoration. Availability can differ per language, so you ask about the exact combination you will use.

Cloning a session for branching work

session.clone() creates a new session with the same accumulated context. This is useful when you want one expensive system setup and many independent questions. You set up context once, then clone per task so tasks do not pollute each other.

Structured output with a JSON Schema

Small models are unreliable at producing parseable text when you merely ask nicely. The Prompt API accepts a responseConstraint, a JSON Schema that constrains the model's output.

const schema = {
  type: 'object',
  properties: {
    sentiment: { enum: ['positive', 'neutral', 'negative'] },
    confidence: { type: 'number' },
  },
  required: ['sentiment', 'confidence'],
};

const raw = await session.prompt(
  'Classify the sentiment of: "The new build is fast but the docs are thin."',
  { responseConstraint: schema },
);

const result = JSON.parse(raw) as {
  sentiment: 'positive' | 'neutral' | 'negative';
  confidence: number;
};

Constrained decoding means the model can only emit tokens that keep the output valid against the schema, so JSON.parse is far safer than parsing free text. You should still wrap it in try/catch and validate with a runtime schema library, because a type assertion proves nothing at runtime. The Prompt API also supports image and audio inputs, with audio requiring GPU acceleration, but text covers most frontend use cases.

Diagram
100%
sequenceDiagram participant UI as React UI participant API as LanguageModel participant Model as Gemini Nano on disk UI->>API: availability(options) API-->>UI: available UI->>API: create(signal) API->>Model: load into memory UI->>API: prompt(text, responseConstraint) API->>Model: constrained decoding Model-->>API: schema-valid JSON API-->>UI: string result UI->>API: session.destroy()
visualized byIOCombats

A production pattern: one interface, two engines

Because the model is optional, define your feature against an interface and plug in either engine. This is the Strategy pattern, and it keeps components unaware of where the intelligence runs.

export interface Summarize {
  run(
    text: string,
    onChunk: (s: string) => void,
    signal?: AbortSignal,
  ): Promise<string>;
}

class OnDeviceSummarizer implements Summarize {
  constructor(private inner: Summarizer) {}
  run(text: string, onChunk: (s: string) => void, signal?: AbortSignal) {
    return streamSummary(this.inner, text, onChunk, signal);
  }
}

class ServerSummarizer implements Summarize {
  async run(text: string, onChunk: (s: string) => void, signal?: AbortSignal) {
    const res = await fetch('/api/summarize', {
      method: 'POST',
      body: JSON.stringify({ text }),
      signal,
    });
    const data = await res.json();
    onChunk(data.summary);
    return data.summary as string;
  }
}

export async function pickSummarizer(): Promise<Summarize | null> {
  if ((await summarizerStatus()) === 'available') {
    return new OnDeviceSummarizer(await createSummarizer(() => {}));
  }
  return new ServerSummarizer();
}

A component calls pickSummarizer() once and uses the result. If you have no server fallback, return null and hide the button entirely. A missing button is better than a button that errors.

Diagram
100%
flowchart LR C[Component] --> P[pickSummarizer] P -->|model available| O[OnDeviceSummarizer] P -->|otherwise| S[ServerSummarizer] O --> R[Streamed text] S --> R
visualized byIOCombats

Trade-offs you should be able to defend

ConcernOn-device (built-in)Server LLM
PrivacyText stays on the deviceText sent to a provider
Cost per requestNone for youBilled per token
OfflineWorks after model downloadNeeds network
ReachNarrow, hardware gated, Chrome and EdgeEvery browser
QualitySmall model, weaker on hard tasksLarger, stronger models
First useLarge one-time downloadNone

Notice what this table implies. On-device models suit small, private, high-volume tasks: summarizing a page, classifying a comment, translating a UI string, cleaning up a draft. They do not suit complex reasoning or code generation. Picking the right tool for the task is the engineering judgment interviewers look for.

Also remember access rules. The Summarizer docs state it works in top-level windows and same-origin iframes, cross-origin iframes need the allow="summarizer" permission policy, and Web Workers are not currently supported. If you planned to run inference off the main thread, check current support first.

Practical takeaway

When you add a built-in AI feature at work, follow this order. Detect the global with in self. Call availability() with the exact options you will use. Only trigger create() from a user gesture, and show download progress through monitor. Pass an AbortSignal everywhere. Destroy sessions on unmount. Put both engines behind one interface, and hide the feature when neither exists.

In an interview, if asked "how would you add summarization to a content site while keeping user data private and costs low?", a strong answer names on-device models as the first tier, explains the hardware gating and the four availability states, and describes a server fallback behind a shared interface. That answer shows system thinking, not tool familiarity.

Conclusion

Chrome's built-in AI turns a language model into a browser capability, with the cost and privacy profile of local code and the reach limits of a hardware-gated feature. The APIs are small, but the engineering lives in the edges: availability states, user activation, streaming, cancellation and graceful fallback. Build the fallback first, add the on-device path as an enhancement, and you get an AI feature that is cheap, private when it can be, and never broken when it cannot.

Sources and further reading

API status and requirements change often, so check these pages before shipping.

Advertisement

Ready to practice?

Test your skills with our interactive UI challenges and build your portfolio.

Start Coding Challenge