How to explain a screenshot with AI Vision

By AI Vision project · Updated

To explain a screenshot in Chrome, install AI Vision, save your Gemini API key, select the relevant area, and ask Gemini what it shows. The answer stays with the capture so you can ask a follow-up question.

Public Store v2.5: The live install opens with Capture. The GitHub-build v2.8 preview has newer labels such as Screenshot and Select an area; use the public controls below when following this guide.

Capture the detail that matters

First, get a Gemini key and save it in AI Vision. Open a supported webpage and use the toolbar icon or Alt + Shift + V to open the extension.

  1. Leave the default Capture mode selected. Drag over the chart, error message, or image detail you want Gemini to see. A single click opens the panel without a picture if you want to ask about the page first.
  2. Include the full labels, units, and nearby context that affect the answer. Text on Chrome internal pages and the Chrome Web Store cannot be captured.
  3. Type a question, or choose a quick action such as Explain, Summarize, or Answer. Wait for the response before sending another request.
  4. Read the answer against the selected pixels. Use Copy for a reusable response, then Follow up to ask about one detail or correct a mistaken reading.

GitHub-build preview: The unreleased v2.8 screenshots on this site show a compact Screenshot selector, Retake, and a direct follow-up field. The Store listing currently serves v2.5, so its labels and layout may differ.

AI Vision answer state showing a selected chart, an explanation, and a follow-up field
Illustrative GitHub-build v2.8 answer state. The public Store listing currently serves v2.5; the workflow is the same, but control labels may differ.

Worked example: explain a reading chart

Imagine a chart with weekly totals of 40, 60, 90, and 100 pages read. Capture the bars and their week labels, then ask:

“Which week had the biggest increase? Use the numbers shown and explain the calculation.”

Illustrative answer: The biggest increase is from week 2 to week 3: 90 − 60 = 30 pages. Week 4 has the highest total, but its increase is only 100 − 90 = 10 pages.

This distinction makes a better question than “What is this?” The highest total and the largest increase are different things. Follow up with “What would you need to know before calling this a lasting trend?” to surface missing context.

Extract text from an image

GitHub-build v2.8 preview: Use Extract text under More actions when a page shows text as an image. In the public Store v2.5 build, ask Gemini to transcribe the selected text with a typed question, then use the available Copy action. For paired source and transcription examples, read the screenshot-to-text guide.

“Transcribe the visible text in reading order. Preserve names, numbers, and punctuation. Mark unreadable words instead of guessing.”

Check every important number or name. AI text extraction can make mistakes, especially with handwriting, small type, or low contrast.

Understand an error message

Capture the full message with enough surrounding context to identify the affected feature. Ask for an explanation before asking for a fix:

“Explain this error in plain English. List two checks I can make without changing settings or running commands.”

If the capture does not work

Choose a different context when it helps

A screenshot is useful for a visual detail. For an entire article, summarize the webpage instead. If you only need editable words, follow the screenshot text extraction guide. For information spread across sources, compare Chrome tabs. Your selected screenshot and question are sent to Google for the response; review the privacy notice before sharing sensitive content.

Add AI Vision to Chrome ↗Back to AI Vision
AI Vision interface preview