AI Vision: Gemini Screenshot, Webpage & Multi-Tab AI Assistant for Chrome
3 ratings
)


Overview
Ask Gemini about screenshots, webpages, articles, products, research, and up to 20 tabs—then run safe approval-based browser tasks.
AI Vision: Your Gemini Screenshot & Browser Assistant Capture part of a webpage, ask questions about the page you are viewing, or analyze supported tabs across one Chrome window using Google's Gemini AI. What it does: - Lets you screenshot any visible part of a website - Sends your selected image and question from the service worker directly to the Gemini API for analysis - Reads supported content from the current tab when you choose The Tab mode - Summarizes, compares, and answers questions across up to 20 tabs in the same Chrome window with All Tabs mode - Provides summaries, explanations, and direct answers about images, articles, documents, products, research, and other webpage content - Lets you choose Balanced, Concise, Formal, Casual, Detailed, or Bullet-oriented responses - Discovers Gemini models that are available to the user's API key - Can open AI Vision without attaching a picture when you click instead of dragging in Capture mode - Keeps Capture as the simple first screen, with The Tab, All Tabs, and Agent Mode under More modes and tools Agent Mode: - Turn Agent Mode on or off directly below the mode selector in Capture, The Tab, or All Tabs - In Capture mode, Agent Mode uses the selected screenshot and acts only in The Tab where the capture started - In The Tab mode, Agent Mode can read, navigate, click, type, and scroll only in that tab - In All Tabs mode, Agent Mode can search, switch tabs, navigate, click, type, and scroll across supported tabs in the starting Chrome window - The Google ADK browser runtime is bundled in the extension, so every planning step rotates through five configured Gemini models without a terminal, Node.js install, companion process, or download - Opening a new tab, going back or forward, and reloading are supported only after approval; new tabs stay in the starting window - All Tabs Agent Mode stops if the source tab moves to another window - Reading, waiting, and scrolling can proceed automatically; every click, text entry, and model-generated navigation requires explicit approval - Sensitive actions such as passwords, credentials, payments, purchases, deletions, uploads, posts, sign-ins, permissions, and acceptance of legal terms are permanently blocked - Every task stops after a 12-step safety limit so you can review what happened How to use: - Right-click on a page or click the extension icon to open AI Vision - Capture is ready as soon as AI Vision opens; use More modes and tools for The Tab, All Tabs, or Agent Mode - In Capture mode, drag to select the area you want to analyze; click once to open AI Vision without a picture - Type a question or use Summarize, Explain, or Answer for an AI-powered response - Optionally expand preferences in Settings to choose a response style or available Gemini model - Your mode, Agent Mode, and response settings remain selected until you manually change them Getting started: The only setup is a Google Gemini API key from Google AI Studio. Visit aistudio.google.com/app/apikey, create a key, then paste it into AI Vision's Settings menu. Capture needs no other setup; API availability, free-tier limits, model access, and pricing are controlled by Google. Privacy and control: - AI Vision runs only when you open it or start a task - Your API key is stored by the service worker with preferences locally in your Chrome profile; the panel receives only masked key status - Requests go directly from the extension to Google's Gemini API over HTTPS - AI Vision has no developer-operated analytics, advertising, tracking, or proxy server - Chrome internal pages, the Chrome Web Store, and other restricted pages cannot be analyzed Source and support: The source is publicly available at github.com/stiwarilbj/AI_Vision. If AI Vision saves you time, please leave an honest rating on the Chrome Web Store. Ratings are never required or rewarded. AI Vision is an independent project and is not affiliated with or endorsed by Google. Gemini and Chrome are trademarks of Google LLC.
5 out of 53 ratings
Details
- Version2.5
- UpdatedSeptember 3, 2026
- Offered byGitchub
- Size32.17MiB
- LanguagesEnglish
- DeveloperShantanu Tiwari
Arthur St Framingham, MA 01702 USEmail
gitchub@gmail.com - Non-traderThis developer has not identified itself as a trader. For consumers in the European Union, please note that consumer rights do not apply to contracts between you and this developer.
Privacy
AI Vision: Gemini Screenshot, Webpage & Multi-Tab AI Assistant for Chrome has disclosed the following information regarding the collection and usage of your data. More detailed information can be found in the developer's privacy policy.
AI Vision: Gemini Screenshot, Webpage & Multi-Tab AI Assistant for Chrome handles the following:
This developer declares that your data is
- Not being sold to third parties, outside of the approved use cases
- Not being used or transferred for purposes that are unrelated to the item's core functionality
- Not being used or transferred to determine creditworthiness or for lending purposes
Support
For help with questions, suggestions, or problems, visit the developer's support site