Page to Markdown Scraper
Overview
Capture page content and OCR text from any page or PDF. Saves as Markdown/ZIP for ChatGPT and LLMs. 100% local processing.
Page to Markdown Scraper turns any web page into clean Markdown, ready to paste or upload into ChatGPT, Claude, Gemini or any other AI tool. VERSION 2.0 - FASTER TO USE, SAFER BY DESIGN - Copy as Markdown: one click puts the page on your clipboard, ready to paste into your AI chat - Save selection: right-click any selected text and choose "Save selection as Markdown" - Single .md file or ZIP: choose a plain Markdown file, or a ZIP with all images included - Token estimate: see roughly how many tokens a page will use before you paste it - Danish OCR: text recognition now reads Danish (æ, ø, å) as well as English - Better Markdown: links, nested lists, tables, code blocks and images in the right place; menus, footers and hidden text are left out - Safer and lighter: fewer permissions, and nothing runs on your pages unless you capture THREE CAPTURE MODES MANUAL MODE Capture the current page with one click as a ZIP or a single .md file, or copy it straight to your clipboard. AUTO MODE Start a recording session and browse normally. Every page you visit in that window is captured automatically. Other windows are never recorded. A red border and REC badge show that recording is active. Stop the session to download everything in one ZIP. OCR MODE Extract text from what's on screen using built-in text recognition. Works where normal capture can't: images, scanned documents, PDFs, canvas elements and pages that block text selection. - OCR Viewport: extract text from the visible area - OCR Full Page: scrolls through the whole page or PDF and reads each section - Choose English, Danish, or both - Keyboard shortcut: Ctrl+Shift+O (Cmd+Shift+O on Mac) - Results panel with word count, confidence score and a one-click download button PRIVACY - All processing happens locally in your browser. Your captured content is never sent to us or anyone else. - The only network requests are for downloading a captured page's images from the website that hosts them. - OCR runs via Tesseract.js (WebAssembly) entirely on your device. No cloud APIs. - No analytics, no tracking, no accounts, no data collection. - Open source: https://github.com/SlambertDK/page-to-markdown-extension OUTPUT - content.md: clean Markdown optimised for AI tools, with a token estimate - images/ folder: all page images and OCR screenshots - README.txt: usage instructions - Everything in a single ZIP, or a single .md file if you prefer DISCLAIMER You are responsible for making sure you have the right to capture and use any content. A terms notice is shown before every capture. Requires Chrome 116 or newer.
0 out of 5No ratings
Details
- Version2.0.1
- UpdatedOctober 1, 2026
- Offered byhenriklambert1979
- Size4.69MiB
- LanguagesEnglish
- Developer
Email
henriklambert@proton.me - Non-traderThis developer has not identified itself as a trader. For consumers in the European Union, please note that consumer rights do not apply to contracts between you and this developer.
Privacy
This developer declares that your data is
- Not being sold to third parties, outside of the approved use cases
- Not being used or transferred for purposes that are unrelated to the item's core functionality
- Not being used or transferred to determine creditworthiness or for lending purposes