Downloads web pages and extracts the main article text and metadata (title, author, date) without menus, ads, footers or other boilerplate, as plain text, Markdown, JSON, CSV or XML. Use when a user asks to extract article text from a URL, get clean text from HTML, turn a web page into Markdown for an LLM, build a text corpus from a site, list URLs from a sitemap or RSS feed, scrape blog posts with their publication dates, or mentions Trafilatura.