MCP topic guide
Web Scraping MCP Servers
Turn URLs into structured content for agents. Compare scraping, crawling, and search servers with indexed tool lists and setup steps.
Showing 60 of 2131 matching servers from our catalog. Search full directory →
SparkForge
Access a versatile collection of tools for media generation, web scraping, and cryptocurrency research. Convert between data formats, summarize long documents, and extract text from images or PDFs with ease. Streamline your workflow by automating tasks like SEO metadata extraction, code reviews, and content creation.
Paper Search
Search and download academic papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic Scholar, and IACR. Fetch PDFs and extract full text to accelerate literature reviews. Get consistent metadata for easier filtering, citation, and analysis.
paper-search-mcp-openai-v2
Find and download academic papers from leading sources like arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic Scholar, CrossRef, and IACR. Get standardized results and fetch full-text PDFs when available. Accelerate literature reviews with deep search and effortless retrieval.
ia-qa.com/mcp llm and RAG testing - Dev/QA toolbox
IA-QA as an MCP Server Use IA-QA's developer tools to test your llm agents RAG ai tools, directly from Cursor, Claude Desktop, Windsurf, or any AI agent without leaving your IDE. Many classical testing tools too ! Enjoy ! No API key. No signup. Free.
Apify
[Actors MCP Server](https://apify.com/apify/actors-mcp-server): Use 3,000+ pre-built cloud tools to extract data from websites, e-commerce, social media, search engines, maps, and more
x711io - universal gas station market and intelligence
Pay-per-call tool API for autonomous agents. x402 payments on Base. 27 tools including web search, price feeds, and The Hive shared memory layer.
Gapup MCP — 270+ Agent-payable AI Tools
270+ agent-payable C-suite expertises (competitive intel, SEC filings, sanctions multi-jurisdiction, KYC, pentest scope, clinical evidence GRADE-graded, real-estate deals EU-first, ESG, NAICS classifier, patent landscape, research paper Q&A). x402 USDC/EURC micro-payments. Free tier 100 calls/mo, no credit card. Dual-audience (human/agent).
Framesail
Create long-form (faceless YouTube) videos end to end from any MCP client: script, locked character references, storyboard, voiceover, and final video editing — with characters and style held consistent across every shot. Making long-form AI video today means 8+ tabs stitched by hand — an LLM for the script, a voice model, an image model, a video model — with characters drifting between tools and style resetting at every export. Framesail replaces the patchwork: the whole pipeline runs in one place and manages your video's context end to end. Six stages: Style (paste images, videos, or YouTube links and Framesail reverse-engineers the look, voice, and direction), Script (write it yourself or generate it in your narrative style), Reference images (auto-generated for every character, place, and prop), Voiceover (one narrator or many characters, with word-level timing), Storyboard (planned scene by scene), and Editor (captions, music, SFX, then export). No black box: you control every prompt, asset, model, and setting.
Receiptor MCP
Receiptor connects AI agents to your bookkeeping workspace so they can find, review, and organize financial documents such as receipts, bills, and invoices. Use it to inspect workspace context, monitor intake sources, review extracted document data, check integrations, and prepare accountant-ready workflows through secure OAuth access to Receiptor AI.
The STALL
293 AI-callable finance and data tools over MCP. No API keys required. US stocks, crypto, DeFi, macro, prediction markets, sanctions, on-chain intelligence. x402 USDC pay-per-call on Base. Built by IntuiTek1.
AI Answer Copier
AI Answer Copier is a Model Context Protocol (MCP) server that solves the "Final Mile" friction in educational content creation. It enables AI models to move beyond just writing questions to actually generating the files required for teaching and assessment. By functioning as a native MCP server, this tool allows Claude or any MCP-enabled IDE to directly pipe its output into specialized educational formats. No more manually fixing bullet points in Word or wrestling with CSV headers for your LMS. Key Capabilities:Zero-Friction Pipeline: Your AI can now "see" your local export tools. Ask it to: "Generate 10 biology questions and send them directly to my Quizizz CSV." Multi-Format Exporting: Seamlessly convert AI responses into upload-ready files for Kahoot, Quizizz, Canvas (JSON), Moodle (XML), and professionally formatted PDFs. Intelligent Smart-Parse: Automatically identifies question stems, multiple-choice distractors, and correct answer keys from raw AI text. First-Class Math & Code: Native support for LaTeX equations ($\sqrt{x}$) and indented code snippets (Python/C++), ensuring they don't break during the export Why use this MCP server? Generating questions takes seconds, but formatting them takes hours. This server reclaims those 5 hours of your week by removing the technical barrier between AI intelligence and classroom delivery.
Brave Search
Search the web with Brave's independent index — web, news, images, and videos. Bring your own subscription token from the [Brave Search API dashboard](https://api-dashboard.search.brave.com).
ToolSnap MCP
Context-efficient microtools for AI agents. 33 tools total. Flagship fetch_extract: 98.1% median token reduction (53,820 → 2,001 tokens) — saves ~$0.156/call at Sonnet pricing vs loading raw HTML. Paid tools: web extraction, full-page screenshots, background removal (fal.ai rembg), keyword research (DataForSEO), PDF text extract, HTML-to-Markdown, page asset/link inventory, temporary file upload. 20 always-free utility tools: UUID, hash, Base64, URL encode/decode, JSON format/query, timestamps, text stats, regex extract, CSV/JSON query, RSS/sitemap parse, and more. Pricing: $0.02 USDC on Base via x402 (EIP-3009), or $0.01 prepaid (deposit ≥$0.50 USDC, off-chain debit, no per-call gas). First call free per wallet on most tools. Includes pay-proxy: local stdio bridge that auto-signs x402 challenges — any MCP client gains payment capability without code changes or API keys.
Firecrawl
URL-to-clean-Markdown scraping for LLM-ready web content extraction.
Web Scraper — Clean Markdown from Any URL
Web content extraction API for AI agents. Scrape any URL and get clean, structured Markdown content with navigation, ads, and scripts stripped. Full JavaScript rendering via headless Chromium. Single and batch (10 URLs) modes. Built for RAG pipelines and AI research. Tools: web_scrape_to_markdown (single), web_scrape_batch (up to 10 URLs). Use this for RAG ingestion, research, content analysis, data extraction, or competitive intelligence. IMPORTANT: For screenshots/PDFs of pages, use capture_screenshot instead. For SEO analysis, use seo_audit_page. Returns: {markdown, title, wordCount, links[]}. No API key required — x402 micropayment $0.005/call on Base L2.
NEXUS Intelligence API
58-endpoint utility hub - DNS lookup, web scraping, CVE scanning, AI translation, PII detection, domain health audit, company intelligence. Pay per call via x402 on Base.
pipeworx
Live data for AI — one connector to SEC filings, economics, FDA, patents, weather, prediction markets, and more across 750+ official sources, with citations. The hero tool ask_pipeworx routes any question to the right source and returns structured, cited data.
Korean Law Search
Search and retrieve Korean statutes and administrative rules with precise filters. Access English translations and drill down to articles, paragraphs, and sub-items. Explore linkages with local ordinances and delegated authority to speed up legal research.
SteadyFetch
Reliable web fetching MCP server with built-in retry logic, circuit breaker patterns, caching, and anti-bot bypass. Fetches URLs as raw HTML or clean markdown optimized for LLM consumption. Includes domain health checks and cache management tools.
emblem-mcp
# EmblemAI MCP EmblemAI is a hosted Model Context Protocol server that gives AI agents a full-featured crypto wallet plus **200+ tools** for trading, DeFi, NFTs, and on-chain analytics across **7 blockchains**: - **Bitcoin** — native BTC, Ordinals inscriptions, Runes, BRC-20, Stamps / SRC-20, Alkanes, rare sats - **Solana** — SOL + SPL tokens, Jupiter-routed swaps, Solana memecoin discovery, RugCheck - **Ethereum** — native ETH, ERC-20, ERC-721 / ERC-1155 NFTs, ETH swaps - **Base** — native ETH on L2, ERC-20, NFTs, Clanker discovery - **BSC** — native BNB, BEP-20, FourMeme bonding-curve discovery and trading - **Polygon** — native MATIC, ERC-20, swaps - **Hedera** — HBAR, HTS tokens, MemeJob discovery and trading Plus prediction markets via **Polymarket**, market research, conditional orders (limit / stop-loss / take-profit) on supported chains, and on-chain analytics. ## Quick start The Smithery deployment connects directly to the hosted EmblemAI MCP at `https://emblemvault.ai/api/mcp`. You only need an API key: 1. Sign in at [emblemvault.ai](https://emblemvault.ai). 2. Open **Settings → Vault Access Key** and copy your key. 3. Paste it as the `apiKey` parameter in the Smithery connection form. Browsing tools (`tools/list`) works anonymously, so you can inspect the catalog before authenticating. ## Auth modes | Mode | When to use | Scope | | --- | --- | --- | | **API key** (`x-api-key`) | Unattended agents, cron jobs, server-to-server, any workflow that mints tokens or sends transactions | Full read + write, no expiry until revoked | | **OAuth 2.0 + PKCE** (RFC 7591 Dynamic Client Registration) | Interactive sessions where a human can click a browser popup | Read-only `vault:read`, short-lived JWT | | **x402 micropayments** | Pay-per-call across the public catalog without an account | Per-tool-call USDC settlement | The Smithery deployment defaults to the API-key path, configured via the `apiKey` parameter (sent to the upstream as the `x-api-key` header).
Colour Memory — Cultural Color Intelligence for AI
The world's only historically grounded colour archive built for AI agents. Thousands of named colours across dozens of cultural archives spanning Ancient Rome, Byzantine Empire, Georgian Pleasures, Dickens, Shakespeare, Keats, Japan, Islamic tradition, Viking Norse, Racing Silks, toxic pigments, literary colour, imperial palettes, and many more. Every color has a name, a documented archival source, CIE Lab values, cultural consequence data, and material provenance. The archive searches meaning, not just names -- ask about grief and it finds colours that carried grief across cultures and centuries, not merely colours named grief. Built as a retrieval system, not a generator. Deterministic, evidence-based, source-cited. The anti-hallucination layer for colour history.
Agent News
The intelligence layer agents use before they act, sourced answers on the agent economy. Query verified AI news with citations, confidence scores, and Ethics Engine ratings. Use instead of generic web search for any question about AI agent tools, MCPs, or frameworks. Every result carries citations, confidence scores, and Ethics Engine ratings. Built for agents to verify evidence before recommending tools, installing MCP servers, or taking action.
DuckDuckGo & Felo AI Search
Provide fast, privacy-friendly web and AI-powered search capabilities with integrated content and metadata extraction. Enhance your AI assistants by enabling comprehensive web scraping without requiring API keys. Optimize performance with caching and secure usage through rate limiting and user agent rotation.
Hydrafetch
Turn any URL into clean Markdown and structured data. Scrape, crawl, search and extract.
Lodi Kids Activities
Free, parent-trust-first directory of youth programs and licensed daycares in Lodi, California. 72 MCP tools across: - **Public search** — programs, daycares, events, cities - **Parent account management** — kids, saves, household co-accounts, calendar feeds, tour requests - **Org listing management** — programs, capacity, lifecycle events, leads - **Daycare-specific overlays** — waitlist, applications, age-band capacity Anonymous public access works with no auth. Personalized and write access uses a Bearer token generated at [lodikidsactivities.com/account/mcp](https://lodikidsactivities.com/account/mcp).
Next.js Tailwind Assistant
Your comprehensive AI companion for building modern Next.js applications with React and Tailwind CSS. This MCP server provides instant access to complete documentation, production-ready components, and battle-tested design patterns abstracted from professional templates.
menjometre
Menjometre is an independent observatory of Catalan public spending. This MCP server gives any LLM client read-only access to the full dataset and the statistical scoring layer built on top of it. ## Coverage - **~19.5 M** grant records from the Registre d'Ajuts i Subvencions de Catalunya (RAISC), 2016–present - **~1.7 M** public procurement contracts from Dades Obertes de Catalunya - **~496 K** beneficiary entities with normalised identities - A **public-figure graph** linking Parlament deputies, Generalitat officials, and the boards of publicly-funded entities - The **Menjometre score** — a statistical indicator flagging concentration and recurrence patterns in public spending ## What the 48 tools cover | Domain | What you can ask | |---|---| | `entities` | Search, profile, network, activity timeline for any funded organisation | | `grants` | RAISC lookups, concentration analysis, purpose trees, year-over-year trends | | `contractes` | Procurement lookups, sole-source detection, fragmentation patterns | | `organs` | Granting-body rankings and profiles | | `xarxa` | Publ
wikipedia-mcp-server
Search Wikipedia, read summaries and full text, target sections, find nearby pages, list languages.
Jina Reader
Web content extraction optimized for readability and RAG pipelines.
JieBang Tools
20 free developer tools - JSON/YAML, SQL, XML, Cron, QR code, SEO check, URL shortener, image converter, timezone, and more. No API key required.
Paper Search (arXiv + Semantic Scholar + OpenAlex)
Unified academic paper search for AI agents. search_all queries arXiv, Semantic Scholar and OpenAlex at once, de-duplicates the same work across corpora (by DOI/title) and re-ranks with Reciprocal Rank Fusion so papers found by several sources rank highest. read_paper returns full arXiv text with formulas as LaTeX. Plus citation graphs, author metrics (h-index), recommendations and full-text snippet search across 220M+ papers. Remote streamable-HTTP, no auth required.
AI Research Assistant
The server provides immediate access to millions of academic papers through Semantic Scholar and arXiv, enabling AI-powered research with comprehensive search, citation analysis, and full-text PDF extraction from multiple sources (arXiv and Wiley open-access). - No API key is required.
WebLens
Scrape, crawl, map and extract the web. Pay per call in USDC, no account or API key.
ArXiv Scout
Search and retrieve academic papers directly from arXiv with advanced query capabilities. Extract full text from PDFs to generate summaries, literature reviews, and side-by-side comparisons. Track citations and references to discover related research and map out academic trends.
Screenshot & PDF Capture — Full-Page Chromium Rendering
Web capture API for AI agents. Take full-page screenshots (PNG/JPEG/WebP) and generate PDFs from any URL using headless Chromium. Custom viewport size, device emulation, and scroll capture. Tools: capture_screenshot (image), webpage_to_pdf (PDF). Use this for visual regression testing, documentation, archiving web pages, or generating reports. IMPORTANT: For text extraction from pages, use web_scrape_to_markdown. For custom PDF generation from data, use document_generate_pdf. Returns: base64-encoded image or PDF. No API key required — x402 micropayment $0.008/call on Base L2.
uk-due-diligence
Tools across five UK public registers. All official APIs. Give an agent a company name and it pulls corporate status, filing compliance, director networks, beneficial ownership chains, disqualification checks, insolvency notices, VAT validation, and property transactions.
Google Drive
Search Google Drive, upload and download files, create folders, move and rename items, and manage sharing and file permissions.
ai.smithery/arjunkmrm-scrapermcp_el
Extract and parse web pages into clean HTML, links, or Markdown. Handle dynamic, complex, or block…
stagenth · 网页数据
Web scraping to clean Markdown with JS rendering, multi-page crawl, structured extract, sitemaps.
Google Maps Lead Gen MCP — Local Business Enrichment
B2B lead generation tool: search Google Maps by 'plumbers in Austin', get back business profiles with emails, phone numbers, ratings, websites. The differentiator over plain Google Maps API is the contact-enrichment layer (scraped from each business's site). Built for agency prospecting workflows.
Web Auditor
Analyze web pages and entire sites for SEO, accessibility, performance, and security compliance. Generate detailed reports and actionable issue lists to improve site quality and user experience. Track historical audit data to monitor improvements across multiple crawl sessions.
ko-financial-data
Real SEC, 13F, insider, congress & macro data your AI agent can cite. 24 tools — institutional 13F holdings (85M+ rows), insider & congress trades, spot-Bitcoin-ETF ownership, company financials, and Treasury/Fed/BLS macro series. Every answer traces back to the original SEC filing. Free tier: 200 calls/day.
Web Scout
Search the web and extract clean, readable text from webpages. Process multiple URLs at once to speed up research with reliable throttling and error handling. Quickly compile sources and summaries for briefs, reports, or competitive analysis.
FavCRM — Agentic CRM
**Agentic CRM for service businesses** — 136 typed tools across customers, bookings, loyalty, invoices, and WhatsApp/SMS. Hosted at `api.favcrm.io/mcp`. ### What it does - **Customers & loyalty** — segments, memberships, rewards, points - **Bookings** — create, confirm, cancel, no-show; capacity-aware schedules - **Commerce** — invoices, payments, products, orders, inventory - **Communications** — WhatsApp, SMS, email with templates - **CRM core** — contacts, tasks, tickets, deals, campaigns ### Auth API key (`fav_mcp_...`). Sign up free at [favcrm.io](https://favcrm.io) → Settings → API Keys. **Free tier:** 100 customers + 200 bookings/month. No credit card. ### Quickstart Install snippets, examples, and smoke tests: [github.com/favcrm/mcp](https://github.com/favcrm/mcp)
primitive.dev
Email infrastructure for AI agents — send, receive, search, and manage email over a clean HTTP API. Connect verified domains, route inbound mail, set up webhooks, and automate email workflows. Authenticate with a Primitive API key or via OAuth.
MTG MCP Server
Magic: The Gathering card search, combo lookup, draft analytics, and Commander tools.
com.apify/apify-mcp-server
Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store ⚡
What Do They Know?
Search UK FOI requests, public authorities, and responses. Draft and submit Freedom of Information requests.
Google Sheets
Search and inspect Google Sheets, create and edit spreadsheets, scan for data issues, review edit history, and manage comments.
Research Report Generator — Multi-Source Analysis
AI-powered research report generator API for AI agents. Generate structured research reports on any topic: multi-source web research, key findings with citations, analysis sections, and recommendations in clean Markdown. Tools: research_generate_report. Use this for market research, competitive analysis, due diligence, or preparing briefing documents. Returns publication-ready Markdown. IMPORTANT: For quick fact-checking, use research_check_fact instead. Returns: {report (markdown), sources[], wordCount}. No API key required — x402 micropayment $0.02/call on Base L2.
pubmed-mcp-server
Search PubMed/Europe PMC, fetch articles and full text (PMC/EPMC/Unpaywall), citations, MeSH terms.
Brainiall Web
Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
com.thenextgennexus/web-scraping-mcp-server
Generic URL crawl + HTML extraction — fallback for sites without dedicated MCPs.
io.github.Cal-Dev-Tech/perpage-mcp-server
Give AI agents clean, LLM-ready web data — scrape any URL to markdown or extract structured JSON.
Spider Cloud
Crawl, scrape, search the web, and automate browsers at scale with anti-bot bypass.
agent-registry
The Registry for the Agent Economy Discover, verify, and connect with AI agents. The first protocol-aware directory built on the A2A standard.
Mirabello Consultancy — Investment Migration & Wealth Protection
Authoritative, source-cited data across five layers: citizenship- & residency-by-investment (CBI/RBI/golden visa) programmes with the Mirabello Investment Migration Index; country immigration pathways via the origin-aware Freedom Compass planner; HNWI tax & wealth-protection (income/CGT/inheritance/wealth tax, trusts, succession, treaties); qualifying investment real estate; and consultation hand-off. 44 tools. By Mirabello Consultancy (Zurich + Dubai).
DialogBrain
DialogBrain gives AI agents access to all your messaging channels — Telegram, WhatsApp, Instagram, Email, and more — through a single MCP interface. Read conversations, send messages, search contacts, and manage threads across 10+ channels. Perfect for LangGraph, Claude, and any MCP-compatible agent framework.
hostdefi-x402
Token-safety intelligence and non-custodial EVM swap quotes for agents. scan_token is FREE and keyless (fair use 100/day/IP): pass any Solana mint or EVM contract address and get an A+ to F safety grade computed from on-chain checks — mint and freeze authority, liquidity depth, holder concentration and contract flags. The deeper tools are machine-payable per call over the x402 protocol in USDC on Solana, Base, Polygon, Arbitrum or Avalanche, charge-on-success only, with no account or API key needed.
UnClick
AI agent tool marketplace. 60+ tools for social, e-commerce, accounting, messaging, media, and more. Search and call tools for Slack, Reddit, Discord, Shopify, Xero, Bluesky, TMDB, Steam, and 50+ more services.
Setup checklist
- Decide whether you need raw HTML, clean markdown, or structured JSON from pages.
- Check rate limits, auth requirements, and whether the server uses headless browsers.
- Add the MCP server to your client and test on a single URL before batch jobs.
- Treat scraped content as untrusted input — never pipe it straight into prod config.
How to choose
- Firecrawl-style servers excel at URL → markdown for LLM consumption.
- Search + scrape combos (Exa, Jina) work well for research agents.
- Browser-based scrapers handle JavaScript-heavy sites; HTTP-only servers are faster for static pages.
Related topics
Where Web Scraping MCP fits
Turn URLs into structured content for agents. Compare scraping, crawling, and search servers with indexed tool lists and setup steps.