Crawl4AI RAG
by coleam00 · Databases · mcp-server, database, ai
Integrates web crawling with RAG functionality to enable website content retrieval, storage in vector databases, and semantic searching over crawled data for enhanced knowledge access
This MCP server integrates web crawling capabilities from Crawl4AI with RAG (Retrieval Augmented Generation) functionality, enabling AI agents and coding assistants to crawl websites, store content in Supabase vector databases, and perform semantic searches over the crawled data. Developed by Cole Medin, it provides four essential tools: single page crawling, intelligent recursive website crawling with automatic URL type detection, source retrieval, and vector-based RAG queries with optional source filtering. The implementation is designed to be incorporated into Archon as a knowledge engine for AI coding assistants, with future plans to support different embedding models and local deployment options using Ollama.
Source: https://github.com/coleam00/mcp-crawl4ai-rag
Install
git clone https://github.com/coleam00/mcp-crawl4ai-ragTags: mcp-server, database, ai
⭐ 2,040 GitHub stars · Source: pulsemcp
About Databases MCP servers and Claude skills
Databases MCP servers extend what AI agents can do inside Claude Code, Cursor, Copilot, Codex, and Windsurf. The Skiln directory indexes 16,000+ such integrations across 22 categories.
Crawl4AI RAG is one of hundreds of Databases entries indexed on Skiln. Browse the full Databases category or the complete directory of Claude skills, MCP servers, agents, commands, and hooks.
Related Databases MCPs and skills
- 🐳 Crawl4AI+SearXNG MCP Server by coleam00
A Docker-based MCP server combining Crawl4AI, SearXNG, and Supabase to provide AI agents with web search, crawling, and Retrieval Augmented Generation (RAG) capabilities. Requires external data files such as .env for API keys and configuration. Database setup requires running the SQL script 'crawled_pages.sql' in Supabase.
- PostgreSQL with GitHub OAuth by coleam00
Provides secure PostgreSQL database access with GitHub OAuth authentication, enabling read-only operations for all authenticated users while restricting write operations to allowlisted GitHub usernames through role-based access control.
- Mem0 (Long-Term Memory) by coleam00
Provides persistent long-term memory capabilities through semantic indexing, retrieval, and search functions with support for multiple LLM providers and PostgreSQL vector storage.
- Snowflake by snowflake-labs
Bridges AI applications with Snowflake's data platform for database interaction
- Ultra (Multi-AI Provider) by realmikechong
Unified server providing access to OpenAI O3, Google Gemini 2.5 Pro, and Azure OpenAI models with automatic usage tracking, cost estimation, and nine specialized development tools for code analysis, debugging, and documentation generation.
- Apache Doris by apache
Enables direct SQL query execution and metadata retrieval from Apache Doris databases without switching contexts.
- SQL Server Performance Monitor by erikdarlingdata
SQL Server performance monitoring with DuckDB storage and natural language queries for CPU, wait stats, blocking, query performance, memory, and I/O.
- MySQL Database Manager by wenb1n-dev
Provides direct access to MySQL databases with advanced features like multiple SQL execution, table metadata querying, execution plan analysis, and Chinese field to pinyin conversion through a configurable Python-based server.
Frequently asked questions
How do I install Crawl4AI RAG?
Add the install command above to your Claude Code, Cursor, or Windsurf MCP configuration. Most servers register via npx, a local command, or a Docker image. Refer to the source repository for environment variables and credential requirements.
Which clients support Crawl4AI RAG?
Any MCP-compatible client works: Claude Desktop, Claude Code CLI, Cursor, Windsurf, Zed, and VS Code with the official MCP extension. OpenAI Codex and GitHub Copilot increasingly support MCP via adapter bridges.
Is Crawl4AI RAG free?
The server itself is typically open source. Any upstream service (API keys, paid tiers, hosted infrastructure) may have its own pricing. Check the source repository for details.