Data & search harnesses for AI agents
89 open-source Data & search harnesses an AI agent can use — MCP servers, SDKs, and adapters. Browse them on Loadbay. An agent can search these over Loadbay's MCP:
claude mcp add --transport http loadbay https://loadbay.xyz/api/mcp
→ Best Data & search harnesses (top picks, ranked)
- markitdown — Python tool that converts Office documents, PDFs, and other files to Markdown for LLM ingestion.
- Firecrawl — Search, scrape, and crawl the web at scale and get clean, structured content. The data layer behind a lot of agents.
- worldmonitor — Real-time global intelligence platform exposing AI-curated news, geopolitical monitoring, and infrastructure tracking via an MCP server, REST API, and SDKs. Agents can query the Country Instability Index, cross-stream correlations, and finance radar programmatically.
- RAGFlow — Open-source RAG engine built on deep document understanding, with grounded citations and an end-to-end retrieval pipeline.
- Scrapling — An adaptive web scraping MCP server that lets AI agents extract targeted content from dynamic websites with stealth browser capabilities. Exposes HTML parsing, CSS/XPath selector extraction, browser session management, and screenshot capture as MCP tools.
- crawl4ai — Open-source LLM-friendly web crawler and scraper that outputs clean markdown and structured data for agents.
- MinerU — Converts PDFs and documents into machine-readable Markdown and JSON, extracting formulas, tables, and reading order for RAG.
- TrendRadar — AI-driven news and trend aggregator with MCP integration — agents query aggregated multi-platform hot topics, RSS feeds, and LLM-analyzed news briefings via natural language.
- Scrapy — Fast asynchronous Python framework for large-scale web crawling and structured data extraction.
- docling — Document parsing and ingestion toolkit that converts PDFs, Office files, and images into structured data for gen AI.
- AnythingLLM — All-in-one app turning documents into a private RAG chatbot, bundling ingestion, vector storage, and retrieval; MCP-compatible.
- mem0 — Universal memory layer that gives AI agents persistent long-term memory across sessions via SDK and MCP.
- mempalace — MemPalace is a local-first AI memory MCP server with 44 tools, storing conversation history as verbatim text with semantic search via ChromaDB. Agents use it as a persistent, hierarchically organized knowledge store (wings, rooms, drawers) across sessions.
- Meilisearch — Fast, typo-tolerant open-source search engine with built-in hybrid keyword and vector/semantic search.
- llama_index — Data framework for building document agents and RAG pipelines that connect LLMs to private and external data sources.
- milvus — High-performance, cloud-native vector database built for scalable vector ANN search over billions of vectors.
- Faiss — Meta library for efficient similarity search and clustering of dense vectors; the de-facto ANN index under many vector DBs.
- Open Notebook — Local-first research platform with a built-in MCP server that agents connect to via Claude Desktop, VS Code, or other MCP clients to search and query notebooks containing PDFs, videos, audio, and web pages. Supports 18+ AI providers including Anthropic, OpenAI, and Ollama for in-notebook AI queries.
- Marker — Fast, high-accuracy converter of PDF, EPUB, and docs to Markdown and JSON with table, equation, and layout handling.
- PageIndex — Vectorless RAG system that builds hierarchical tree indexes from documents and uses LLM reasoning for context-aware retrieval, eliminating vector databases and chunking. Exposes an MCP server for agent integration and supports multi-LLM backends via LiteLLM.
- graphrag — Modular graph-based retrieval-augmented generation system that builds knowledge graphs from documents for agent queries.
- qdrant — High-performance, massive-scale vector database and vector search engine with REST and gRPC APIs.
- searxng — Free, self-hostable metasearch engine that aggregates results from many services without tracking users.
- Supermemory — Memory and context engine for AI agents with a fast local-capable Memory API, SDKs, and tools for long-term recall across conversations.
- chroma — Open-source embedding database and search infrastructure for building AI apps with retrieval and memory.
- gbrain — Production-grade knowledge synthesis layer that powers autonomous AI agents by indexing pages into a graph, synthesizing answers across 43 curated skills, and exposing results as an MCP layer for Claude Code, Codex, and Cursor.
- graphiti — Framework for building real-time temporal knowledge graphs as memory for AI agents, with an MCP server.
- Scrapegraph-ai — AI-powered Python scraper that uses LLMs and graph pipelines to extract data from websites and documents.
- OpenViking — A unified context database for AI agents that manages memory, knowledge retrieval, and skills through a filesystem-like hierarchy with tiered loading (abstract summaries, overviews, full details). Reduces token consumption while improving retrieval accuracy via recursive directory-based search strategies.
- Typesense — Open-source typo-tolerant search engine with native vector and hybrid semantic search, a lightweight retrieval backend.
- haystack — Orchestration framework for production LLM applications with modular pipelines for retrieval, RAG, and agent workflows.
- Crawlee — Web scraping and browser-automation library (Node and Python) built to extract data for AI, LLMs, and RAG.
- letta — Platform for stateful agents with advanced memory (formerly MemGPT) that learn and self-improve over time.
- MaxKB — Open-source enterprise agent platform with RAG knowledge base and MCP server integration — agents query a vector database of uploaded documents and data, with multi-LLM support.
- pgvector — Postgres extension adding vector types and HNSW/IVFFlat similarity search, for vector retrieval inside a relational DB.
- Turso — A modern SQLite-compatible in-process database written in Rust with a built-in MCP server that exposes nine tools (query, insert, update, schema management) to any AI agent. Run `tursodb your_database.db --mcp` to give Claude Code, Claude Desktop, or Cursor full read-write access to structured data.
- DBX — Lightweight multi-database client with a built-in AI assistant and MCP server so agents can query and manage 90+ databases from one tool.
- cognee — Open-source AI memory platform giving agents persistent long-term memory via a self-hosted knowledge-graph engine.
- Hindsight — A biomimetic agent memory system that organizes long-term knowledge into worlds (facts), experiences, and mental models so agents can retain, recall, and reflect on information across sessions. Achieves state-of-the-art on LongMemEval and deploys via Docker, bare-metal, or embedded Python mode.
- weaviate — Open-source vector database storing objects and vectors, combining vector search with structured filtering.
- mcp-toolbox — Google's open-source MCP server that connects AI agents directly to enterprise databases including BigQuery, PostgreSQL, MySQL, Oracle, and Spanner. Exposes prebuilt SQL exploration tools and a framework for custom parameterized database tools with built-in auth, connection pooling, and OpenTelemetry observability.
- memvid — A serverless agent memory layer that packages data, embeddings, and search structures into portable .mv2 files for persistent long-term memory without databases. Supports vector similarity search, full-text search, and temporal queries via SDKs in Rust, Python, Node.js, and CLI.
- unstructured — Open-source ETL library that transforms complex documents into clean structured data for language models.
- OpenMetadata — Open Context Layer for Data and AI: an open platform for metadata management, data discovery, and semantic context that exposes an MCP server so AI agents can query and navigate enterprise data assets.
- PentestGPT — Automated penetration-testing agentic framework powered by LLMs that guides and runs offensive-security workflows.
- txtai — All-in-one embeddings framework for semantic search, RAG, and language-model workflows over your own data.
- Open Deep Research — Open-source deep research agent built on LangGraph that autonomously searches, synthesizes, and generates comprehensive reports; supports MCP servers and works across many LLM providers and search APIs.
- lancedb — Developer-friendly embedded retrieval library and vector database for multimodal AI search.
- cocoindex — Incremental data-indexing engine for long-horizon AI agents — keeps RAG pipelines and agent context always fresh by only reprocessing changed data from codebases, Slack, docs, and databases.
- hexstrike-ai — MCP server that lets AI agents autonomously run 150+ cybersecurity tools for pentesting, vuln discovery, and bug-bounty automation.
- unstract — No-code LLM platform that turns unstructured documents (PDFs, images, scans) into structured JSON, exposing an MCP server agents can connect to for document extraction. Supports 9+ LLM providers, 5+ vector databases, and ETL connectors to S3, Snowflake, BigQuery, and more.
- MindSearch — An open multi-agent web-search framework, in the spirit of Perplexity Pro, that plans queries and synthesizes answers.
- firecrawl-mcp-server — Official Firecrawl MCP Server that adds powerful web scraping, crawling, and search capabilities to Claude Desktop, Cursor, and other MCP clients via the Firecrawl API. Agents can crawl entire sites or search the web and receive clean Markdown output.
- gemini-notebook-mcp-cli — An MCP server and CLI that gives AI agents programmatic access to Google Gemini Notebook for creating, querying, and managing notebooks. Agents use MCP tool calls to read and write notebook content, enabling research workflows driven by an LLM.
- whodb — Database management platform with an integrated MCP server that gives AI agents text-to-SQL access to PostgreSQL, MySQL, SQLite, Oracle, SQL Server, MongoDB, and Redis. Includes schema exploration, ER diagram generation, and a visual query builder alongside the MCP tools.
- Integuru — AI agent that reverse-engineers a platform's internal APIs from browser traffic to build permissionless integrations.
- exa-mcp-server — MCP server letting agents perform web search and crawling through the Exa neural search API.
- RuVector — High-performance Rust-based vector database and agent memory substrate that combines semantic embeddings, graph relationships, and self-learning; ships an mcp-brain MCP interface for shared agent memory. Agents connect via the MCP server to persist working, episodic, semantic, and procedural memory across sessions without needing an external API key.
- AI-Infra-Guard — Full-stack AI red-teaming platform for agent scan, MCP scan, AI infra scan, and LLM jailbreak evaluation.
- semantica — A graph-native AI infrastructure platform that exposes a full MCP server (via semantica-mcp) with tools for entity extraction, decision recording, graph queries, rule-based reasoning (Rete, Datalog, SPARQL), and W3C PROV-O audit trails. Designed for multi-agent systems in regulated industries requiring provenance-tracked, explainable AI decisions.
- tabularis — Desktop SQL workspace for 14+ databases—PostgreSQL, MySQL, SQLite, DuckDB, ClickHouse—with a built-in MCP server so Claude, Cursor, and Devin can query and manage databases directly.
- flint-chart — A declarative chart specification language and MCP server by Microsoft that lets AI agents produce expressive, human-editable charts without hallucinating chart syntax or configuration.
- ai-memory — Rust service that gives AI coding CLIs a long-term shared memory so you can quit Claude Code mid-task and continue in Codex without re-explaining the architecture. Exposed to agents through MCP config and lifecycle hooks across Claude Code, Codex, Cursor, Gemini CLI, OpenCode, Devin, Kimi Code, Oh My Pi, Pi, and Command Code.
- hister — Personal search engine with MCP server that indexes browser history and local files, letting AI agents query full-text indexed web pages and documents using field filters, phrases, wildcards, negation, and optional semantic search via a configurable embeddings endpoint.
- DBHub — A zero-dependency database MCP server for Postgres, MySQL, SQL Server, and more — query your data in natural language.
- spiceai — Portable Rust runtime that operates as an MCP server sidecar providing SQL query federation, hybrid vector/BM25 search, and LLM inference across 30+ data connectors, with text-to-SQL tools and agent skills for Claude Code and Cursor. Enables agents to query federated data sources at millisecond latency.
- open-seo — An open-source SEO platform with an MCP server and pre-built agent skills for keyword research, rank tracking, and domain insights — agents connect via MCP or install skills to run full SEO workflows against DataForSEO and Google Search Console.
- sie — A self-hosted inference server that consolidates embedding, retrieval, document-to-markdown, structured extraction, and content safety into a single OpenAI-compatible cluster, with an MCP edge component that offloads document work from Claude and other MCP clients. Integrates natively with LangChain, LlamaIndex, CrewAI, DSPy, Haystack, and vector stores like Chroma and Qdrant.
- google-analytics-mcp — An official Google Analytics MCP server that enables language models to query Analytics account data and run reports through the Google Analytics Admin and Data APIs. Exposes tools for account summaries, standard reports, and real-time analytics.
- agent-scan — Security scanner for AI agents, MCP servers, and agent skills that detects vulnerabilities and misconfigurations.
- modelcontextprotocol — Official Perplexity MCP server that gives AI assistants web-wide search and answers through the Perplexity API.
- korean-law-mcp — MCP server wrapping 42 Korean Ministry of Government Legislation APIs into 10 tools for statutes, precedents, ordinances, and treaties. Includes citation hallucination verification, article impact graphs, and temporal law comparison.
- tavily-mcp — Production MCP server giving agents real-time web search, extract, map, and crawl via the Tavily API.
- webclaw — Fast, local-first web content extraction Rust binary that scrapes, crawls, and converts web pages to structured Markdown; ships a CLI, REST API, and MCP server for LLM and agent pipelines.
- apify-mcp-server — Official Apify MCP server that gives AI agents access to thousands of ready-made web scrapers and automation Actors — extract data from social media, search engines, maps, and e-commerce sites via OAuth or API token.
- mcp-server-qdrant — Official Qdrant MCP server exposing vector storage and semantic search as a memory layer for agents.
- memanto — Pluggable long-term memory layer for AI agents with semantic and episodic recall, designed to integrate directly with CrewAI, LangChain, and other agent frameworks. Agents read and write memories through a simple API that supports RAG-based retrieval and stateful multi-session continuity.
- jupyter-mcp-server — MCP server giving AI agents full control over Jupyter notebooks — create cells, execute code, read outputs, and manage kernel state. Enables coding agents to run and inspect computational notebooks interactively.
- brave-search-mcp-server — Official Brave Search MCP server providing web, image, video, news, and local search.
- OpenOSINT — AI-powered OSINT agent with an integrated MCP server exposing 16 intelligence-gathering tools including Sherlock, Holehe, and Maigret for authorized security research. Works with Claude, GPT-4, or local models via an interactive REPL or CLI.
- harvey-labs — Open-source evaluation benchmark and execution harness for scoring LLM agents on realistic legal work tasks across 24+ practice areas, enabling researchers to measure and advance agent capabilities in legal support workflows.
- retentioneering-tools — Python toolkit, MCP server, and agent skills for reproducible clickstream and event-log analytics. AI agents connect via MCP to build customer journey maps, run behavioral segmentation, A/B tests, and Markov chain simulations over product event data.
- Ryze SEO MCP — Open-source MCP server that gives Claude and other agents direct access to SEO and GEO data from Google Search Console, GA4, and ads — exposing keyword research, rank tracking, backlinks, and site audits as callable tools.
- skills-for-fabric — Reusable agent skills and MCP server configurations that let AI coding tools (Claude Code, GitHub Copilot, Cursor) query and operate Microsoft Fabric workloads including SQL, Spark, Power BI, and KQL endpoints via Azure credentials.
- mcp-filesystem-server — A filesystem MCP server — give an agent scoped read/write access to files and directories.
- seo — A local MCP server exposing 70+ SEO audit tools to any AI agent, using the agent's own crawl data, Google Search Console, and GA4 analytics for comprehensive technical SEO analysis and optimization.
- kindly-web-search — Web search MCP server for AI coding tools and agents, supporting Serper, Tavily, and SearXNG backends. Agents call it to search the web and retrieve full page content in a structured format, compatible with Claude Code, Codex, Cursor, and 40+ others.
- mcp-google-map — MCP server for Google Maps including geocoding, place search, directions, and distance calculations.
- livetennisapi-mcp — An MCP server for the Live Tennis API that gives agents access to real-time tennis scores, match odds, and model-computed win probabilities. Works with Claude, Cursor, and any MCP-compatible agent.