Browser & computer harnesses for AI agents
52 open-source Browser & computer harnesses an AI agent can use — MCP servers, SDKs, and adapters. Browse them on Loadbay. An agent can search these over Loadbay's MCP:
claude mcp add --transport http loadbay https://loadbay.xyz/api/mcp
→ Best Browser & computer harnesses (top picks, ranked)
- browser-use — Make any website usable by an agent. It drives a real browser to click, type, and finish tasks online — the most-starred browser harness by a wide margin.
- puppeteer — Node.js library providing a high-level API to control Chrome and Firefox headlessly, enabling AI agents to automate web browsing, scrape pages, take screenshots, fill forms, and execute JavaScript in a real browser environment.
- open-interpreter — Lets LLMs run code and control the local computer through a natural-language interface for OS-level automation.
- Chrome DevTools MCP — Official Chrome DevTools MCP server letting an agent control and inspect a live Chrome browser for automation and debugging.
- UI-TARS Desktop — ByteDance multimodal agent stack (Agent TARS + UI-TARS Desktop) that controls computer and browser operators via natural language.
- agent-browser — Native Rust CLI that exposes browser automation to AI agents via Chrome DevTools Protocol, providing accessibility tree snapshots, session management, screenshot capture, and a built-in MCP server. Integrates with cloud browser providers like Browserbase and Browserless for headless operation.
- Lightpanda — Headless browser written in Zig from scratch for AI agents and automation, with CDP compatibility and a native MCP server. Far lighter and faster than headless Chrome for agent browsing workloads.
- playwright-mcp — Microsoft's Playwright MCP server — drive a real browser (navigate, click, fill, assert) from any agent.
- AIHawk — Open-source AI browser agent for web automation — drives a real browser to complete tasks on any website using plain English instructions, supporting computer-use and web browsing workflows.
- OpenCLI — Converts websites and browser sessions into deterministic CLI interfaces for AI agents. Install the opencli-browser skill in any coding agent to navigate, fill forms, click, extract and inspect pages through the user logged-in Chrome via CDP.
- OmniParser — Screen-parsing tool that converts UI screenshots into structured elements to enable pure vision-based GUI agents.
- stagehand — SDK for browser agents that adds act, extract, and observe primitives on top of Playwright for AI-driven web automation.
- skyvern — Automates browser-based workflows using LLMs and computer vision to operate websites without site-specific scripts.
- page-agent — JavaScript library that embeds a GUI agent directly inside any webpage, letting AI control web interfaces via natural-language text-DOM commands without screenshots or headless browsers. Also ships an MCP server for agent-driven multi-page automation across browser tabs.
- cua — Infrastructure for computer-use agents: sandboxes, SDKs, and benchmarks so an agent can drive a whole desktop without escaping it.
- Anthropic computer-use demo — Anthropic official computer-use reference: a containerized Linux desktop where Claude controls the GUI via screenshots and tool calls.
- web-ui — Browser-based UI for running web-automation agents with support for custom models and persistent browser sessions.
- browser-harness — A thin, self-editing CDP harness that connects an LLM directly to a real Chrome browser via WebSocket. The agent writes missing helpers at runtime, improving the harness each run — no middleware between the model and the page.
- maxun — No-code platform that turns websites into structured APIs through browser-based scraping and AI data extraction.
- midscene — Vision-driven UI automation that drives web and mobile interfaces from natural language for AI agents.
- nanobrowser — An open-source Chrome extension that runs multi-agent web-automation workflows right in your browser.
- mcp-chrome — Chrome MCP Server is a Chrome extension that exposes browser tabs, history, bookmarks, and live page content as MCP tools, letting an agent read and control the user's Chrome instance via the Model Context Protocol.
- Agent-S — An open framework that lets an agent use a computer the way a person does: read the screen, move the mouse, type.
- BrowserOS — Open-source agentic web browser (Chromium fork) that runs AI agents natively inside the browser.
- Bytebot — A self-hosted AI desktop agent that automates computer tasks inside its own containerized desktop.
- self-operating-computer — A framework that lets a multimodal model operate a computer by looking at the screen and moving the mouse.
- camofox-browser — Stealth headless Firefox browser server for AI agents, built on Camoufox fingerprint spoofing. Drop-in Puppeteer/Playwright replacement that helps agents browse sites that block ordinary headless Chrome.
- UFO — UI-focused agent that operates Windows applications via natural language using GUI and API actions.
- ego-lite — Browser automation runtime for AI agents that runs tasks in isolated parallel workspaces while inheriting the user's Chrome session; exposes a JavaScript tool interface for Claude Code, Codex, and Hermes agents to browse and automate web tasks concurrently.
- browser-tools-mcp — An MCP server that streams browser console logs, network requests, and screenshots directly into Cursor, Claude Code, and other MCP-compatible IDEs so agents can monitor and debug live web apps in real time.
- steel-browser — Open-source browser API and sandbox that lets AI agents automate the web without managing browser infrastructure.
- browsermcp — An MCP server that gives AI agents full browser control — navigating pages, clicking elements, filling forms, and extracting content via Playwright. Agents issue browser commands through standard MCP tool calls.
- Osaurus — Native macOS agent harness built in Swift that runs AI agents fully offline with persistent memory and MCP server support. Compatible with OpenAI, Anthropic, Ollama, and Apple Foundation Models, with agents living locally and accessible to any MCP-connected client.
- LaVague — Large Action Model framework that turns natural-language objectives into executable web automation for AI agents.
- Playwright MCP (ExecuteAutomation) — Popular community Playwright MCP server enabling agents to automate browsers and APIs, with screenshots and codegen.
- mobile-mcp — MCP server for mobile automation that lets agents drive iOS and Android devices via ADB and Appium, enabling automated app testing, UI interaction, and scraping of mobile applications through the Model Context Protocol.
- openagent — Single-binary personal AI assistant with computer-use, browser-use, and coding-agent loops; connects to 30+ LLM providers and ships a built-in tool-management UI and activity monitor.
- Peekaboo — A macOS CLI and optional MCP server that lets AI agents capture screenshots of any application or the entire desktop, with optional visual question-answering via local or remote AI models. Agents invoke it to observe screen state during computer-use tasks.
- agent-device — An MCP server, CLI, and typed Node.js API for AI coding agents to automate and verify mobile apps across iOS, Android, HarmonyOS, TV, and web. Agents can tap, swipe, type, capture screenshots, and run UI tests against connected simulators or real devices.
- Open-Interface — Controls any computer using LLMs by simulating keyboard and mouse to complete user-specified tasks across apps.
- notte — Framework to build web agents and deploy serverless web automation functions on managed browser infrastructure.
- OpenAdapt — Generative RPA that records desktop screen and input activity and replays it with multimodal models to automate GUI tasks.
- open-codex-computer-use — An open-source alternative to Codex Computer Use that gives agents cross-platform GUI automation — mouse clicks, keyboard input, and screen capture — via MCP on macOS and other platforms using accessibility APIs.
- stealth-browser-mcp — Python MCP server for browser automation using Chrome DevTools Protocol, enabling AI agents to navigate pages, intercept network traffic, and control Chromium in stealth mode. Exposes web-automation tools that bypass standard bot-detection systems.
- WebArena — Self-hostable realistic web environment (e-commerce, forums, GitLab, CMS) for building and benchmarking autonomous web agents.
- BrowserSkill — A CLI and browser extension that lets AI agents use your real, logged-in browser sessions without interrupting your work. Agents invoke it via the bsk CLI to automate sites where you are already authenticated, with human-in-the-loop support for CAPTCHAs and confirmations.
- Agent-E — Agent-driven browser automation built on the AutoGen framework for autonomous web task execution via natural language.
- WebVoyager — End-to-end multimodal web agent that completes instructions on real live websites with set-of-mark prompting.
- OpenCUA — Open foundation models and framework for computer-use agents, including dataset, benchmark, and end-to-end CUA models.
- computer-use-linux — Rust MCP server for controlling a real Linux desktop from any MCP host. Reads accessibility trees, takes screenshots, and drives clicks, scrolls, and keystrokes across GNOME, KDE, Hyprland, and i3 — Wayland-first.
- nuphus-mcp — Desktop and browser automation MCP server that exposes screen, window, keyboard/mouse, and Chrome control as standard MCP tools over stdio (JSON-RPC 2.0). Any MCP client — Claude Desktop, Cursor, or custom agents — can connect and immediately control the desktop without any network service or installation daemon.
- winremote-mcp — Windows Remote MCP server exposing 40+ tools for desktop automation, process management, and file operations via FastMCP. AI agents connect to control a local or remote Windows machine through natural language, covering window management, mouse/keyboard, and shell execution.