lemonade
AMD-backed local LLM server optimizing and serving open-source models on consumer GPUs and NPUs, exposing them via an MCP server so agents can run inference entirely on-device.
Connects to: AMD GPUs, ONNX Runtime, MCP tools, OpenAI API, Ollama-compatible clients · · Apache-2.0 4,751★
Use it with an AI agent
Loadbay is an MCP server, so an agent can search the catalog and find this harness:
claude mcp add --transport http loadbay https://loadbay.xyz/api/mcp
- Source: https://github.com/lemonade-sdk/lemonade
- This harness as JSON: /api/harnesses/lemonade
- Agent setup: /setup.md
- Browse all 700+ harnesses on Loadbay