Data & search · harness

harvey-labs

Open-source evaluation benchmark and execution harness for scoring LLM agents on realistic legal work tasks across 24+ practice areas, enabling researchers to measure and advance agent capabilities in legal support workflows.

Connects to: LLM agent runtimes, legal task datasets, scoring pipelines · · MIT 1,003★

Use it with an AI agent

Loadbay is an MCP server, so an agent can search the catalog and find this harness:

claude mcp add --transport http loadbay https://loadbay.xyz/api/mcp