Check a changed AI feature against a reference run and your written requirements, and keep the evidence for each case. This endpoint gives Claude and other agents access to your RedCrown records over the Model Context Protocol.
prove_taskGive a plain-language task and a few examples. RedCrown runs several models, ranks them
against your quality bar, and returns the ranking with the report's caveats. When the scoring method is
not clear, it runs nothing and asks you to confirm a method with quality_metric. Pass
publish: true to also create a share link. Leave the expected output blank to rank against
the model you use now.
try_sampleReturns a stored example report, with its ranking and caveats. It needs no input, no keys and no setup, and it starts no new run.
Sixteen more advanced tools drive the full loop (import results, scaffold and run experiments, live proxy capture, and the reviewer decision report).
claude mcp add --transport http redcrown https://mcp.redcrown.ai
{
"mcpServers": {
"redcrown": { "type": "http", "url": "https://mcp.redcrown.ai" }
}
}Once connected, try: "Use RedCrown to rank models for classifying these
support tickets, and publish the result." The agent calls prove_task with
publish: true and gives you the share link.