ToolCabana

Agent tools / Agent evaluations

Score exact and normalized answer matches

Compare caller-supplied actual and expected values for every case. Both fields are required; this deterministic comparison does not judge semantic correctness.

Status: implemented. Runtime: browser-or-local. Version: 1.0.0.

Keep deterministic scores separate from model judgments; show rubric, model version, evidence and uncertainty.

Supported profile

This tool accepts the schema below. Its result includes data and versioned runtime metadata. Read warnings in returned data for algorithm coverage and assumptions.

Input schema
{
  "type": "object",
  "properties": {
    "cases": {
      "type": "array",
      "maxItems": 2000,
      "items": {
        "type": "object",
        "properties": {
          "actual": {},
          "expected": {}
        },
        "required": [
          "actual",
          "expected"
        ]
      }
    },
    "normalize": {
      "type": "boolean"
    }
  },
  "required": [
    "cases"
  ],
  "additionalProperties": false
}
Output schema
{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "type": "object",
  "properties": {
    "data": {
      "type": "object",
      "properties": {
        "scores": {
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "passed": {
                "type": "boolean"
              }
            },
            "required": [
              "passed"
            ],
            "additionalProperties": false,
            "maxProperties": 2000
          },
          "maxItems": 2000
        },
        "accuracy": {
          "anyOf": [
            {
              "type": "number",
              "minimum": 0,
              "maximum": 1
            },
            {
              "type": "null"
            }
          ]
        }
      },
      "required": [
        "scores",
        "accuracy"
      ],
      "additionalProperties": false,
      "maxProperties": 2000
    },
    "metadata": {
      "type": "object",
      "properties": {
        "tool": {
          "const": "eval.exact_match"
        },
        "version": {
          "const": "1.0.0"
        },
        "runtime": {
          "type": "string",
          "enum": [
            "local",
            "browser",
            "server"
          ]
        },
        "elapsedMs": {
          "type": "integer",
          "minimum": 0,
          "maximum": 9007199254740991
        },
        "retained": {
          "const": false
        }
      },
      "required": [
        "tool",
        "version",
        "runtime",
        "retained"
      ],
      "additionalProperties": false,
      "maxProperties": 2000
    }
  },
  "required": [
    "data",
    "metadata"
  ],
  "additionalProperties": false,
  "$defs": {
    "json": {
      "anyOf": [
        {
          "type": "null"
        },
        {
          "type": "boolean"
        },
        {
          "type": "number"
        },
        {
          "type": "string",
          "maxLength": 200000
        },
        {
          "type": "array",
          "items": {
            "$ref": "#/$defs/json"
          },
          "maxItems": 2000
        },
        {
          "type": "object",
          "maxProperties": 2000,
          "additionalProperties": {
            "$ref": "#/$defs/json"
          }
        }
      ]
    }
  }
}

Example arguments

{
  "cases": [
    {
      "actual": " A ",
      "expected": "a"
    }
  ],
  "normalize": true
}

Local MCP tool name: eval__exact_match. Browser and local execution work without a key.

Hosted REST and MCP execution are disabled in this release. Run the example in the browser console below, or use the local MCP server from this repository.

Privacy and limits

256 KiB arguments, 512 KiB output, 20 nesting levels and 2,000 array entries. No input/output logging in this layer. Browser/local utilities do not submit inputs remotely. Optional network operations disclose destination requests. Caller-supplied policies and guardrail findings never grant execution privileges.

136 matching tools

Score exact and normalized answer matches

Compare caller-supplied actual and expected values for every case. Both fields are required; this deterministic comparison does not judge semantic correctness.

Keep deterministic scores separate from model judgments; show rubric, model version, evidence and uncertainty.

Contract and limitations

Result

Run a tool to see its structured result.

Let a browser agent use this console

Enable WebMCP when your browser supports it. Inputs run locally and results remain visible here.

Full integration guide