> ## Documentation Index
> Fetch the complete documentation index at: https://penseapp.vercel.app/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Run agent benchmark

> Run a multi-model benchmark on an agent's linked tests as a background job.

## Usage

```python theme={null}
from calibrate import Calibrate

client = Calibrate(
    api_key="your_api_key",
)

client.agent_tests.benchmark(
    agent_uuid="f47ac10b-58cc-4372-a567-0e02b2c3d479",
    models=[
        "openai/gpt-4.1",
        "anthropic/claude-sonnet-4"
    ],
)
```

## API endpoint

See **POST /agent-tests/agent/\{agent\_uuid}/benchmark** in the [API reference](/docs/api-reference/introduction) tab.
