Changelog

August 2026

  • Fix what a shared link shows #375
  • Fail the build when a new page is missing from the sitemap #374
  • Add the Claude and Calibrate tutorial to the Learn page #373
  • Add three slide decks to the Learn page #372
  • Add a why Calibrate section to the landing page #371
  • Add a Learn page with our AI evaluation sessions #370
  • Add a coding agents section to the landing page #368
  • Remember the landing how-to card in the URL #369
  • Add a continuous improvement section to the landing page #367
  • Make the changelog workflow able to push to main #366
  • The changelog page lists what has changed in the app, and gets a new line every time something ships #365
  • See how many items an evaluator run has finished while it is still going #364
  • Duplicate an agent from its own page instead of building a copy by hand #363
  • Mark an evaluator as optional so annotators can finish an item without scoring it #361
  • A page you cannot open now says it is blocked instead of showing a number #360
  • A Google sign-in that does not go through comes back to the Calibrate login page with a message instead of an error page #358
  • See how many tests each model has finished while a benchmark is running #357
  • Traces that only made tool calls are left out when you submit traces for labelling, because there is no reply to score #355
  • Search an agent's traces by anything said in the conversation, the reply, or the trace's own details #354
  • A trace's extra details now read as a table with room for long values, and the traces list shows its count and page controls above the table #353
  • Stop the traces setup steps disappearing while you check for traces, which threw away the key you had just created #352
  • See the conversations your live agent handled on a new Traces tab, open one in full, add them to your tests, or send them for labelling #340
  • The open test run stays in the address bar on the agent page, so a reload or a shared link opens the same run #347
  • The item count and page controls sit directly above the items table on a labelling task #349
  • One click clears a part-filled selection in the labelling dialogs instead of ticking everything first #346
  • Moving to the next item in a labelling task carries on past the end of the page, and your position counts against every item in the task #348
  • Evaluators with no human labels yet are marked on the evaluation run page, so a missing agreement number explains itself #343
  • Send the items your filters leave on screen for review from the labelling job page and the evaluation run page #342
  • The workspace is now part of the address, so a shared link opens in the workspace it came from #337
  • Wait on one loading spinner until your workspace is ready instead of seeing the page load in pieces #335
  • Open your workspace API keys straight from the profile menu or the workspace switcher #334
  • Sign in from a shared link and land on the page it pointed at instead of the agents list #333
  • Read Noora Health's own words on the landing page, with a link to their workshop slides #329
  • Open a link to something kept in another of your workspaces and Calibrate switches to that workspace instead of saying the page is blocked #327
  • Read the rewritten Kabakoo story on the landing page #325
  • Keep your item filters after you reload the page or share the link, and see the score cards again on the task overview #320
  • Filter items by the score an evaluator gave them, on both the evaluation run page and the labelling job page #319
  • See the task overview even when only evaluators have scored the items and no one has labelled them yet #318
  • See each evaluator's overall score on the task overview, beside how often it agrees with the annotators #317
  • See each evaluator's overall score on the evaluation run page, beside how often it agrees with the annotators #312
  • Read what other teams use Calibrate for on the landing page, with the header staying in view as you scroll #303
  • The empty gap above the conversation on an evaluation run item is gone #311
  • Add a new annotator without leaving the assign annotators dialog #310
  • Each answer option on the annotation card now says what it means #309
  • Hear the clip before reading the text on a TTS labelling item, with the audio now above the text #302

July 2026

  • Add custom fields to a connection agent so every request carries them, and give a single test its own values for those fields #291
  • See on the Connection tab how your agent can report the cost, latency, and tokens of each run so Calibrate can show them #301
  • An item with every evaluator answered saves itself when the annotator moves to another item, and a part answered one warns before leaving #300
  • Rank benchmark models by how much you care about cost, quality, and speed, using sliders #298
  • Annotators now see the whole conversation, including any tool the agent called, when a test run is sent for labelling #295
  • Compare speech to text and text to speech providers on quality against cost and speed in a new Model selection tab #278
  • Pick the best value model from a new Top picks tab in benchmark results #286
  • See what each speech to text and text to speech provider cost for a run, with its own price and currency #277
  • The Talk to us button now sits at the bottom of the sidebar instead of floating over the page #285
  • Save an agent with Cmd+S or Ctrl+S #284
  • The About tab on speech to text and text to speech results now explains the Sarvam judge scores, and columns with no scores are hidden #282
  • Check that a connection agent answers before its tests run, with a way to jump to its connection settings when it does not #280
  • See how long each speech to text provider took, as a new Latency column with its own chart and About entry #274
  • Stop the tour card and its highlight hanging over the wrong part of the screen while the next step loads #281
  • Read what each score on a test run or benchmark means in a new About tab #275
  • Stop a rerun on the tests page opening the old run again #273
  • Rerun a test run or a benchmark straight from its results window #266
  • Save a test and run it in one step from the test window #265
  • The address bar now holds the open test, so a reload or a shared link opens the same test #264
  • The guided tour now works when the standard evaluators have been renamed or deleted, reusing what is there or creating what it needs #262
  • New users get a guided tour through their first evaluation #252
  • See Semantic WER on speech to text results whenever the run works it out, with the judge's reasoning on each row #261
  • Pick a task on the Human alignment page straight away, instead of waiting while every task is checked first #258
  • Select several agents on the agents list and delete them in one go #257
  • Delete speech to text and text to speech evaluations, one at a time or several at once #251
  • A leaderboard with only one chart keeps it at half width instead of stretching across #250
  • Run a speech to text evaluation without picking any evaluator #245
  • Switch on built-in judges that score speech to text transcripts on meaning rather than exact word matches, and read their scores and reasoning in the results #247
  • Speech to text and text to speech setup is split into Dataset, Models, and Settings tabs so the models are no longer mixed in with everything else #249
  • The Talk to us button sits at the bottom left of the screen instead of the bottom right #248
  • See CER next to WER on speech to text results #246
  • Edit your workspace's default evaluators, add versions to them and delete them, and find them listed under Default instead of among your own #244
  • Submit rows from a text to speech run for labelling #243
  • The speech to text and text to speech pickers show only the providers your workspace is set up to use #242
  • Create a text to speech labelling task, add items with their audio, and have annotators score them #236
  • Compare the models in a benchmark on a chart of pass rate against cost and speed, and download the chart as an image #241
  • The view switch above a conversation now sits flush at the top while you scroll through a long conversation #240
  • Submit speech to text results and simulation run transcripts for labelling, the way test and benchmark results already could be #235
  • See a spinner while an agent's tabs are still loading, instead of a message saying there is nothing there yet #237
  • Pick which evaluators an agent uses from a new Evaluators tab on the agent page, and new tests start from those evaluators #231
  • Changing which evaluators a labelling task uses now saves in one step, so a failure part way through cannot leave the task with the wrong evaluators #234
  • The Data Extraction tab no longer appears on the agent page #230
  • Benchmark Gemini for speech to text and text to speech, and see for each provider whether it works in real time or sends the whole audio at once #228
  • Retry a speech to text or text to speech run that failed before producing any rows, which used to say it could not be retried #229
  • Evaluator kinds are now called LLM reply and LLM output instead of Conversational reply and LLM response #222
  • Tests are listed and searched by name only, the description no longer shows under each name #221
  • Choose how a test search matches the name: contains, starts with, ends with, or exact #219
  • Filter the tests list by test type: response, tool call, or conversation #220
  • Groq is no longer offered as a speech to text or text to speech provider #218
  • Smallest AI text to speech now uses the newer lightning v3.1 voice by default #217

June 2026

  • Test and benchmark results show a typical response time instead of an average, with the slowest response times noted underneath #213
  • Large speech datasets upload their audio much faster, every row can still play its clip, and long row numbers no longer overflow their circle #211
  • Filter by test type when picking which tests to add to an agent #210
  • Select some tests on the agent Tests tab and compare models on just those tests #209
  • Test and benchmark results show a separate pass rate for tool call tests instead of folding them into the overall one #208
  • Switch a test result's conversation between the chat view and its raw text, and copy the raw text in one click #207
  • Test and benchmark results show which version of the evaluator gave each score #206
  • Send the results of a test run or benchmark run to a labelling task so annotators can score the same replies #205
  • Choose which of a task's evaluators annotators are asked to score when you assign them items #203
  • Create an evaluator that judges a single piece of text rather than a conversation, and build labelling tasks of that kind #201
  • Accept any value for a tool call setting with the new "Is any" match option, instead of naming the exact value #202
  • Test and benchmark results show how long each response took and what it cost #200
  • Choose the evaluators for each row when uploading many tests at once #199
  • The test count on the agent Tests tab now matches the tests you are actually looking at after filtering, and the test run title has more room #198
  • A dot next to the run name shows while a test run is still going, the passed and failed counts sit beside it, and long model names in benchmark results are no longer cut short #196
  • Each test run in the tests list shows how many tests passed, failed, and errored instead of one overall label #195
  • Search test run and benchmark results by test name, and see tests that hit an error grouped apart from the ones that failed #194
  • Add several existing tests to an agent in one go, and close the test dialog straight away when you have not changed anything #193
  • Choose, for each expected tool call argument, whether it must match exactly or be judged by an evaluator #190
  • Stop the test type picker appearing when you duplicate a test, and stop hover labels sticking on screen after a dialog opens #191
  • Mark expected tool call arguments as required or optional, and set values that sit inside other arguments or that the tool does not list #189
  • Switch between the form and a plain text view when adding a tool, so you can paste a whole tool definition at once #188
  • A page you cannot open now says it was not found or that you do not have access, and switching workspace lands you on the list page for the section you were in #187
  • Create and revoke keys for a workspace under Workspace settings, so automated runs can reach Calibrate without anyone signing in #184
  • Stop benchmarks failing to start, or starting twice, when your sign-in was still loading #185
  • See what each tool returned in tool call test results, next to the tool name and its arguments #183
  • The landing page now has a Human alignment section showing how human labels are compared with evaluators #181

May 2026

  • Speech to Text comes first when you pick what kind of labelling task to create #178
  • The tool call list in the Add test dialog no longer gets cut off at the edge of the dialog #177
  • Create a Conversation test that checks a whole conversation, and pick the evaluators that score it #172
  • Duplicate a test or a labelling item straight from its row #173
  • Select a range of labelling items by holding shift and clicking, and land on Tasks when the Human alignment overview has nothing to show yet #169
  • Search the items in a labelling task, move through them a page at a time, and keep the page size you chose #168
  • Drag the evaluators on a labelling task into the order you want annotators to see them #159
  • Leave a comment on an item while labelling, and read every annotator's comments on that item #160
  • Download the items in a labelling task as a CSV file, and refresh them without reloading the page #158
  • Stop the labelling task page showing dashes and jumping to another tab while it is still loading #155
  • Sort the items in a labelling task by when they were last updated, and the choice is kept for next time #152
  • See why renaming an annotator failed right under the name box instead of at the top of the page #151
  • Name the two labels a yes or no evaluator gives, and see those names everywhere its scores appear #148
  • Move to the next or previous item without leaving the item view, pick all annotators at once, and rename an annotator from the list #136
  • Narrow an item down to chosen annotators so you can compare just their labels #133
  • Open any item in a labelling task to see its content, the labels annotators gave it, and what the evaluators scored #131
  • Switch workspaces from the sidebar, create a new one, and manage who belongs to it #117
  • Delete labelling jobs one at a time or several at once, and the benchmarking switch on a connected agent saves by itself #116
  • Add a description to a labelling item, see the time of each turn in a conversation, and upload a file with curly quotes without it failing #107
  • A connected agent that was never verified keeps your changes as you type, and after you verify it again it asks before saving the new settings #92
  • Create and edit tests from an agent's Tests tab, filter them by type, run only the ones you pick, and download speech test results as a zip file #88
  • Share a labelling job or an evaluator run with a link, run an evaluator again, and delete several tests at once #78
  • Filter an evaluator run to the items where the evaluator and the annotators disagreed, and download the run as a spreadsheet #66
  • Uploading items in bulk marks the rows that match items you already have, in red where that annotator's existing labels will be replaced #65
  • See how often an evaluator agreed with the annotators, and what each annotator scored, on the evaluator run page #64
  • Rating buttons while labelling use the evaluator's own scale instead of always 1 to 5, and annotators can see the criteria values for each item #58