The About tab on speech to text and text to speech results now explains the Sarvam judge scores, and columns with no scores are hidden #282
Check that a connection agent answers before its tests run, with a way to jump to its connection settings when it does not #280
See how long each speech to text provider took, as a new Latency column with its own chart and About entry #274
Stop the tour card and its highlight hanging over the wrong part of the screen while the next step loads #281
Read what each score on a test run or benchmark means in a new About tab #275
Stop a rerun on the tests page opening the old run again #273
Rerun a test run or a benchmark straight from its results window #266
Save a test and run it in one step from the test window #265
The address bar now holds the open test, so a reload or a shared link opens the same test #264
The guided tour now works when the standard evaluators have been renamed or deleted, reusing what is there or creating what it needs #262
New users get a guided tour through their first evaluation #252
See Semantic WER on speech to text results whenever the run works it out, with the judge's reasoning on each row #261
Pick a task on the Human alignment page straight away, instead of waiting while every task is checked first #258
Select several agents on the agents list and delete them in one go #257
Delete speech to text and text to speech evaluations, one at a time or several at once #251
A leaderboard with only one chart keeps it at half width instead of stretching across #250
Run a speech to text evaluation without picking any evaluator #245
Switch on built-in judges that score speech to text transcripts on meaning rather than exact word matches, and read their scores and reasoning in the results #247
Speech to text and text to speech setup is split into Dataset, Models, and Settings tabs so the models are no longer mixed in with everything else #249
The Talk to us button sits at the bottom left of the screen instead of the bottom right #248
See CER next to WER on speech to text results #246
Edit your workspace's default evaluators, add versions to them and delete them, and find them listed under Default instead of among your own #244
Submit rows from a text to speech run for labelling #243
The speech to text and text to speech pickers show only the providers your workspace is set up to use #242
Create a text to speech labelling task, add items with their audio, and have annotators score them #236
Compare the models in a benchmark on a chart of pass rate against cost and speed, and download the chart as an image #241
The view switch above a conversation now sits flush at the top while you scroll through a long conversation #240
Submit speech to text results and simulation run transcripts for labelling, the way test and benchmark results already could be #235
See a spinner while an agent's tabs are still loading, instead of a message saying there is nothing there yet #237
Pick which evaluators an agent uses from a new Evaluators tab on the agent page, and new tests start from those evaluators #231
Changing which evaluators a labelling task uses now saves in one step, so a failure part way through cannot leave the task with the wrong evaluators #234
The Data Extraction tab no longer appears on the agent page #230
Benchmark Gemini for speech to text and text to speech, and see for each provider whether it works in real time or sends the whole audio at once #228
Retry a speech to text or text to speech run that failed before producing any rows, which used to say it could not be retried #229
Evaluator kinds are now called LLM reply and LLM output instead of Conversational reply and LLM response #222
Tests are listed and searched by name only, the description no longer shows under each name #221
Choose how a test search matches the name: contains, starts with, ends with, or exact #219
Filter the tests list by test type: response, tool call, or conversation #220
Groq is no longer offered as a speech to text or text to speech provider #218
Smallest AI text to speech now uses the newer lightning v3.1 voice by default #217
June 2026
Test and benchmark results show a typical response time instead of an average, with the slowest response times noted underneath #213
Large speech datasets upload their audio much faster, every row can still play its clip, and long row numbers no longer overflow their circle #211
Filter by test type when picking which tests to add to an agent #210
Select some tests on the agent Tests tab and compare models on just those tests #209
Test and benchmark results show a separate pass rate for tool call tests instead of folding them into the overall one #208
Switch a test result's conversation between the chat view and its raw text, and copy the raw text in one click #207
Test and benchmark results show which version of the evaluator gave each score #206
Send the results of a test run or benchmark run to a labelling task so annotators can score the same replies #205
Choose which of a task's evaluators annotators are asked to score when you assign them items #203
Create an evaluator that judges a single piece of text rather than a conversation, and build labelling tasks of that kind #201
Accept any value for a tool call setting with the new "Is any" match option, instead of naming the exact value #202
Test and benchmark results show how long each response took and what it cost #200
Choose the evaluators for each row when uploading many tests at once #199
The test count on the agent Tests tab now matches the tests you are actually looking at after filtering, and the test run title has more room #198
A dot next to the run name shows while a test run is still going, the passed and failed counts sit beside it, and long model names in benchmark results are no longer cut short #196
Each test run in the tests list shows how many tests passed, failed, and errored instead of one overall label #195
Search test run and benchmark results by test name, and see tests that hit an error grouped apart from the ones that failed #194
Add several existing tests to an agent in one go, and close the test dialog straight away when you have not changed anything #193
Choose, for each expected tool call argument, whether it must match exactly or be judged by an evaluator #190
Stop the test type picker appearing when you duplicate a test, and stop hover labels sticking on screen after a dialog opens #191
Mark expected tool call arguments as required or optional, and set values that sit inside other arguments or that the tool does not list #189
Switch between the form and a plain text view when adding a tool, so you can paste a whole tool definition at once #188
A page you cannot open now says it was not found or that you do not have access, and switching workspace lands you on the list page for the section you were in #187
Create and revoke keys for a workspace under Workspace settings, so automated runs can reach Calibrate without anyone signing in #184
Stop benchmarks failing to start, or starting twice, when your sign-in was still loading #185
See what each tool returned in tool call test results, next to the tool name and its arguments #183
The landing page now has a Human alignment section showing how human labels are compared with evaluators #181
May 2026
Speech to Text comes first when you pick what kind of labelling task to create #178
The tool call list in the Add test dialog no longer gets cut off at the edge of the dialog #177
Create a Conversation test that checks a whole conversation, and pick the evaluators that score it #172
Duplicate a test or a labelling item straight from its row #173
Select a range of labelling items by holding shift and clicking, and land on Tasks when the Human alignment overview has nothing to show yet #169
Search the items in a labelling task, move through them a page at a time, and keep the page size you chose #168
Drag the evaluators on a labelling task into the order you want annotators to see them #159
Leave a comment on an item while labelling, and read every annotator's comments on that item #160
Download the items in a labelling task as a CSV file, and refresh them without reloading the page #158
Stop the labelling task page showing dashes and jumping to another tab while it is still loading #155
Sort the items in a labelling task by when they were last updated, and the choice is kept for next time #152
See why renaming an annotator failed right under the name box instead of at the top of the page #151
Name the two labels a yes or no evaluator gives, and see those names everywhere its scores appear #148
Move to the next or previous item without leaving the item view, pick all annotators at once, and rename an annotator from the list #136
Narrow an item down to chosen annotators so you can compare just their labels #133
Open any item in a labelling task to see its content, the labels annotators gave it, and what the evaluators scored #131
Switch workspaces from the sidebar, create a new one, and manage who belongs to it #117
Delete labelling jobs one at a time or several at once, and the benchmarking switch on a connected agent saves by itself #116
Add a description to a labelling item, see the time of each turn in a conversation, and upload a file with curly quotes without it failing #107
A connected agent that was never verified keeps your changes as you type, and after you verify it again it asks before saving the new settings #92
Create and edit tests from an agent's Tests tab, filter them by type, run only the ones you pick, and download speech test results as a zip file #88
Share a labelling job or an evaluator run with a link, run an evaluator again, and delete several tests at once #78
Filter an evaluator run to the items where the evaluator and the annotators disagreed, and download the run as a spreadsheet #66
Uploading items in bulk marks the rows that match items you already have, in red where that annotator's existing labels will be replaced #65
See how often an evaluator agreed with the annotators, and what each annotator scored, on the evaluator run page #64
Rating buttons while labelling use the evaluator's own scale instead of always 1 to 5, and annotators can see the criteria values for each item #58