Experiments
A/B-test prompt versions head-to-head and pick a statistical winner. Compare a baseline against any number of variants, optionally over a real dataset.
Three ways to start one
All three land the experiment in DRAFT — nothing runs automatically until you press Run.
Manually
Experiments → New Experiment. Pick a prompt (it needs at least one saved version), choose one version as the baseline and any number of other versions as variants, optionally pick a dataset to run them against, name it, and create. Skip the dataset and the prompt just runs once as-is.
From the Playground
After a dataset batch run, use Save as experiment. This saves your current editor content as a new version (the baseline) and creates an experiment against the dataset and row count you just ran.
From Vera
Open a prompt, go to Suggestions, and use “Ask Vera to fix this prompt”. Once Vera proposes a fix you can Stage A/B experiment (one candidate vs. the live prompt) or Search several candidates at once, which stages them all as variants on one experiment for you to compare.
Running and results
Open a DRAFT experiment and click Run experiment. It moves to RUNNING and the page polls automatically until it's done — no manual refresh needed. Each variant executes against every row of the attached dataset (or once, if there's no dataset).
Once COMPLETED, you get:
If a variant wins
Ship to staging promotes immediately. Promote to production respects the prompt's approval gate — it opens a release request instead of shipping directly when approvals are required. If the baseline holds up, there's nothing to ship.
Before you rely on the numbers
Scoring is currently fixed
The model is currently fixed too
gpt-4o-mini, independent of which provider or model your prompt actually targets in production. An OpenAI API key must be connected under Settings → Providers for an experiment to run — even if the prompt itself is served against Anthropic or Google. If a run shows FAILED, this is the first thing to check.Need a dataset to run variants against first? See Evaluations & Datasets.