Docs

Experiments

A/B-test prompt versions head-to-head and pick a statistical winner. Compare a baseline against any number of variants, optionally over a real dataset.

Three ways to start one

All three land the experiment in DRAFT — nothing runs automatically until you press Run.

1

Manually

Experiments → New Experiment. Pick a prompt (it needs at least one saved version), choose one version as the baseline and any number of other versions as variants, optionally pick a dataset to run them against, name it, and create. Skip the dataset and the prompt just runs once as-is.

2

From the Playground

After a dataset batch run, use Save as experiment. This saves your current editor content as a new version (the baseline) and creates an experiment against the dataset and row count you just ran.

3

From Vera

Open a prompt, go to Suggestions, and use “Ask Vera to fix this prompt”. Once Vera proposes a fix you can Stage A/B experiment (one candidate vs. the live prompt) or Search several candidates at once, which stages them all as variants on one experiment for you to compare.

Running and results

Open a DRAFT experiment and click Run experiment. It moves to RUNNING and the page polls automatically until it's done — no manual refresh needed. Each variant executes against every row of the attached dataset (or once, if there's no dataset).

Once COMPLETED, you get:

Results table
Runs, pass rate, average score, average latency, and cost — per variant.
Statistical analysis
p-value, confidence level, and effect size, when the sample size supports one.
Reliability
pass@k vs. pass^k when the run repeated each input multiple times — pass@k means at least one of k trials succeeded; pass^k means all k did, the bar for a customer-facing agent.

If a variant wins

Ship to staging promotes immediately. Promote to production respects the prompt's approval gate — it opens a release request instead of shipping directly when approvals are required. If the baseline holds up, there's nothing to ship.

Before you rely on the numbers

Scoring is currently fixed

Every experiment run is scored with a deterministic length check (output between 20 and 8,000 characters) — not by whatever evaluators you've configured elsewhere in Evaluations. Treat the winner as a signal on output length, latency, and cost, not quality, until custom scoring ships.

The model is currently fixed too

Runs execute against OpenAI's gpt-4o-mini, independent of which provider or model your prompt actually targets in production. An OpenAI API key must be connected under Settings → Providers for an experiment to run — even if the prompt itself is served against Anthropic or Google. If a run shows FAILED, this is the first thing to check.

Need a dataset to run variants against first? See Evaluations & Datasets.

Experiments - Enprompta