Docs

Eval Runs

The TypeScript SDK's evaluate() one-liner — run your own function against a set of examples and score the outputs, the same shape as LangSmith's evaluate(fn, dataset). Every named run is persisted automatically; there's nothing to set up on the dashboard page itself.

Run it

examples can be an inline array of { input, expectedOutput? }, or a saved dataset's id as a string.

client.evaluations.evaluate()
const results = await client.evaluations.evaluate(
  async ({ input }) => myLLMFunction(input),
  [
    { input: 'What is 2+2?', expectedOutput: '4' },
    { input: 'Capital of France?', expectedOutput: 'Paris' },
  ],
  {
    evaluators: [{ type: 'contains' }],
    experimentName: 'baseline-check', // names this run so it's findable later
    concurrency: 4,
  }
)

console.log(`Pass rate: ${results.passRate}%`)

recordTraces (default true)

Records each example's execution as a real Enprompta trace, tagged with experimentName — this counts against your normal trace quota, the same tradeoff LangSmith's evaluate() makes by default. Set recordTraces: false to run purely in-process with no server calls beyond any ctx.judge() scoring.

Named runs are comparable over time

When both recordTraces and experimentName are set, the aggregate (pass rate, per-scorer breakdown, up to 500 redacted per-example results) is persisted as a dated run under that name — that's what populates the Eval Runs page, filterable by project and searchable by name.

Custom scorers

Mix declarative matchers with your own logic — call an LLM-as-judge evaluator directly from inside a scorer function via ctx.judge().

Custom scorer calling a judge
const results = await client.evaluations.evaluate(
  async ({ input }) => myLLMFunction(input),
  myDataset,
  {
    evaluators: [
      { type: 'non_empty' },
      async (ctx) => {
        const r = await ctx.judge({ evaluatorId: 'ev_groundedness' })
        return { name: 'groundedness', score: r.score, passed: r.passed, reasoning: r.reasoning }
      },
    ],
  }
)

Before you rely on it

Naming collisions

History keys off experimentName verbatim, with no project or date namespacing — two unrelated runs that happen to share a name will have their traces and history blended together. Pick a name specific enough not to collide.

TypeScript-only, for now

This function is currently TypeScript-only — the Python SDK does not yet have an evaluate() equivalent.
Eval Runs - Enprompta