Docs
Eval Runs
The TypeScript SDK's evaluate() one-liner — run your own function against a set of examples and score the outputs, the same shape as LangSmith's evaluate(fn, dataset). Every named run is persisted automatically; there's nothing to set up on the dashboard page itself.
Run it
examples can be an inline array of { input, expectedOutput? }, or a saved dataset's id as a string.
client.evaluations.evaluate()
const results = await client.evaluations.evaluate(
async ({ input }) => myLLMFunction(input),
[
{ input: 'What is 2+2?', expectedOutput: '4' },
{ input: 'Capital of France?', expectedOutput: 'Paris' },
],
{
evaluators: [{ type: 'contains' }],
experimentName: 'baseline-check', // names this run so it's findable later
concurrency: 4,
}
)
console.log(`Pass rate: ${results.passRate}%`)recordTraces (default true)
Records each example's execution as a real Enprompta trace, tagged with
experimentName — this counts against your normal trace quota, the same tradeoff LangSmith's evaluate() makes by default. Set recordTraces: false to run purely in-process with no server calls beyond any ctx.judge() scoring.Named runs are comparable over time
When both
recordTraces and experimentName are set, the aggregate (pass rate, per-scorer breakdown, up to 500 redacted per-example results) is persisted as a dated run under that name — that's what populates the Eval Runs page, filterable by project and searchable by name.Custom scorers
Mix declarative matchers with your own logic — call an LLM-as-judge evaluator directly from inside a scorer function via ctx.judge().
Custom scorer calling a judge
const results = await client.evaluations.evaluate(
async ({ input }) => myLLMFunction(input),
myDataset,
{
evaluators: [
{ type: 'non_empty' },
async (ctx) => {
const r = await ctx.judge({ evaluatorId: 'ev_groundedness' })
return { name: 'groundedness', score: r.score, passed: r.passed, reasoning: r.reasoning }
},
],
}
)Before you rely on it
Naming collisions
History keys off
experimentName verbatim, with no project or date namespacing — two unrelated runs that happen to share a name will have their traces and history blended together. Pick a name specific enough not to collide.TypeScript-only, for now
This function is currently TypeScript-only — the Python SDK does not yet have an
evaluate() equivalent.