​
Overview

An evaluation is defined by three things: a completed dataset revision, one deployed task @chalkcompute.function version, and one or more deployed scorer @chalkcompute.function versions. Creating an evaluation saves that definition in the active environment. Running it queues a separate execution that calls the task once per dataset row, passes each output to every scorer, and writes the scores to a result dataset.

The definition is immutable, so two runs of the same evaluation stay comparable: the data, the task, and the scorers are fixed at creation time. To evaluate a new version of the task, create a new evaluation.

┌──────────────────┐   define                 execute
│ dataset revision │──┐       ┌──────────────┐       ┌─────────────────┐
│ task version     │──┼──────▸│  Evaluation  │──────▸│  EvaluationRun  │
│ scorer versions  │──┘       │  immutable   │       │   repeatable    │
└──────────────────┘          └──────────────┘       └────────┬────────┘
   pinned inputs                                              │ produces
                                                              ▾
                                                   result dataset revision

Tasks and scorers are ordinary Chalk Functions, so everything on that page applies to them: resource configuration, secrets, retries, rate limits, and tracing.

Evaluations are offline. They score a pinned dataset revision, not live traffic.

Contact our support team to enable evaluations in your environment.


​
Quick start

import chalkcompute
import pyarrow.compute as pc

dataset = chalkcompute.DatasetClient().upload("support-triage", [
    {"message": "I was charged twice for order 1182, please refund",
     "expected": '{"intent": "refund", "amount": 49.99}', "expected_intent": "refund"},
    {"message": "Where is my order? It was due Tuesday.",
     "expected": '{"intent": "status", "amount": 0.0}', "expected_intent": "status"},
    {"message": "Please close my account and refund the last charge",
     "expected": '{"intent": "cancel", "amount": 12.0}', "expected_intent": "cancel"},
    {"message": "Following up on my last email",
     "expected": '{"intent": "status", "amount": 0.0}', "expected_intent": "status"},
])

suite = chalkcompute.EvaluationSuite.create("Release")

@chalkcompute.function
def triage(message: str) -> str:
    import json, re
    text = message.lower()
    if "refund" in text or "money back" in text:
        intent = "refund"
    elif "cancel" in text or "close my account" in text:
        intent = "cancel"
    elif "invoice" in text or "billed" in text or "card" in text:
        intent = "billing"
    elif "where is" in text or "waiting" in text or "arrived" in text:
        intent = "status"
    else:
        return "unclassified"
    amount = re.search(r"\d+(?:\.\d+)?", message)
    return json.dumps({"intent": intent, "amount": float(amount.group()) if amount else 0.0})

evaluation = chalkcompute.Evaluation.create(
    "support-triage",
    dataset=dataset,
    task=triage,
    scorers=[chalkcompute.scorers.json_valid(name="valid-json")],
    suite_id=suite.id,
)

run = evaluation.run(metadata={"git_sha": "abc123"}).wait()
print(run.status)

table = chalkcompute.DatasetClient().read(run.result_dataset)
print(pc.mean(table["valid-json_value"]).as_py())
EvaluationRunStatus.SUCCEEDED
0.75

That runs as written. Rows inline are one of the forms upload accepts, so there is no file to prepare first. A row needs whatever columns the task and scorers read: message for the task, and expected and expected_intent for the scorers further down this page.

triage reads the first number it finds, so the order number in the first row becomes a dollar amount. It checks for “refund” before “cancel”, so the third row is classified as a refund. It returns text that is not JSON when nothing matches, as in the fourth. That last row is the one valid-json marks down, giving the 0.75 above.

Most of this page continues with that example, through scorers and on to reading the results. Where a feature is clearer with something else, such as LLM judges, those sections use an example of their own.


​
Datasets

Upload a dataset with DatasetClient.upload(name, data). It registers a dataset revision and returns a DatasetRevisionRef you pass straight to Evaluation.create. Your data is ordinary tabular data, so this needs neither ChalkPy nor feature definitions.

Pass any of the following as data:

  • a CSV or Parquet path, or a sequence of them, whose Arrow schemas match exactly,
  • a PyArrow table or record batch,
  • a mapping of column name to values, or a sequence of row mappings,
  • any dataframe exposing to_arrow or the dataframe interchange protocol.

CSV files and in-memory values are converted to temporary Parquet files locally before upload. Uploading under a name that already exists creates a new revision of that dataset.

import chalkcompute

client = chalkcompute.DatasetClient()

# From files
revision = client.upload("support-triage", ["jan.parquet", "feb.parquet"])

# From rows, with progress
revision = client.upload(
    "support-triage",
    [{"message": "where is my order?", "expected": '{"intent": "status", "amount": 0.0}'}],
    on_progress=lambda done, total: print(f"{done}/{total}"),
)

upload also accepts part_size_bytes to set the multipart chunk size and timeout to bound each request, in seconds.

Six column names are reserved by the evaluation runtime: output, trace, __chalk_evaluation_output, __chalk_evaluation_row_id, __chalk__session_id__ and __chalk__task_session_id__. A dataset carrying any of them is rejected by Evaluation.create with dataset column "<name>" is reserved by the evaluation runtime.

​
Reading a dataset

DatasetClient.read returns a revision as a PyArrow table. Pass it the DatasetRevisionRef that upload returned, or a revision id:

import chalkcompute

table = chalkcompute.DatasetClient().read(revision)

timeout bounds the request in seconds and defaults to 300. A revision with no output raises DatasetError, which is what you get for an evaluation run that is still going or that failed.

This is how you read a run’s scores as well; see Reading results for the columns a run writes.

​
Referring to a dataset

Pass Evaluation.create(dataset=...) the reference upload returned, or build a DatasetRevisionRef for a dataset that is already there:

import chalkcompute

# What upload returns: the revision it just registered
revision = chalkcompute.DatasetClient().upload("support-triage", "support_messages.csv")

# An exact revision, by id
chalkcompute.DatasetRevisionRef(dataset_id="...", revision_id="...")

# The latest revision of a dataset, resolved when the evaluation is created
chalkcompute.DatasetRevisionRef(dataset_name="support-triage")

A name resolves once, when you create the evaluation, and the resolved revision is then pinned. Later uploads under the same name leave existing evaluations alone. DatasetRevisionRef rejects a mix of the two forms: pass either dataset_name, or dataset_id together with revision_id.

create also takes any object exposing dataset_id and revision_id, which pins that revision, or one exposing dataset_name, whose latest revision it resolves. A dataset handle from another Chalk client therefore passes straight through.


​
Tasks

Write the task as an ordinary @chalkcompute.function. Its parameters bind to your dataset’s columns by name, so triage above declares message: str and requires the dataset to have a message column.

import chalkcompute

@chalkcompute.function
def triage(message: str) -> str:
    ...

@chalkcompute.function starts deploying in the background as soon as the module runs, so consecutive definitions build concurrently. Evaluation.create waits on those handles before reading their immutable version IDs, so you do not need to call deploy() yourself for decorated functions.


​
Scorers

A scorer reads one row and returns a number between 0 and 1. chalkcompute.scorers ships eleven of them, and you can write your own. Most evaluations mix several.

​
Choosing a scorer

The built-in scorers differ in how they arrive at a score:

  • Deterministic comparisons grade one column against another by a fixed rule, covering text, numbers, JSON, and regular expressions. They are cheap and repeatable, and they are the place to start.
  • Model-based scorers judge meaning rather than form. embedding_similarity catches an answer that is right but worded differently, and llm_judge grades qualities you can describe but not compute, such as tone or whether an answer actually helped.
  • SQL expressions score from the dataset’s own columns through scorers.sql, the one built-in that deploys nothing.

Writing your own function is the fourth option, and the only one outside chalkcompute.scorers. Reach for it for business rules, and anything that needs real computation.

​
Built-in scorers

Every built-in is a function you call. Calling it deploys the scorer and returns a handle to pass to scorers=, so call it even when you change none of its options. scorers.sql is the exception, deploying nothing.

llm_judge and sql are configured differently from the rest, and have their own sections below. The options for the other nine follow the table.

import chalkcompute

evaluation = chalkcompute.Evaluation.create(
    "support-triage",
    dataset=dataset,
    task=triage,
    scorers=[
        chalkcompute.scorers.json_valid(inputs=("output",), name="valid-json"),
        chalkcompute.scorers.json_match(inputs=("output", "expected"), name="full-match"),
        chalkcompute.scorers.json_match(
            inputs=("output", "expected"), paths=["intent"], name="intent-match"
        ),
        chalkcompute.scorers.contains(inputs=("output", "expected_intent"), name="mentions-intent"),
    ],
)

On the four rows from the quick start, those four score valid-json 0.75, intent-match 0.5, mentions-intent 0.5 and full-match 0.25. full-match is lowest because it compares the whole document, so the wrong amount in the first row counts against it even though the intent is right.

ScorerWhat it measures
exact_matchTwo columns hold the same text
containsExpected substrings appear in a column
regex_matchA column matches a regular expression
levenshteinEdit distance, normalized into a score
string_similarityA named similarity measure, such as Jaro-Winkler
json_validA column parses as JSON
json_matchTwo columns agree as JSON documents
numeric_closeTwo columns agree as numbers
embedding_similarityTwo columns mean the same thing
llm_judgeA quality you describe in a pydantic model
sqlA Chalk SQL expression you write yourself

Each of the nine keyword-configured scorers takes name, which labels the scorer, and inputs, which names the columns it reads. inputs defaults to the columns that scorer needs: ("output", "expected") for the seven that compare two columns, and ("output",) for regex_match and json_valid, which read one. Any remaining keywords go to @chalkcompute.function, so a built-in takes the same image, resource, and retry options as any other Chalk Function. Each one records metadata explaining its score.

The options worth knowing per scorer:

  • exact_match strips each side by default. normalize_whitespace collapses runs of whitespace instead, and case_insensitive folds case.
  • contains reads one substring, or a JSON array of them, from the reference column. mode scores them together: "all" demands every one, "any" one of them, and "fraction" gives the share that appeared.
  • regex_match takes its pattern positionally, and full_match requires the pattern to cover the whole column. Named groups are recorded as metadata, so a pattern can also pull a value out of the row.
  • string_similarity takes its measure positionally: "jaro_winkler", "jaccard", "token_set", "token_sort", "partial", or "sequence".
  • numeric_close with a tolerance scores 1.0 or 0.0 on whether the two are within it, taken as a share of the larger value when relative. Without one, the score falls off with the gap between them.
  • json_match compares whole documents unless you pass paths, so object key order does not matter but list order does. With paths, written dotted as in "user.id", the score is the share of those paths whose values agree.
  • embedding_similarity embeds both columns with model, defaulting to text-embedding-3-small, and scores them by cosine similarity. It catches an answer that is right but worded differently, which the string scorers mark wrong. base_url and api_key name the endpoint and the Secret holding its key.

​
LLM judges

chalkcompute.scorers.llm_judge deploys a scorer described by a pydantic model (v1 or v2) instead of a function body. The model’s score field becomes the score, and every other field is recorded as row metadata. The judge requests the reply through the OpenAI client using structured outputs and validates it against the model, so a malformed grade fails the row.

import chalkcompute
from pydantic import BaseModel, Field

class TraceQuality(BaseModel):
    score: float = Field(ge=0, le=1, description="Overall quality.")
    directness: int = Field(ge=0, le=2, description="Shortest reasonable path to the goal.")
    task_correctness: int = Field(ge=0, le=2, description="Was the task actually completed?")
    reason: str

trace_quality = chalkcompute.scorers.llm_judge(
    TraceQuality,
    model="gpt-5",
    instructions="You are grading a browser agent's login attempt.",
    api_key=chalkcompute.Secret.from_chalk_env("OPENAI_API_KEY"),
)

The scorer deploys under the model’s class name in kebab case, so TraceQuality becomes trace-quality. Pass name= to choose another. The remaining keywords:

  • inputs names the dataset columns the judge reads, defaulting to ("output",).
  • prompt_fn replaces the default prompt, and parse_fn replaces structured outputs for an endpoint that does not support them.
  • completion_kwargs, for example {"temperature": 0, "max_tokens": 400}, are sent with every request.
  • base_url points at any OpenAI-compatible endpoint.
  • api_key takes a Secret that the judge reads in the pod.
  • image supplies a base image, and any other keyword goes to @chalkcompute.function.

Every function deployed from a module imports that module, so give sibling functions in the same file an image that installs pydantic as well. A judge is generated at import time, which means it cannot be deployed in strip mode.

​
SQL expression scorers

scorers.sql scores a row with a Chalk SQL expression rather than a deployed function. Nothing is deployed, so there is no image to build and no cold start.

import chalkcompute

answered = chalkcompute.scorers.sql(
    "answered",
    '''CASE WHEN "output" IS NOT NULL AND "output" <> '' THEN 1.0 ELSE 0.0 END''',
)

The expression is a select item over the evaluation’s own columns: every dataset column by name, plus output, the task’s result for that row. It must yield a number. name labels the result column and the row in the compare table.

Quote column names, since output and other bare identifiers can collide with SQL keywords.

​
Writing your own scorer

Write a scorer function the same way as a task function. Its parameters bind to dataset columns by name, plus one argument the runner supplies: output, the task’s return value for that row.

import chalkcompute

@chalkcompute.function
def amount_is_plausible(message: str, output: str) -> float:
    import json
    try:
        amount = json.loads(output)["amount"]
    except Exception:
        return 0.0
    return 1.0 if amount == 0.0 or str(amount).rstrip("0").rstrip(".") in message else 0.0

Here message comes from the dataset and output comes from the task. A scorer can read any column in the dataset, not only the one holding the expected answer, which is what lets this one check the extracted amount against the message it came from.

​
Scorer results

A scorer must return a floating point number between 0 and 1 inclusive, one EvaluationScorerResult, or a list[EvaluationScorerResult]. A list lets one scorer emit several scores from shared computation, such as a single model call you grade on three axes. An empty list emits no scores for that row.

import chalkcompute

@chalkcompute.function
def per_field(output: str, expected: str) -> list[chalkcompute.EvaluationScorerResult]:
    import json
    try:
        got, want = json.loads(output), json.loads(expected)
    except ValueError:
        return []
    return [
        chalkcompute.EvaluationScorerResult(
            score=1.0 if got.get(field) == value else 0.0,
            metadata={"field": field},
        )
        for field, value in want.items()
    ]

EvaluationScorerResult takes a score and optional metadata:

  • score must be a finite number between 0 and 1, inclusive. A boolean raises TypeError, and a score outside the range raises ValueError. This check runs when EvaluationScorerResult is constructed, so returning the dataclass is what enforces the range. A scorer returning a plain number, and a SQL expression scorer, are not checked.
  • metadata must be JSON-serializable, with no NaN or infinity. Chalk stores it per row alongside the score.

The return annotation declares the Arrow schema for the function, and EvaluationScorerResult carries its own Arrow contract. Return the dataclass rather than a plain dictionary, which has no such contract and fails to serialize.


​
Attaching deployed functions by reference

Attach functions that already exist rather than redefining them:

import chalkcompute

evaluation = chalkcompute.Evaluation.create(
    "support-triage",
    dataset=dataset,
    task=chalkcompute.RemoteFunction.from_name("triage"),
    scorers=[chalkcompute.RemoteFunction.from_version_id("fn_intent_match_v2")],
)

RemoteFunction.from_name(name) resolves the currently selected version at lookup time, and creating the evaluation then pins that version. A decorated function registers under its Python name, so triage here is the def triage from the quick start. RemoteFunction.from_version_id(id) attaches an exact, immutable version; from_id is a compatibility alias for it.

Deploy an imperative RemoteFunction before you use it. Evaluation.create never deploys on your behalf: an undeployed handle raises EvaluationError before any request goes out, and a bare callable raises TypeError.


​
Creating an evaluation

Pass the dataset, the task, and the scorers to Evaluation.create, along with any metadata you want stored on the definition:

import chalkcompute

evaluation = chalkcompute.Evaluation.create(
    "support-triage",
    dataset=dataset,
    task=triage,
    scorers=[chalkcompute.scorers.json_valid(name="valid-json")],
    metadata={"owner": "support-eng"},
    suite_id=suite.id,
)

name is required and must be non-empty, and at least one scorer is required. metadata must be a JSON-serializable mapping with string keys; Chalk stores it on the definition and returns it on every read. suite_id files the evaluation under a suite.

The returned Evaluation exposes the pinned identifiers as dataset_id, dataset_revision_id, task_function_version_id, and scorers, a tuple of EvaluationScorer carrying an evaluation-local id and then whichever of function_version_id, builtin_id, or sql_name and sql_expression matches how that scorer was defined.

Attach to an existing definition with chalkcompute.Evaluation.from_id("..."), and read the latest server state with evaluation.refresh().


​
Running an evaluation

evaluation.run() is an asynchronous process that queues a run and returns a handle immediately:

run = evaluation.run(metadata={"git_sha": "abc123"})
print(run.status)          # queued, not yet finished
final = run.wait()
print(final.status, final.result_dataset)

wait() polls until the run reaches a terminal state and returns the final handle:

ParameterTypeDefaultDescription
timeoutfloat | None600.0Seconds to wait; None waits indefinitely. Exceeding it raises TimeoutError.
poll_intervalfloat2.0Seconds between status polls.
raise_on_failureboolTrueRaise EvaluationRunFailedError for a failed or canceled run. Pass False to inspect the terminal handle yourself.

EvaluationRunStatus covers PENDING, RUNNING, FINALIZING, SUCCEEDED, FAILED, CANCELED, and UNSPECIFIED. run.is_terminal is true for the three that end a run: SUCCEEDED, FAILED, and CANCELED.

A successful run exposes run.result_dataset, a DatasetRevisionRef for the dataset revision holding the task output and each scorer’s score and metadata per row. It stays None until the run produces one. run.error_message carries the failure detail, and a raised EvaluationRunFailedError exposes the terminal handle as error.run.

Read a definition’s history with evaluation.runs(limit=...), newest first, or attach to one run with chalkcompute.EvaluationRun.from_id("...") and run.refresh().


​
Rescoring a run

The task is usually the slow, expensive half of an evaluation, and it stays the same while you iterate on how its outputs are graded. run.rescore() scores a finished run’s outputs again without calling the task. It reads each row’s output back out of the run’s result dataset and calls only the scorers:

import chalkcompute

run = chalkcompute.EvaluationRun.from_id("...")

# Score the same outputs with new scorers.
rescored = run.rescore(
    scorers=[chalkcompute.scorers.json_valid(name="valid-json"), strict_judge],
).wait()

An evaluation’s scorers are fixed when it is created, so rescore(scorers=...) creates a new evaluation. It copies the run’s dataset revision, task and suite, replaces the scorers, and records the rescored run under it. That evaluation is named "<evaluation name> (rescored)" unless you pass name=. Because the new evaluation is a normal one, later calls to evaluation.run() on it do run the task.

To rescore into an evaluation that already exists, pass it instead:

strict = chalkcompute.Evaluation.from_id("...")
rescored = run.rescore(evaluation=strict)   # or strict.rescore(run)

That evaluation must use the run’s dataset revision and task function version, since the rescored run is credited with that task’s outputs on those rows. The server rejects anything else. run.rescore() with no arguments scores the outputs again with the run’s own evaluation’s scorers, which is useful when a scorer failed on some rows.

Only a run that succeeded has outputs to rescore. The rescored run is an ordinary run of its evaluation: wait(), result_dataset and the dashboard all treat it like any other. Each scorer’s traces go to a session of their own, and a scorer that takes trace still reads the original task’s spans.


​
Reading results

A run writes one row per dataset row into its result dataset. Read it with DatasetClient.read, which returns a PyArrow table:

import chalkcompute

run = evaluation.run().wait()
table = chalkcompute.DatasetClient().read(run.result_dataset)

result_dataset is the DatasetRevisionRef for that run. See Reading a dataset for read’s other arguments.

The result carries the dataset’s own columns, plus:

ColumnHolds
outputthe task’s return value for that row
output_errorthe failure detail, when the task raised
<scorer>_valuethat scorer’s score
<scorer>_errorthe failure detail, when that scorer raised
<scorer>_metadatawhatever the scorer recorded alongside its score

A SQL expression scorer contributes only <name>_value, since nothing is deployed to record an error or metadata against it.

The result also carries __chalk__-prefixed columns that the runtime uses to tie a row to its telemetry. These are internal and not part of the reference above.

table is an ordinary PyArrow table, so summarize or filter it in place:

import pyarrow.compute as pc

pc.mean(table["intent-match_value"]).as_py()      # 0.5, on the rows above
table.filter(pc.equal(table["intent-match_value"], 0.0))

table.to_pandas() works too, if you would rather go on in pandas.

Index a score column with brackets rather than an attribute. A scorer’s name becomes its column prefix, and a built-in names itself in kebab case, so intent-match_value is not a valid Python identifier.

That is the per-row detail. For a run’s average score, its per-scorer averages, and its cost and duration, use the dashboard, which computes those aggregates for you.


​
Comparing runs in the dashboard

Evaluations, runs and scores also appear in the Chalk dashboard, on an environment’s Evaluations page. Three things live there and nowhere else.

Compare evaluations side by side. Each column is one evaluation, averaged over its succeeded runs. Alongside the scores, the table compares cost, tokens, duration, and LLM calls, and reads every value against the first column, where cheaper and faster count as improvements.

Choose a baseline run. At most one run per environment can be the baseline. Picking the run that represents what you ship turns the evaluation into a release gate, since every later run is then read as an improvement or a regression against that fixed point.

Archive a run. Archiving removes a run from the list. A rescored run lists and reads like any other.

The dashboard also shows the summaries behind those comparisons: an average per run, an average per scorer, and how the scores are distributed across a run’s rows.


​
Suites

A suite is a named collection of evaluations and their runs, such as a release gate, a regression set, or one team’s work.

import chalkcompute

suite = chalkcompute.EvaluationSuite.create("Release")

for existing in chalkcompute.EvaluationSuite.list_all():
    print(existing.id, existing.name)

for evaluation in suite.evaluations(limit=20):
    print(evaluation.name)

for run in suite.runs(limit=20):
    print(run.id, run.status)

suite.evaluations() and suite.runs() both yield newest first and page transparently.


​
Using EvaluationClient directly

The class methods above each construct a client for a single call. Hold an EvaluationClient when you make several calls, or when you want it to borrow credentials from a ConnectClient or a ChalkPy client you already have:

import chalkcompute

client = chalkcompute.EvaluationClient(chalk_client=my_chalk_client)

suite = client.create_suite("Release")
evaluation = client.create("Nightly", dataset=dataset, task=answer, scorers=[exact_match])
run = client.run(evaluation.id, metadata={"git_sha": "abc123"})

for candidate in client.list(suite_id=suite.id, limit=50):
    print(candidate.name)

for previous in client.list_runs(evaluation_id=evaluation.id, limit=10):
    print(previous.id, previous.status)

The client carries create_suite and list_suites for suites, create, get, and list for definitions, and run, get_run, list_runs, and rescore for runs. client.rescore(run_id, evaluation_id=...) is the call behind run.rescore(), and evaluation_id defaults to the run’s own evaluation. list and list_runs are generators that page in batches of up to 100 and stop at limit. Pass only one of chalk_client or connect_client to the constructor.

Every object a client returns keeps a reference to that client, which is what makes evaluation.run(), run.refresh(), and suite.evaluations() work. A handle you construct yourself has no client and raises EvaluationError on those methods, so fetch it through the client instead.


​
Errors

A failed request raises EvaluationError, which carries status_code, the HTTP equivalent of the underlying code. A missing evaluation or run raises EvaluationNotFoundError, a subclass of it:

import chalkcompute

try:
    evaluation = chalkcompute.Evaluation.from_id("does-not-exist")
except chalkcompute.EvaluationNotFoundError as e:
    print(e.status_code)  # 404

A run that fails or is canceled raises EvaluationRunFailedError out of wait(), with the terminal handle attached so you can read its status and error message:

try:
    run = evaluation.run().wait()
except chalkcompute.EvaluationRunFailedError as e:
    print(e.run.status, e.run.error_message)

wait() raises TimeoutError when it runs out of time, as does a Trace read that never settles. Local validation raises ValueError or TypeError for an empty name, an empty scorer list, metadata that is not JSON-serializable, a score outside [0, 1], or a bare callable passed as a task or scorer. That validation runs before any request goes out, so a malformed call fails locally.