Compute
Evaluate any LLM workload against historical production data and compare outcomes based on criteria you define.
An evaluation is defined by three things: a completed dataset revision, one
deployed task @chalkcompute.function version, and
one or more deployed scorer @chalkcompute.function
versions. Creating an evaluation saves that definition in the active environment.
Running it queues a separate execution that calls the task once per dataset row,
passes each output to every scorer, and writes the scores to a result dataset.
The definition is immutable, so two runs of the same evaluation stay comparable: the data, the task, and the scorers are fixed at creation time. To evaluate a new version of the task, create a new evaluation.
┌──────────────────┐ define execute
│ dataset revision │──┐ ┌──────────────┐ ┌─────────────────┐
│ task version │──┼──────▸│ Evaluation │──────▸│ EvaluationRun │
│ scorer versions │──┘ │ immutable │ │ repeatable │
└──────────────────┘ └──────────────┘ └────────┬────────┘
pinned inputs │ produces
▾
result dataset revision
Tasks and scorers are ordinary Chalk Functions, so everything on that page applies to them: resource configuration, secrets, retries, rate limits, and tracing.
Evaluations are offline. They score a pinned dataset revision, not live traffic.
Contact our support team to enable evaluations in your environment.
import chalkcompute
import pyarrow.compute as pc
dataset = chalkcompute.DatasetClient().upload("support-triage", [
{"message": "I was charged twice for order 1182, please refund",
"expected": '{"intent": "refund", "amount": 49.99}', "expected_intent": "refund"},
{"message": "Where is my order? It was due Tuesday.",
"expected": '{"intent": "status", "amount": 0.0}', "expected_intent": "status"},
{"message": "Please close my account and refund the last charge",
"expected": '{"intent": "cancel", "amount": 12.0}', "expected_intent": "cancel"},
{"message": "Following up on my last email",
"expected": '{"intent": "status", "amount": 0.0}', "expected_intent": "status"},
])
suite = chalkcompute.EvaluationSuite.create("Release")
@chalkcompute.function
def triage(message: str) -> str:
import json, re
text = message.lower()
if "refund" in text or "money back" in text:
intent = "refund"
elif "cancel" in text or "close my account" in text:
intent = "cancel"
elif "invoice" in text or "billed" in text or "card" in text:
intent = "billing"
elif "where is" in text or "waiting" in text or "arrived" in text:
intent = "status"
else:
return "unclassified"
amount = re.search(r"\d+(?:\.\d+)?", message)
return json.dumps({"intent": intent, "amount": float(amount.group()) if amount else 0.0})
evaluation = chalkcompute.Evaluation.create(
"support-triage",
dataset=dataset,
task=triage,
scorers=[chalkcompute.scorers.json_valid(name="valid-json")],
suite_id=suite.id,
)
run = evaluation.run(metadata={"git_sha": "abc123"}).wait()
print(run.status)
table = chalkcompute.DatasetClient().read(run.result_dataset)
print(pc.mean(table["valid-json_value"]).as_py())
EvaluationRunStatus.SUCCEEDED
0.75
That runs as written. Rows inline are one of the forms
upload accepts, so there is no file to prepare first. A row needs
whatever columns the task and scorers read: message for the task, and
expected and expected_intent for the scorers further down this page.
triage reads the first number it finds, so the order number in the first row
becomes a dollar amount. It checks for “refund” before “cancel”, so the third row
is classified as a refund. It returns text that is not JSON when nothing matches,
as in the fourth. That last row is the one valid-json marks down, giving the
0.75 above.
Most of this page continues with that example, through scorers and on to reading the results. Where a feature is clearer with something else, such as LLM judges, those sections use an example of their own.
Upload a dataset with DatasetClient.upload(name, data). It registers a dataset
revision and returns a DatasetRevisionRef you pass straight to
Evaluation.create. Your data is ordinary tabular data, so this needs neither
ChalkPy nor feature definitions.
Pass any of the following as data:
to_arrow or the dataframe interchange protocol.CSV files and in-memory values are converted to temporary Parquet files locally before upload. Uploading under a name that already exists creates a new revision of that dataset.
import chalkcompute
client = chalkcompute.DatasetClient()
# From files
revision = client.upload("support-triage", ["jan.parquet", "feb.parquet"])
# From rows, with progress
revision = client.upload(
"support-triage",
[{"message": "where is my order?", "expected": '{"intent": "status", "amount": 0.0}'}],
on_progress=lambda done, total: print(f"{done}/{total}"),
)
upload also accepts part_size_bytes to set the multipart chunk size and
timeout to bound each request, in seconds.
Six column names are reserved by the evaluation runtime: output, trace,
__chalk_evaluation_output, __chalk_evaluation_row_id,
__chalk__session_id__ and __chalk__task_session_id__. A dataset carrying any
of them is rejected by Evaluation.create with dataset column "<name>" is reserved by the evaluation runtime.
DatasetClient.read returns a revision as a PyArrow table. Pass it the
DatasetRevisionRef that upload returned, or a revision id:
import chalkcompute
table = chalkcompute.DatasetClient().read(revision)
timeout bounds the request in seconds and defaults to 300. A revision with no
output raises DatasetError, which is what you get for an evaluation run that is
still going or that failed.
This is how you read a run’s scores as well; see Reading results for the columns a run writes.
Pass Evaluation.create(dataset=...) the reference upload returned, or build a
DatasetRevisionRef for a dataset that is already there:
import chalkcompute
# What upload returns: the revision it just registered
revision = chalkcompute.DatasetClient().upload("support-triage", "support_messages.csv")
# An exact revision, by id
chalkcompute.DatasetRevisionRef(dataset_id="...", revision_id="...")
# The latest revision of a dataset, resolved when the evaluation is created
chalkcompute.DatasetRevisionRef(dataset_name="support-triage")
A name resolves once, when you create the evaluation, and the resolved revision
is then pinned. Later uploads under the same name leave existing evaluations
alone. DatasetRevisionRef rejects a mix of the two forms: pass either
dataset_name, or dataset_id together with revision_id.
create also takes any object exposing dataset_id and revision_id, which
pins that revision, or one exposing dataset_name, whose latest revision it
resolves. A dataset handle from another Chalk client therefore passes straight
through.
Write the task as an ordinary @chalkcompute.function.
Its parameters bind to your dataset’s columns by name, so triage above
declares message: str and requires the dataset to have a message column.
import chalkcompute
@chalkcompute.function
def triage(message: str) -> str:
...
@chalkcompute.function starts deploying in the background as soon as the
module runs, so consecutive definitions build concurrently. Evaluation.create
waits on those handles before reading their immutable version IDs, so you do not
need to call deploy() yourself for decorated functions.
A scorer reads one row and returns a number between 0 and 1.
chalkcompute.scorers ships eleven of them, and you can write your own. Most
evaluations mix several.
The built-in scorers differ in how they arrive at a score:
embedding_similarity
catches an answer that is right but worded differently, and
llm_judge grades qualities you can describe but not compute,
such as tone or whether an answer actually helped.scorers.sql, the one built-in that deploys
nothing.Writing your own function is the fourth option, and the only one outside
chalkcompute.scorers. Reach for it for business rules, and anything that needs
real computation.
Every built-in is a function you call. Calling it deploys the scorer and returns
a handle to pass to scorers=, so call it even when you change none of its
options. scorers.sql is the exception, deploying nothing.
llm_judge and sql are configured differently from the rest, and have their
own sections below. The options for the other nine follow the table.
import chalkcompute
evaluation = chalkcompute.Evaluation.create(
"support-triage",
dataset=dataset,
task=triage,
scorers=[
chalkcompute.scorers.json_valid(inputs=("output",), name="valid-json"),
chalkcompute.scorers.json_match(inputs=("output", "expected"), name="full-match"),
chalkcompute.scorers.json_match(
inputs=("output", "expected"), paths=["intent"], name="intent-match"
),
chalkcompute.scorers.contains(inputs=("output", "expected_intent"), name="mentions-intent"),
],
)
On the four rows from the quick start, those four score valid-json 0.75,
intent-match 0.5, mentions-intent 0.5 and full-match 0.25. full-match is
lowest because it compares the whole document, so the wrong amount in the first
row counts against it even though the intent is right.
| Scorer | What it measures |
|---|---|
exact_match | Two columns hold the same text |
contains | Expected substrings appear in a column |
regex_match | A column matches a regular expression |
levenshtein | Edit distance, normalized into a score |
string_similarity | A named similarity measure, such as Jaro-Winkler |
json_valid | A column parses as JSON |
json_match | Two columns agree as JSON documents |
numeric_close | Two columns agree as numbers |
embedding_similarity | Two columns mean the same thing |
llm_judge | A quality you describe in a pydantic model |
sql | A Chalk SQL expression you write yourself |
Each of the nine keyword-configured scorers takes name, which labels the
scorer, and inputs, which names the columns it reads. inputs defaults to the
columns that scorer needs: ("output", "expected") for the seven that compare
two columns, and ("output",) for regex_match and json_valid, which read
one. Any remaining keywords go to
@chalkcompute.function, so a built-in takes the same
image, resource, and retry options as any other Chalk Function. Each one records
metadata explaining its score.
The options worth knowing per scorer:
exact_match strips each side by default. normalize_whitespace collapses
runs of whitespace instead, and case_insensitive folds case.contains reads one substring, or a JSON array of them, from the reference
column. mode scores them together: "all" demands every one, "any" one of
them, and "fraction" gives the share that appeared.regex_match takes its pattern positionally, and full_match requires the
pattern to cover the whole column. Named groups are recorded as metadata, so a
pattern can also pull a value out of the row.string_similarity takes its measure positionally: "jaro_winkler",
"jaccard", "token_set", "token_sort", "partial", or "sequence".numeric_close with a tolerance scores 1.0 or 0.0 on whether the two are
within it, taken as a share of the larger value when relative. Without one,
the score falls off with the gap between them.json_match compares whole documents unless you pass paths, so object key
order does not matter but list order does. With paths, written dotted as in
"user.id", the score is the share of those paths whose values agree.embedding_similarity embeds both columns with model, defaulting to
text-embedding-3-small, and scores them by cosine similarity. It catches an
answer that is right but worded differently, which the string scorers mark
wrong. base_url and api_key name the endpoint and the
Secret holding its key.chalkcompute.scorers.llm_judge deploys a scorer described by a pydantic model
(v1 or v2) instead of a function body. The model’s score field becomes the
score, and every other field is recorded as row metadata. The judge requests the
reply through the OpenAI client using structured outputs and validates it against
the model, so a malformed grade fails the row.
import chalkcompute
from pydantic import BaseModel, Field
class TraceQuality(BaseModel):
score: float = Field(ge=0, le=1, description="Overall quality.")
directness: int = Field(ge=0, le=2, description="Shortest reasonable path to the goal.")
task_correctness: int = Field(ge=0, le=2, description="Was the task actually completed?")
reason: str
trace_quality = chalkcompute.scorers.llm_judge(
TraceQuality,
model="gpt-5",
instructions="You are grading a browser agent's login attempt.",
api_key=chalkcompute.Secret.from_chalk_env("OPENAI_API_KEY"),
)
The scorer deploys under the model’s class name in kebab case, so TraceQuality
becomes trace-quality. Pass name= to choose another. The remaining keywords:
inputs names the dataset columns the judge reads, defaulting to ("output",).prompt_fn replaces the default prompt, and parse_fn replaces structured
outputs for an endpoint that does not support them.completion_kwargs, for example {"temperature": 0, "max_tokens": 400}, are
sent with every request.base_url points at any OpenAI-compatible endpoint.api_key takes a Secret that the judge reads in the pod.image supplies a base image, and any other keyword goes to
@chalkcompute.function.Every function deployed from a module imports that module, so give sibling
functions in the same file an image that installs pydantic as well. A judge is
generated at import time, which means it cannot be deployed in strip mode.
scorers.sql scores a row with a Chalk SQL expression rather than a deployed
function. Nothing is deployed, so there is no image to build and no cold start.
import chalkcompute
answered = chalkcompute.scorers.sql(
"answered",
'''CASE WHEN "output" IS NOT NULL AND "output" <> '' THEN 1.0 ELSE 0.0 END''',
)
The expression is a select item over the evaluation’s own columns: every dataset
column by name, plus output, the task’s result for that row. It must yield a
number. name labels the result column and the row in the compare table.
Quote column names, since output and other bare identifiers can collide with
SQL keywords.
Write a scorer function the same way as a task function. Its parameters bind to
dataset columns by name, plus one argument the runner supplies: output, the
task’s return value for that row.
import chalkcompute
@chalkcompute.function
def amount_is_plausible(message: str, output: str) -> float:
import json
try:
amount = json.loads(output)["amount"]
except Exception:
return 0.0
return 1.0 if amount == 0.0 or str(amount).rstrip("0").rstrip(".") in message else 0.0
Here message comes from the dataset and output comes from the task. A scorer
can read any column in the dataset, not only the one holding the expected answer,
which is what lets this one check the extracted amount against the message it
came from.
A scorer must return a floating point number between 0 and 1 inclusive, one EvaluationScorerResult, or a
list[EvaluationScorerResult]. A list lets one scorer emit several scores from
shared computation, such as a single model call you grade on three axes. An
empty list emits no scores for that row.
import chalkcompute
@chalkcompute.function
def per_field(output: str, expected: str) -> list[chalkcompute.EvaluationScorerResult]:
import json
try:
got, want = json.loads(output), json.loads(expected)
except ValueError:
return []
return [
chalkcompute.EvaluationScorerResult(
score=1.0 if got.get(field) == value else 0.0,
metadata={"field": field},
)
for field, value in want.items()
]
EvaluationScorerResult takes a score and optional metadata:
score must be a finite number between 0 and 1, inclusive. A boolean raises
TypeError, and a score outside the range raises ValueError. This check
runs when EvaluationScorerResult is constructed, so returning the dataclass
is what enforces the range. A scorer returning a plain number, and a
SQL expression scorer, are not checked.metadata must be JSON-serializable, with no NaN or infinity. Chalk stores it
per row alongside the score.The return annotation declares the Arrow schema for the function, and
EvaluationScorerResult carries its own Arrow contract. Return the dataclass
rather than a plain dictionary, which has no such contract and fails to
serialize.
Attach functions that already exist rather than redefining them:
import chalkcompute
evaluation = chalkcompute.Evaluation.create(
"support-triage",
dataset=dataset,
task=chalkcompute.RemoteFunction.from_name("triage"),
scorers=[chalkcompute.RemoteFunction.from_version_id("fn_intent_match_v2")],
)
RemoteFunction.from_name(name) resolves the currently selected version at
lookup time, and creating the evaluation then pins that version. A decorated
function registers under its Python name, so triage here is the def triage
from the quick start.
RemoteFunction.from_version_id(id) attaches an exact, immutable version;
from_id is a compatibility alias for it.
Deploy an imperative RemoteFunction before you use it. Evaluation.create
never deploys on your behalf: an undeployed handle raises EvaluationError
before any request goes out, and a bare callable raises TypeError.
Pass the dataset, the task, and the scorers to Evaluation.create, along with
any metadata you want stored on the definition:
import chalkcompute
evaluation = chalkcompute.Evaluation.create(
"support-triage",
dataset=dataset,
task=triage,
scorers=[chalkcompute.scorers.json_valid(name="valid-json")],
metadata={"owner": "support-eng"},
suite_id=suite.id,
)
name is required and must be non-empty, and at least one scorer is required.
metadata must be a JSON-serializable mapping with string keys; Chalk stores it
on the definition and returns it on every read. suite_id files the evaluation
under a suite.
The returned Evaluation exposes the pinned identifiers as dataset_id,
dataset_revision_id, task_function_version_id, and scorers, a tuple of
EvaluationScorer carrying an evaluation-local id and then whichever of
function_version_id, builtin_id, or sql_name and sql_expression matches
how that scorer was defined.
Attach to an existing definition with chalkcompute.Evaluation.from_id("..."),
and read the latest server state with evaluation.refresh().
evaluation.run() is an asynchronous process that queues a run and returns a handle immediately:
run = evaluation.run(metadata={"git_sha": "abc123"})
print(run.status) # queued, not yet finished
final = run.wait()
print(final.status, final.result_dataset)
wait() polls until the run reaches a terminal state and returns the final
handle:
| Parameter | Type | Default | Description |
|---|---|---|---|
timeout | float | None | 600.0 | Seconds to wait; None waits indefinitely. Exceeding it raises TimeoutError. |
poll_interval | float | 2.0 | Seconds between status polls. |
raise_on_failure | bool | True | Raise EvaluationRunFailedError for a failed or canceled run. Pass False to inspect the terminal handle yourself. |
EvaluationRunStatus covers PENDING, RUNNING, FINALIZING, SUCCEEDED,
FAILED, CANCELED, and UNSPECIFIED. run.is_terminal is true for the three
that end a run: SUCCEEDED, FAILED, and CANCELED.
A successful run exposes run.result_dataset, a DatasetRevisionRef for the
dataset revision holding the task output and each scorer’s score and metadata per
row. It stays None until the run produces one. run.error_message carries the
failure detail, and a raised EvaluationRunFailedError exposes the terminal
handle as error.run.
Read a definition’s history with evaluation.runs(limit=...), newest first, or
attach to one run with chalkcompute.EvaluationRun.from_id("...") and
run.refresh().
The task is usually the slow, expensive half of an evaluation, and it stays the
same while you iterate on how its outputs are graded. run.rescore() scores a
finished run’s outputs again without calling the task. It reads each row’s
output back out of the run’s result dataset and calls only the scorers:
import chalkcompute
run = chalkcompute.EvaluationRun.from_id("...")
# Score the same outputs with new scorers.
rescored = run.rescore(
scorers=[chalkcompute.scorers.json_valid(name="valid-json"), strict_judge],
).wait()
An evaluation’s scorers are fixed when it is created, so rescore(scorers=...)
creates a new evaluation. It copies the run’s dataset revision, task and suite,
replaces the scorers, and records the rescored run under it. That evaluation is
named "<evaluation name> (rescored)" unless you pass name=. Because the new
evaluation is a normal one, later calls to evaluation.run() on it do run the
task.
To rescore into an evaluation that already exists, pass it instead:
strict = chalkcompute.Evaluation.from_id("...")
rescored = run.rescore(evaluation=strict) # or strict.rescore(run)
That evaluation must use the run’s dataset revision and task function version,
since the rescored run is credited with that task’s outputs on those rows. The
server rejects anything else. run.rescore() with no arguments scores the
outputs again with the run’s own evaluation’s scorers, which is useful when a
scorer failed on some rows.
Only a run that succeeded has outputs to rescore. The rescored run is an
ordinary run of its evaluation: wait(), result_dataset and the dashboard all
treat it like any other. Each scorer’s traces go to a session of their own, and
a scorer that takes trace still reads the original task’s spans.
A run writes one row per dataset row into its result dataset. Read it with
DatasetClient.read, which returns a PyArrow table:
import chalkcompute
run = evaluation.run().wait()
table = chalkcompute.DatasetClient().read(run.result_dataset)
result_dataset is the DatasetRevisionRef for that run. See
Reading a dataset for read’s other arguments.
The result carries the dataset’s own columns, plus:
| Column | Holds |
|---|---|
output | the task’s return value for that row |
output_error | the failure detail, when the task raised |
<scorer>_value | that scorer’s score |
<scorer>_error | the failure detail, when that scorer raised |
<scorer>_metadata | whatever the scorer recorded alongside its score |
A SQL expression scorer contributes only
<name>_value, since nothing is deployed to record an error or metadata against
it.
The result also carries __chalk__-prefixed columns that the runtime uses to
tie a row to its telemetry. These are internal and not part of the reference
above.
table is an ordinary PyArrow table, so summarize or filter it in place:
import pyarrow.compute as pc
pc.mean(table["intent-match_value"]).as_py() # 0.5, on the rows above
table.filter(pc.equal(table["intent-match_value"], 0.0))
table.to_pandas() works too, if you would rather go on in pandas.
Index a score column with brackets rather than an attribute. A scorer’s name
becomes its column prefix, and a built-in names itself in kebab case, so
intent-match_value is not a valid Python identifier.
That is the per-row detail. For a run’s average score, its per-scorer averages, and its cost and duration, use the dashboard, which computes those aggregates for you.
Evaluations, runs and scores also appear in the Chalk dashboard, on an environment’s Evaluations page. Three things live there and nowhere else.
Compare evaluations side by side. Each column is one evaluation, averaged over its succeeded runs. Alongside the scores, the table compares cost, tokens, duration, and LLM calls, and reads every value against the first column, where cheaper and faster count as improvements.
Choose a baseline run. At most one run per environment can be the baseline. Picking the run that represents what you ship turns the evaluation into a release gate, since every later run is then read as an improvement or a regression against that fixed point.
Archive a run. Archiving removes a run from the list. A rescored run lists and reads like any other.
The dashboard also shows the summaries behind those comparisons: an average per run, an average per scorer, and how the scores are distributed across a run’s rows.
A suite is a named collection of evaluations and their runs, such as a release gate, a regression set, or one team’s work.
import chalkcompute
suite = chalkcompute.EvaluationSuite.create("Release")
for existing in chalkcompute.EvaluationSuite.list_all():
print(existing.id, existing.name)
for evaluation in suite.evaluations(limit=20):
print(evaluation.name)
for run in suite.runs(limit=20):
print(run.id, run.status)
suite.evaluations() and suite.runs() both yield newest first and page
transparently.
The class methods above each construct a client for a single call. Hold an
EvaluationClient when you make several calls, or when you want it to borrow
credentials from a ConnectClient or a ChalkPy client you already have:
import chalkcompute
client = chalkcompute.EvaluationClient(chalk_client=my_chalk_client)
suite = client.create_suite("Release")
evaluation = client.create("Nightly", dataset=dataset, task=answer, scorers=[exact_match])
run = client.run(evaluation.id, metadata={"git_sha": "abc123"})
for candidate in client.list(suite_id=suite.id, limit=50):
print(candidate.name)
for previous in client.list_runs(evaluation_id=evaluation.id, limit=10):
print(previous.id, previous.status)
The client carries create_suite and list_suites for suites, create, get,
and list for definitions, and run, get_run, list_runs, and rescore for
runs. client.rescore(run_id, evaluation_id=...) is the call behind
run.rescore(), and evaluation_id defaults to the run’s own evaluation.
list and list_runs are generators that page in batches of up to 100 and stop
at limit. Pass only one of chalk_client or connect_client to the
constructor.
Every object a client returns keeps a reference to that client, which is what
makes evaluation.run(), run.refresh(), and suite.evaluations() work. A
handle you construct yourself has no client and raises EvaluationError on those
methods, so fetch it through the client instead.
A failed request raises EvaluationError, which carries status_code, the HTTP
equivalent of the underlying code. A missing evaluation or run raises
EvaluationNotFoundError, a subclass of it:
import chalkcompute
try:
evaluation = chalkcompute.Evaluation.from_id("does-not-exist")
except chalkcompute.EvaluationNotFoundError as e:
print(e.status_code) # 404
A run that fails or is canceled raises EvaluationRunFailedError out of
wait(), with the terminal handle attached so you can read its status and error
message:
try:
run = evaluation.run().wait()
except chalkcompute.EvaluationRunFailedError as e:
print(e.run.status, e.run.error_message)
wait() raises TimeoutError when it runs out of time, as does a Trace read
that never settles. Local validation raises ValueError or TypeError for an
empty name, an empty scorer list, metadata that is not JSON-serializable, a score
outside [0, 1], or a bare callable passed as a task or scorer. That validation
runs before any request goes out, so a malformed call fails locally.