# Model Deployments
source: https://docs.chalk.ai/docs/model_deployments

## Learn how to deploy and manage machine learning models in Chalk

With Chalk, you can deploy machine learning models as isolated services in dedicated scaling
groups. Your models get their own compute resources, autoscaling policies, and lifecycle
management, so you can scale and update them independently of the Chalk engine.

Model deployments host your models as standalone services that you can call from feature
resolvers or external applications. You can use model inference in your feature pipeline while
managing model serving separately from feature computation.

### When to Use Model Deployments

Model deployments are ideal when you want to:

- Isolate model resources: Give models their own CPU, memory, and GPU resources independent of the engine
- Scale models independently: Auto-scale models based on inference demand without affecting other services
- Version and update models separately: Deploy new model versions without redeploying your entire Chalk system
- Run containerized models: Deploy models as Chalk images or Docker images without converting to Python objects
- Enable high-throughput inference: Run multiple replicas of your model in parallel

### The @model_handler decorator

The fastest way to deploy a model is the @model_handler decorator. You write a single class
with a predict method, hand a trained model object to register_model_version, and Chalk
builds the serving image for you, ships your class as source, serializes and mounts the model,
and wires up the runtime — no Dockerfile, no hand-written Arrow plumbing, no chalkcompute.Image.

```
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class HousePriceModel:
    def predict(self, df):
        return pd.DataFrame({"price": self.model.predict(df.to_pandas())})


rng = np.random.default_rng(0)
X_train = rng.normal(size=(200, 2))  # columns: sqft, rooms
y_train = X_train @ [150.0, 50.0]    # price
rf = RandomForestRegressor().fit(X_train, y_train)

client = ChalkClient()
result = client.register_model_version(
    name="house_price",
    model=HousePriceModel(model=rf),
    input_schema={"sqft": float, "rooms": float},
    output_schema={"price": float},
    dependencies=["scikit-learn", "pandas", "chalkdf"],
)
deployment = client.create_model_deployment(
    name="house-price",
    model_name="house_price",
    model_version=result.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

Pass input_schema, output_schema, and dependencies explicitly. Schemas are
required — naming your columns is what makes the output line up with what
predict returns — and dependencies is the list of pip packages your predict
needs at runtime.

Schema values can be plain Python types — float, int, str, and bool map to
pa.float64(), pa.int64(), pa.string(), and pa.bool_() respectively, so you don't
need to import pyarrow for ordinary tabular models. Reach for a PyArrow type only when
you need one outside those four (for example pa.float32(), pa.large_string(), or a
timestamp type).

The decorated class is a normal Python class with three Chalk-managed attributes injected:

- model: your trained model object. Chalk serializes it with the framework's native
serializer at registration time, uploads it to the model's artifact volume, and deserializes
it back into self.model inside the container before predict runs. The same attribute is
available on both sides.
- files: a list of local file paths (files=["./scaler.pkl"]) that Chalk uploads to the
artifact volume. At runtime self.files is a {basename: Path} mapping, so
self.files["scaler.pkl"] resolves to the mounted path in the container (and to your local
path when testing). Use this for tokenizers, scalers, encoders, lookup tables, etc.
- artifact_path: the mounted artifact volume directory.

### predict and return types

predict(self, df) receives a chalkdf.DataFrame built from the request. Call
df.to_pandas() (or df.to_arrow()) to get the shape your model expects. The return value can
be any of the following — Chalk coerces it to the output columns for you:

| Return type                              | Becomes                      |
| ---------------------------------------- | ---------------------------- |
| `pandas.DataFrame`                       | columns by name              |
| `polars.DataFrame` / `chalkdf.DataFrame` | columns by name              |
| `pyarrow.RecordBatch` / `pyarrow.Table`  | columns by name              |
| `numpy.ndarray` (1-D)                    | a single `prediction` column |
| `numpy.ndarray` (2-D)                    | `col_0`, `col_1`, … columns  |

Returning a named frame (pandas/polars/Arrow) is recommended so your output columns are explicit
and match your output_schema.

### Supported frameworks

The model= object can be any of: scikit-learn, PyTorch, XGBoost, LightGBM,
CatBoost, TensorFlow/Keras, or ONNX. The framework is auto-detected and the right
serializer is used.

### Custom loading with load_model

For models Chalk can't serialize (a custom Python class, a sentence encoder, a lookup table),
leave model=None, ship the artifacts via files=[...], and own the loading in load_model,
which runs once per replica at startup:

```
@model_handler
class CategoryEnricher:
    def load_model(self):
        import pickle
        with open(self.files["table.pkl"], "rb") as f:
            self.table = pickle.load(f)

    def predict(self, df):
        ids = df.to_pandas()["category_id"]
        return pd.DataFrame({"name": [self.table.get(i, "unknown") for i in ids]})
```

When you do have a model= object and also need extra setup (device placement, .eval(),
auxiliary files), call self.default_load_model() inside your override to populate self.model
the default way, then add your own steps:

```
@model_handler
class UserEmbedder:
    def load_model(self):
        self.default_load_model()                    # restores self.model
        self.model.to("cuda").eval()
```

### Local testing

Because the decorated class is plain Python, you can exercise it locally with no container
plumbing — construct it, call load_model() (a no-op outside the container when model= is
already set), and call predict with a chalkdf.DataFrame:

```
import pyarrow as pa
from chalkdf import DataFrame

m = HousePriceModel(model=rf)
m.load_model()
out = m.predict(DataFrame.from_arrow(pa.RecordBatch.from_pydict({"sqft": [1000.0], "rooms": [3.0]})))
```

The predict path hands your code a chalkdf.DataFrame, so the serving image installs chalkdf,
which requires Python below 3.13 (Chalk pins the image to 3.12 automatically). The container also
installs the same chalkpy version you registered from, so the version you run locally must be a
released version that includes @model_handler predict support.

### Examples by model type

Working @model_handler examples for each supported framework. The pattern is identical — write
predict, hand over a trained model, register, deploy — but note that the object you get back as
self.model in the container is whatever that framework's loader returns, which is often not
the class you trained (see the comments in each snippet). Switch tabs to see each one:

```
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class SklearnRF:
    def predict(self, df):
        X = df.to_pandas()[["f0", "f1", "f2", "f3"]]
        return pd.DataFrame({"prediction": self.model.predict(X)})


rng = np.random.default_rng(0)
X_train = rng.normal(size=(200, 4))
y_train = X_train @ [3.0, -2.0, 1.0, 0.5]
rf = RandomForestRegressor().fit(X_train, y_train)

client = ChalkClient()
v = client.register_model_version(
    name="sklearn_rf",
    model=SklearnRF(model=rf),
    input_schema={f"f{i}": float for i in range(4)},
    output_schema={"prediction": float},
    dependencies=["scikit-learn", "pandas", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"sklearn-rf-{v.model_version}", model_name="sklearn_rf", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

```
import numpy as np
import pandas as pd
import xgboost as xgb

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class XGBoostClf:
    def predict(self, df):
        # Reloaded as a raw xgboost.Booster in the container — there is no
        # predict_proba; use the Booster API. binary:logistic returns P(class=1).
        X = df.to_pandas()[["f0", "f1", "f2", "f3"]].to_numpy()
        p = self.model.predict(xgb.DMatrix(X))
        return pd.DataFrame({"p_negative": 1 - p, "p_positive": p})


rng = np.random.default_rng(0)
X_train = rng.normal(size=(200, 4))
y_train = (X_train[:, 0] > 0).astype(int)  # binary label
clf = xgb.XGBClassifier().fit(X_train, y_train)

client = ChalkClient()
v = client.register_model_version(
    name="xgboost_clf",
    model=XGBoostClf(model=clf),
    input_schema={f"f{i}": float for i in range(4)},
    output_schema={"p_negative": float, "p_positive": float},
    dependencies=["xgboost", "pandas", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"xgboost-clf-{v.model_version}", model_name="xgboost_clf", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

```
import pandas as pd
import torch
import torch.nn as nn

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class TorchReg:
    def load_model(self):
        self.default_load_model()  # restores the TorchScript module into self.model
        self.model.eval()

    def predict(self, df):
        X = df.to_pandas()[["f0", "f1", "f2", "f3"]].to_numpy().astype("float32")
        with torch.no_grad():
            out = self.model(torch.from_numpy(X)).numpy().ravel()
        return pd.DataFrame({"prediction": out})


# Must be torch.jit-scriptable — chalkpy serializes via torch.jit.script.
model = nn.Sequential(nn.Linear(4, 8), nn.ReLU(), nn.Linear(8, 1)).eval()

client = ChalkClient()
v = client.register_model_version(
    name="torch_reg",
    model=TorchReg(model=model),
    input_schema={f"f{i}": float for i in range(4)},
    output_schema={"prediction": float},
    dependencies=["torch", "pandas", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"torch-reg-{v.model_version}", model_name="torch_reg", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

```
import numpy as np
import pandas as pd
from catboost import CatBoostClassifier

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class CatBoostClf:
    def predict(self, df):
        # Reloaded as a generic catboost.CatBoost — no predict_proba; use
        # predict(..., prediction_type="Probability").
        X = df.to_pandas()[["f0", "f1", "f2", "f3"]].to_numpy()
        probs = self.model.predict(X, prediction_type="Probability")
        return pd.DataFrame({"p_negative": probs[:, 0], "p_positive": probs[:, 1]})


rng = np.random.default_rng(0)
X_train = rng.normal(size=(200, 4))
y_train = (X_train[:, 0] > 0).astype(int)  # binary label
clf = CatBoostClassifier(verbose=False).fit(X_train, y_train)

client = ChalkClient()
v = client.register_model_version(
    name="catboost_clf",
    model=CatBoostClf(model=clf),
    input_schema={f"f{i}": float for i in range(4)},
    output_schema={"p_negative": float, "p_positive": float},
    dependencies=["catboost", "pandas", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"catboost-clf-{v.model_version}", model_name="catboost_clf", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

```
import numpy as np
import pandas as pd
from skl2onnx import to_onnx
from sklearn.linear_model import LinearRegression

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class OnnxReg:
    def predict(self, df):
        # Reloaded as an onnxruntime.InferenceSession — call .run(), not .predict().
        X = df.to_pandas()[["f0", "f1", "f2", "f3"]].to_numpy().astype("float32")
        out = self.model.run(None, {self.model.get_inputs()[0].name: X})[0]
        return pd.DataFrame({"prediction": np.asarray(out).ravel()})


rng = np.random.default_rng(0)
X_train = rng.normal(size=(200, 4)).astype("float32")
y_train = X_train @ np.array([1.5, -0.5, 2.0, 0.25], dtype="float32")
skl = LinearRegression().fit(X_train, y_train)
onnx_model = to_onnx(skl, X_train[:1])

client = ChalkClient()
v = client.register_model_version(
    name="onnx_reg",
    model=OnnxReg(model=onnx_model),
    input_schema={f"f{i}": float for i in range(4)},
    output_schema={"prediction": float},
    dependencies=["onnxruntime", "numpy", "pandas", "pyarrow", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"onnx-reg-{v.model_version}", model_name="onnx_reg", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

```
import pickle

import pandas as pd

from chalk.client import ChalkClient
from chalk.ml import model_handler
from chalk.scalinggroup import ScalingGroupResourceRequest


@model_handler
class CustomLookup:
    def load_model(self):
        # model=None, so own the loading — no default_load_model() to call.
        with open(self.files["table.pkl"], "rb") as f:
            self.table = pickle.load(f)

    def predict(self, df):
        ids = df.to_pandas()["category_id"]
        return pd.DataFrame({"risk": [self.table.get(c, 0.0) for c in ids]})


# Build and pickle the lookup table this model serves (category_id -> risk).
table = {"groceries": 0.1, "travel": 0.3, "gambling": 0.9}
with open("table.pkl", "wb") as f:
    pickle.dump(table, f)

client = ChalkClient()
# chalkpy can't serialize an arbitrary object — leave model=None, ship artifacts via files=.
v = client.register_model_version(
    name="custom_lookup",
    model=CustomLookup(files=["./table.pkl"]),
    input_schema={"category_id": str},
    output_schema={"risk": float},
    dependencies=["pandas", "pyarrow", "chalkdf"],
)
deployment = client.create_model_deployment(
    name=f"custom-lookup-{v.model_version}", model_name="custom_lookup", model_version=v.model_version,
    resources=ScalingGroupResourceRequest(cpu="1", memory="2Gi"),
)
```

### Container images

When you need full control over the serving container — a custom inference handler, extra system
packages, or a pre-built image — register a model with a container image instead of a decorated
class or a Python model object. Every model image runs the
chalk-remote-call-python shim, which routes
requests and handles PyArrow serialization. You supply a handler and an entrypoint that
points chalk-remote-call at it, then register either a chalkcompute.Image (Chalk builds it) or a
pre-built Docker image (you build it).

### Writing a handler

The handler receives a dict of PyArrow Arrays — one per input_schema column — and returns a
PyArrow Array. Optionally define on_startup to load resources once when the container starts;
artifacts mounted via a volume live at /app/artifacts/ (see
Automatic Volume Upload for Model Artifacts).

```
import json

import pyarrow as pa
import pyarrow.compute as pc

model = None


def on_startup():
    global model
    with open("/app/artifacts/model.json") as f:
        model = json.load(f)


def handler(event: dict[str, pa.Array], context: dict) -> pa.Array:
    factor = model["factor"]
    return pc.multiply(event["x"], pa.scalar(factor, type=pa.float64()))
```

Point the entrypoint at the handler (and optional startup hook):

```
chalk-remote-call --handler model.handler --on-startup model.on_startup --port 8080
```

### Building a Chalk image

Pass a chalkcompute.Image and Chalk builds and manages the container for you. Install
chalk-remote-call-python, add your handler file, and set chalk-remote-call as the entrypoint:

```
from chalk.client import ChalkClient
import chalkcompute

client = ChalkClient()

image = (
    chalkcompute.Image.debian_slim("3.11")
    .pip_install(["chalk-remote-call-python", "pyarrow"])
    .add_local_file("model.py", "/app/model.py", strategy="copy")
    .env({"PYTHONPATH": "/app"})
    .workdir("/app")
    .entrypoint(["chalk-remote-call", "--handler", "model.handler", "--port", "8080"])
)

client.register_model_version(
    name="my-model",
    input_schema={"x": float},
    output_schema={"y": float},
    model_image=image,
)
```

Useful chalkcompute.Image methods:

- .base(image): use a custom base Docker image
- .debian_slim(python_version): slim Debian base with the given Python version
- .pip_install(packages): install Python packages
- .run_commands(commands): run shell commands during the build
- .add_local_file(src, dest, strategy) / .add_local_dir(src, dest, strategy) — copy files in
- .env(vars) / .workdir(path) / .entrypoint(command) — set env vars, working dir, entrypoint

### Using a pre-built Docker image

If you build and push the image yourself, register it by string reference. The image must meet the
same requirements: install chalk-remote-call-python, define a handler (and optional on_startup),
and run chalk-remote-call as the entrypoint.

```
client.register_model_version(
    name="my-model",
    input_schema={"x": float},
    output_schema={"y": float},
    model_image="my-model-image:latest",
)
```

```
FROM python:3.11-slim
WORKDIR /app
RUN pip install --no-cache-dir chalk-remote-call-python pyarrow
COPY model.py /app/model.py
ENV PYTHONPATH=/app
EXPOSE 8080
ENTRYPOINT ["chalk-remote-call", "--handler", "model.handler", "--port", "8080"]
```

```
docker build --platform linux/amd64 -t my-model:latest .
docker push my-model:latest
```

### Automatic Volume Upload for Model Artifacts

When your model files are large (e.g. multi-gigabyte weight files), baking them into the container image is impractical — it slows down builds, increases image pull times, and wastes storage. Instead, Chalk automatically uploads model artifacts to a volume that gets mounted into your container at runtime. If your model artifacts are already baked into the image and you want to skip this automatic upload, pass skip_volume_upload=True during registration:

```
client.register_model_version(
    name="my-large-model",
    input_schema={"x": float},
    output_schema={"y": float},
    model_image="my-large-model-image:latest",
    skip_volume_upload=True,
)
```

The uploaded artifacts are mounted at /app/artifacts/ inside the container. Load them in
on_startup exactly as shown in Writing a handler — open
/app/artifacts/model.json once at startup and reference it from your handler.

### Creating a model deployment

Use ChalkClient.create_model_deployment to deploy a registered model version. It returns a
ModelDeployment with the deployment's ID, name, selected revision_id, status, replica
counts, and public web_url.

```
from chalk.client import ChalkClient
from chalk.scalinggroup import AutoScalingSpec, ScalingGroupResourceRequest

client = ChalkClient()

deployment = client.create_model_deployment(
    name="my-model-sg",
    model_name="my-model",
    model_version=1,
    handler="model.handler",
    scaling=AutoScalingSpec(
        min_replicas=1,
        max_replicas=2,
        target_cpu_utilization_percentage=70,
    ),
    resources=ScalingGroupResourceRequest(cpu="2", memory="4Gi"),
)
print(deployment.id, deployment.status, deployment.web_url)
```

Creation waits for readiness by default, with a wait_timeout of 300 seconds. Set
wait_ready=False to return before the deployment is ready, then call
deployment.wait_ready(timeout=600) when you need to wait. deployment.refresh() returns a
new snapshot of the deployment's state.

To change an existing deployment, use update_model_deployment. Creating with an existing
name also records and selects a new revision of that deployment.

### Auto-Scaling Configuration

Control how your model deployment scales based on demand using AutoScalingSpec.

```
from chalk.scalinggroup import AutoScalingSpec

# Configure auto-scaling behavior
scaling = AutoScalingSpec(
    min_replicas=1,                          # Minimum number of replicas
    max_replicas=5,                          # Maximum number of replicas
    target_cpu_utilization_percentage=70,    # Target CPU utilization (optional)
)
```

Chalk automatically scales the number of replicas based on inference request load and CPU utilization, staying within your min/max bounds. This ensures your models handle traffic spikes efficiently without wasting resources during quiet periods.

### Resource Configuration

Specify CPU, memory, and GPU resources for each replica of your model using ScalingGroupResourceRequest.

```
from chalk.scalinggroup import ScalingGroupResourceRequest

# Request resources per replica
resources = ScalingGroupResourceRequest(
    cpu="2",                          # CPU allocation per replica
    memory="4Gi",                     # Memory allocation per replica
    gpu="nvidia-tesla-t4:1",          # Optional: GPU type and count
)
```

Each replica gets the specified resources. When Chalk scales from 1 to 3 replicas, total resource usage is multiplied accordingly (e.g., 3 replicas × 2 CPU = 6 CPU total).

### Calling Deployed Models

You can call a deployed model three ways: directly from Python with remote() or
defer()/get(), from a feature resolver with F.catalog_call, or from SQL. In every case the
argument order must match your model's input_schema.

### Directly from Python: remote() and defer()/get()

Get a deployment by exactly one of its id or name, then call it directly:

```
from chalk.client import ChalkClient

client = ChalkClient()
deployment = client.get_model_deployment(name="house-price")

price = deployment.remote(sqft=1000.0, rooms=3.0)
price = deployment.remote(1000.0, 3.0)

handle = deployment.defer(sqft=1000.0, rooms=3.0)
price = handle.get(timeout=30)
```

remote() sends one row of feature values and returns the first output value when the model
responds. Pass values by name or position, using the model's input_features order for
positional arguments.

Calls use the deployment's selected revision, so updates and rollbacks apply to subsequent
calls.

defer() enqueues the same inputs on the selected deployment's stored queue and returns a
ModelCallHandle immediately. Pending calls are consumed by the revision selected when they
run. Older revisions can retain a shared, model-wide queue until redeployed; selecting an
older deployment does not isolate its deferred calls from other consumers of that queue.

handle.get() raises ModelRemoteError if the call fails, or TimeoutError if its timeout
expires before a result is available.

### Gather multiple model calls

Submit several calls with defer(), then collect their results with chalk.gather():

```
from chalk import gather
from chalk.client import ChalkClient

deployment = ChalkClient().get_model_deployment(name="house-price")
handles = [
    deployment.defer(sqft=1000.0, rooms=3.0),
    deployment.defer(sqft=1500.0, rooms=4.0),
    deployment.defer(sqft=2000.0, rooms=5.0),
]
prices = gather(handles)
```

gather() waits for the results concurrently and returns a list in the same order as the
handles. In async code, use prices = await gather(handles).

If any call fails, gather() raises the first exception it encounters without returning
partial results. Other remote calls continue running.

### Selecting a deployment through a model version

get_model_version also exposes remote() and defer() on a DeployedModelVersion. Select
a deployment explicitly when a model version is served by several deployments:

```
model = client.get_model_version("house_price", deployment_name="house-price")
price = model.remote(sqft=1000.0, rooms=3.0)
print(model.deployment.id)

model = client.get_model_version(
    "house_price",
    version=model.version,
    deployment_id=model.deployment.id,
)
```

Use deployment_id or deployment_name to choose a deployment; version defaults to the
version it serves. Without a deployment selector, get_model_version uses the requested
version (or the latest registered version) and requires a single active deployment for
inference.

Direct calls require the chalkcompute transport. Install it with pip install
  'chalkpy[compute]'.

### From Resolvers

Address a model deployment by the name passed to create_model_deployment,
referenced as model.{deployment_name}. Calls follow the deployment's selected revision.

### From a Feature Resolver — F.catalog_call

Call the model as part of feature computation, so its output is just another feature:

```
from chalk.features import features, _
from chalk import functions as F

@features
class MyModel:
    id: int
    x_1: float
    x_2: float
    y: float = F.catalog_call(
        "model.my-model-sg",
        _.x_1,
        _.x_2,
    )
```

F.catalog_call is evaluated during chalk query, so the call only takes effect once the
feature graph is applied to your environment with chalk apply.

### Asynchronous Calls — F.catalog_call_async

F.catalog_call_async takes the same arguments as F.catalog_call and produces the same
result, but the engine enqueues the call on the scaling group and polls for the result
instead of holding a blocking connection to the model server for the duration of the call.
Use it for long-running or heavy inference, such as LLM generation or large batch scoring, where
blocking calls would pin engine connections while the model works:

```
y: float = F.catalog_call_async(
    "model.my-model-sg",
    _.x_1,
    _.x_2,
)
```

### From SQL Resolver — chalksql

In a SQL resolver, invoke the model with the catalog_call('model.{deployment_name}', ...)
function, passing the model's qualified name as the first argument:

```
select
    id,
    catalog_call('model.my-model-sg', x_1, x_2) as y
from my_table
```

A deployed model only becomes available to SQL after you redeploy your Chalk deployment with
chalk apply. Registering the model and deploying it to a scaling group is not enough on its own
— the redeploy is what makes the model's qualified name resolvable by catalog_call. If you get
an "unknown function" or unresolved-name error in a SQL resolver right after deploying a model,
run chalk apply and try again.

### Asynchronous SQL calls

catalog_call_async takes the same arguments as catalog_call and is the SQL counterpart of
F.catalog_call_async: we enqueue and poll for a result
instead of holding a connection open for the duration of the call. This is preferred for longer-running inference. Only
model.* and function.* targets can be called like this.

```
select
    id,
    catalog_call_async('model.my-model-sg', x_1, x_2) as y
from my_table
```

### Named arguments

After the call's positional arguments, catalog_call and catalog_call_async accept named
arguments that control how the call is made. They may be combined, and each may be given at
most once.

```
select
    id,
    catalog_call_async(
        'model.my-model-sg', x_1, x_2,
        timeout => interval '90 seconds',
        on_error => 'struct'
    ) as y
from my_table
```

- timeout: how long a single request may take. Write it as an interval literal, or as a
bare number of seconds (timeout => 90). Rows reach the model a batch at a time: one request
per batch, and one request per distinct metadata value within it. Each of those requests
gets the whole budget, so a slow request cannot spend another's. A request that runs
out of time fails; on_error decides what that does to the query. The deadline stops the query
waiting, but does not abort the work: the model is not told, and finishes the request unread.
- on_error ('fail' / 'struct', default 'fail') — what a failed request does.
'fail' propagates it and fails the query, which is what catalog_call has always done.
'struct' reports it in the result instead: the call returns
struct<value, error, started_at, ended_at>, where a successful row has its value and a null
error, and a failed row has a null value and the failure's message. Only the rows that
failed are affected: with catalog_call_async that is the failing row alone, and otherwise the
rows of the request that failed. Every other row, and the query itself, carry on.
started_at and ended_at are UTC timestamps of when the row was sent (for
catalog_call_async, when it was queued, so queue wait is included) and when its result or
failure came back. Note that a model whose
deployment configures retries does not retry under 'struct': the failure is reported rather
than raised, so the retry policy never sees it.

```
-- Keep the rows the model answered, and see why the rest did not.
select
    id,
    y.value as score,
    y.error as failure
from (
    select id, catalog_call('model.my-model-sg', x_1, on_error => 'struct') as y
    from my_table
)
```

### Serving an LLM with vLLM

create_model_deployment runs the image with the chalk-remote-call
entrypoint (Chalk's Arrow request/response runtime), so you can't deploy a stock
vLLM image directly — its own OpenAI HTTP server never starts. Instead,
put the Chalk shim in front of vLLM: ship a handler that loads vLLM at startup and answers
each Arrow request by generating with it. The model then deploys through the normal
registration and model deployment path and is called like any other Chalk model
(F.catalog_call or SQL). This is exactly how curated models are built.

```
import os

import pyarrow as pa

_llm = None


def on_startup():
    global _llm
    from vllm import LLM

    _llm = LLM(model=os.environ.get("VLLM_MODEL", "Qwen/Qwen2.5-0.5B-Instruct"))


def handler(event, context):
    from vllm import SamplingParams

    rb = pa.Table.from_pydict(event).combine_chunks().to_batches()[0]
    prompts = rb.column("prompt").to_pylist()
    outs = _llm.generate(prompts, SamplingParams(max_tokens=64))
    texts = [o.outputs[0].text for o in outs]
    return {"completion": pa.array(texts, type=pa.large_string())}
```

```
import pyarrow as pa
import chalkcompute

from chalk.client import ChalkClient
from chalk.scalinggroup import AutoScalingSpec, ScalingGroupResourceRequest

# vLLM base + the chalk-remote-call runtime + our handler.
image = (
    chalkcompute.Image.base("vllm/vllm-openai:latest")
    .pip_install(["chalk-remote-call-python", "pyarrow"])
    .add_local_file("handler.py", "/app/handler.py", strategy="copy")
    .workdir("/app")
    .env({"PYTHONPATH": "/app"})
)

client = ChalkClient()
result = client.register_model_version(
    name="vllm-qwen",
    model_image=image,
    input_schema={"prompt": pa.large_string()},
    output_schema={"completion": pa.large_string()},
)
deployment = client.create_model_deployment(
    name="vllm-qwen-sg",
    model_name="vllm-qwen",
    model_version=result.model_version,
    scaling=AutoScalingSpec(min_replicas=1, max_replicas=2),
    resources=ScalingGroupResourceRequest(gpu="nvidia-tesla-t4:1"),  # vLLM needs a GPU
    handler="handler.handler",
    env_vars={"PYTHONPATH": "/app"},
)
```

Then call it like any Chalk model — over the Arrow contract, not raw HTTP. For example, from a
feature resolver:

```
from chalk.features import features, _
from chalk import functions as F

@features
class Doc:
    id: int
    prompt: str
    completion: str = F.catalog_call("model.vllm-qwen-sg", _.prompt)
```

Curated models are Chalk's prepackaged version of this idea: a maintained catalog of
open-weight models (e.g. gemma-3-4b-it, mistral-7b-instruct, qwen3-embedding-0-6b,
chronos-2) that serve the same Arrow contract and deploy with one call — no image to build or
maintain. They are not built on this vLLM-in-a-Python-handler recipe, though: the text models
serve through Chalk's native Rust chalk_model_runtime, with -cpu and -gpu image variants
selected at deploy time. The recipe above is how you serve your own vLLM model when a curated one
doesn't fit.

If you instead want vLLM's raw OpenAI-compatible HTTP API — to point an OpenAI client straight at
it — deploy it as a chalkcompute.ScalingGroup, which runs the image's own entrypoint and exposes
the HTTP endpoint directly. See
Model Inference with vLLM.

### Managing deployments

Use the model deployment APIs to inspect a deployment, select a new revision, or delete it.

### Listing deployments

list_model_deployments returns one page. Filter by model_name and, optionally,
model_version. Pass next_cursor back as cursor until it is None:

```
cursor = None
while True:
    page = client.list_model_deployments(model_name="my-model", cursor=cursor, limit=50)
    for item in page.deployments:
        print(item.id, item.name, item.model_version, item.status)
    cursor = page.next_cursor
    if cursor is None:
        break
```

### Updating a deployment

Register the new model version, then replace the deployment's serving spec with
update_model_deployment.

```
import chalkcompute
from chalk.client.model_deployment import ModelDeploymentSpec

new_version = client.register_model_version(
    name="my-model",
    input_schema={"x": float},
    output_schema={"y": float},
    model_image=(
        chalkcompute.Image.debian_slim("3.11")
        .pip_install(["chalk-remote-call-python", "pyarrow"])
        .add_local_file("model_v2.py", "/app/model.py", strategy="copy")
        .env({"PYTHONPATH": "/app"})
        .workdir("/app")
        .entrypoint(["chalk-remote-call", "--handler", "model.handler", "--port", "8080"])
    ),
)

current = client.get_model_deployment(name="my-model-sg")
previous_revision_id = current.revision_id
deployment = client.update_model_deployment(
    id=current.id,
    spec=ModelDeploymentSpec(
        model_version=new_version.model_version,
        scaling=current.scaling,
        resources=current.resources,
        handler="model.handler",
    ),
)
```

ModelDeploymentSpec replaces the entire serving configuration. Supply its required
model_version and scaling, and every optional setting you want to keep. Omitted resource
requests, environment variables, secrets, probes, and workload identity take their defaults;
they are not copied from the current revision. The example preserves the scaling and resources
from the creation example. For a deployment with additional settings, include those settings
from your deployment configuration as well.

Update and rollback wait for readiness by default, and accept wait_ready=False and
wait_timeout just like creation.

### Inspecting revisions and rolling back

list_model_deployment_revisions also returns one page at a time. Use next_cursor to read
all revisions. Each revision has a stable id, a model_version, and a selected flag.
After rollback, the selected revision can be older than the newest revision.

```
cursor = None
while True:
    page = client.list_model_deployment_revisions(
        deployment_id=deployment.id, cursor=cursor, limit=50
    )
    for revision in page.revisions:
        print(revision.id, revision.model_version, revision.selected)
    cursor = page.next_cursor
    if cursor is None:
        break

previous = client.get_model_deployment_revision(
    previous_revision_id, deployment_id=deployment.id
)
deployment = client.rollback_model_deployment(
    id=deployment.id,
    revision_id=previous.id,
)
```

Rollback selects an existing revision's already-built spec without rebuilding its image.
The endpoint and deferred calls follow the selected revision.

### Deleting a deployment

Delete the serving deployment with delete_model_deployment:

```
deleted = client.delete_model_deployment(id=deployment.id)
archived = client.get_model_deployment(id=deleted.id, include_deleted=True)
print(archived.deleted_at)
```

Deletion removes the serving endpoint and leaves archived metadata available with
include_deleted=True. That option is also supported by deployment lists and revision
lookups/lists. Deleting a deployment does not delete the registered model version.

### Structuring Your Model Deployment Code

Model registration and deployment should be controlled manually and separately from your feature definitions. Either:

- Add to .chalkignore to prevent them from running during chalk apply.
- Run in a separate repository dedicated to model management, keeping it independent from your Chalk feature code.

Your chalk apply will fail if it tries to run model registration and deployment code.

Organize your project to keep model management separate from feature definitions:

```
my-chalk-project/
|- models/                          # Model deployment code (add to .chalkignore)
|  |- model.py
|  `- deploy_model.py               # Registration + deployment script
|
|- features/                        # Feature definitions (synced with chalk apply)
|  |- __init__.py
|  `- user_features.py
|
|- .chalkignore
`- chalk.yaml
```

Put the following line in your .chalkignore so chalk apply skips everything under models/.

```
models/
```





