Online and offline queries can combine data from multiple sources, and the way data flows through a query can be complex. Chalk has several tools for debugging queries. This page uses them on a few simple features and resolvers.

First, we define the feature classes. The example uses a User and a Transaction object, representing users performing financial transactions. We then write a simple aggregation that has a bug in its definition, and use various tools to debug that issue.

from datetime import datetime
from chalk.features import features, DataFrame

@features
class Transaction:
    id: int
    user_id: "User.id"
    amount: float
    ts: datetime

@features
class User:
    id: int
    sum_transaction_amount: float
    transactions: DataFrame[Transaction]

Then, we define simple online resolvers to ingest user and transaction data from a Postgres database:

users.chalk.sql
-- resolves: User
-- source: postgres

select id from users
transactions.chalk.sql
-- resolves: Transaction
-- source: postgres

select id, user_id, amount, ts from transactions

Next, we write a simple (buggy) aggregation resolver to compute the sum of a user’s transactions:

from chalk import online

@online
def sum_transaction_amount(txns: User.transactions[Transaction.amount]) -> User.sum_transaction_amount:
    return txns[Transaction.amount].count() # this 'count' is a bug! we debug it below.

This aggregation has a bug. It computes the count of the User’s Transactions, rather than the sum of their amounts.

Run a query:

data = ChalkClient().query(input={User.id: 1}, output=[User.sum_transaction_amount])

print(data.get_feature_value(User.sum_transaction_amount))

# Outputs 100

The output is wrong. From prior knowledge of the data, the transaction sum for User 1 is not 100. To debug this issue, use the query plan visualizer.


​
Query plan visualizer

The query plan visualizer shows the structure of a query and the data that flows through it. It is on the query’s details page. For an online query, go to Online > Online queries, select the query, and open the Plan tab. For an offline query, go to Offline > Offline queries and select the query.

To inspect the data at each stage of the query, run the query again with store_plan_stages=True. This stores all of the data that passes through each stage of the query plan. Storing plan stages significantly increases latency, so use it only for debugging.

ChalkClient().query(input={User.id: 1}, output=[User.sum_transaction_amount], store_plan_stages=True)

# To debug multiple users at once:
ChalkClient().offline_query(input={User.id: [1, 2, 3]}, output=[User.sum_transaction_amount], store_plan_stages=True)

With the CLI, pass the --store-plan-stages flag to chalk query.

Here is the plan for the online query. The plan shows the structure of the query: the transactions resolver runs to fetch the transactions for each user, and then the sum_transaction_amount resolver runs to compute the sum. The number on each edge is the row count, and each operator shows its duration.

The visualizer also shows the data that flows through the query. Select the transactions resolver, then open the Outputs tab to see its output. The Stats tab shows summary statistics for the output.

The transaction.amount column shows that the sum of the transaction amounts is far larger than 100. The incorrect output 100 is identical to the count summary of the transactions resolver, which is a hint that the aggregation computes the count of the transactions rather than the sum of their amounts.

If the match with the count is not obvious, more tools are available. The next section walks through how to execute the aggregation locally using the raw input data, so you can use a debugger.

The visualizer displays the first 20 rows of each stage of the query plan. To see more, select Download full output to download the raw parquet file.


​
Resolver replay

Resolver execution can be examined using the query plan visualizer, but Chalk also allows you to directly execute your resolver on your local computer using the same arguments that were used in your query.

To use the resolver_replay functionality:

  1. Run an offline_query with store_plan_stages=True to store the query plan stages, and recompute_features=True to allow resolver execution.
  2. On the returned dataset, call resolver_replay with your resolver function.

Example:

from chalk.client import ChalkClient

ChalkClient().offline_query(
    input={User.id: [1]},
    output=[User.sum_transaction_amount],
    store_plan_stages=True,
    recompute_features=True
).resolver_replay(sum_transaction_amount)

This will execute the resolver locally, using the same input data that was used in the query. This allows you to debug using your IDE (VS Code, PyCharm, or even pdb):

Debugger

You can edit the definition of the resolver locally, and re-run resolver_replay to see the results of your changes in your terminal or by stepping through with a debugger.


​
Using logs to debug

If you cannot use the query plan visualizer or resolver replay, you can also use the logs to debug your query. Use the chalk_logger from the chalk.clogging just like a standard Python logger. This logger will output to the web interface, and to configured log sinks (for example, a Datadog instance).

from chalk.clogging import chalk_logger

@online
def sum_transaction_amount(txns: User.transactions[Transaction.amount]) -> User.sum_transaction_amount:
    chalk_logger.info(f"Transactions: {txns}")
    return txns[Transaction.amount].count() # this 'count' is a bug! we debug it below.

You can also view these logs using resolver_replay, where they will be emitted to your terminal:

Logs output example


​
Query error categories

In using the debugging techniques above, you may encounter a few different categories of errors that have common root causes:

​
Request error

Request errors are raised before your resolver code executes. They are often caused by invalid feature names in the input or by requests that cannot be satisfied by the resolvers you have defined.

​
Feature error

Feature errors are raised when a specific resolver that maps to a feature fails. For this type of error, you will find a feature and resolver attribute in the error type.

When a resolver crashes, you will receive null value in the response. To differentiate from a resolver returning a null value and a failure in the resolver, you need to check the error schema.

​
Network error

Network errors are thrown outside your resolvers. They can be caused by unauthenticated requests or other connection failures.


​
Error schema

The online query interface for resolvers returns the following schema:

​
Response schema

Attributes
dataFeatureResult[]

The outputs features and any query metadata (discussed in detail at Query Basics.)

errorsChalkError[]?

Errors encountered while running the resolvers. Each element in the list is a ChalkError. If no errors were encountered, this field is empty.

​
ChalkError

Attributes
codeErrorCode

The type of error, matching one of the error codes.

categoryErrorCode.kind

The category of the error, given in the type field for the error codes. This will be one of “REQUEST”, “NETWORK”, and “FIELD”.

messagestring

A readable description of the error message.

exceptionobject?

The exception that caused the failure, if applicable.

Child attributes
kindstring

The name of the class of the exception.

messagestring

The message taken from the exception.

stacktracestring

The stacktrace produced by the code.

featurestring?

The fully qualified name of the failing feature, eg. user.identity.has_voip_phone.

resolverstring?

The fully qualified name of the failing resolver, eg. my.project.get_fraud_score.

​
Error codes

Values
PARSE_FAILEDREQUEST

The query contained features that do not exist.

RESOLVER_NOT_FOUNDREQUEST

A resolver was required as part of running the dependency graph that could not be found.

INVALID_QUERYREQUEST

The query is invalid. All supplied features need to be rooted in the same top-level entity.

VALIDATION_FAILEDFIELD

A feature value did not match the expected schema (eg. incompatible type “int”; expected “str”)

RESOLVER_FAILEDFIELD

The resolver for a feature errored.

RESOLVER_TIMED_OUTFIELD

The resolver for a feature timed out.

UPSTREAM_FAILEDFIELD

A crash in a resolver that was to produce an input for the resolver crashed, and so the resolver could not run.

UNAUTHENTICATEDNETWORK

The request was submitted with an invalid authentication header.

UNAUTHORIZEDNETWORK

The request has credentials that do not provide the required authorization to execute an operation.

INTERNAL_SERVER_ERRORNETWORK

An unspecified error occurred.