Observability
Export Chalk metrics to other monitoring systems.
Chalk’s online dashboard provides a simple way to view metrics about performance of your feature pipelines. However, you may wish to export these metrics from Chalk into other observability tools so that you can view your Chalk-related data alongside data from other systems you maintain.
Chalk tracks various time series metrics that measure the latency and throughput of resolvers and streaming pipelines.
Chalk stores these metrics in your cluster. You can use any OpenMetrics-compatible collector to collect metrics about the execution of your feature pipelines from Chalk. Examples include:
The table below summarizes the metrics that are available for export. The headers in the table are the exported metric name followed by the OpenMetrics metric type (gauge, histogram, summary, or counter).
resolver_latency_secondsSummaryProvides information about the time it takes to compute a resolver.
idStringThe name of the resolver, for example, my.company.get_user
quantile0.5 | 0.75 | 0.95 | 0.99Whether this latency represents the median, 75th percentile, 95th percentile, or 99th percentile of the latency
resolver_typeonline | offline | streamThe type of the resolver - online, offline, or stream.
query_latency_secondsSummaryProvides information about the time it takes to execute an online query.
idStringThe name of the query, for example, eligibility_query_v2. Queries without names are
labeled “Unnamed”
quantile0.5 | 0.75 | 0.95 | 0.99Whether this latency represents the median, 75th percentile, 95th percentile, or 99th percentile of the latency
cron_run_latency_secondsSummaryProvides information about the time it takes to execute a cron run.
idStringThe name of the resolver executed by the cron run, for example, my.company.get_user
quantile0.5 | 0.75 | 0.95 | 0.99Whether this latency represents the median, 75th percentile, 95th percentile, or 99th percentile of the latency
feature_requestCounterProvides information about the number of times a feature was computed.
idStringThe name of the feature, for example, user.age
statussuccess | failureThe status of the computed feature (success or failure)
contextinference | cron | migration | streamingThe context in which the feature was generated
resolver_requestCounterProvides information about the number of times a resolver was computed. This metric informs the number of times that resolvers are being called and the context in which they are called, for example in a cron run as part of a scheduled job or in inference as part of a query plan.
idStringThe name of the resolver, for example, my.company.get_user
statussuccess | failureThe status of the resolver run (success or failure)
contextinference | cron | migration | streamingThe context in which the resolver ran
resolver_typeonline | offline | streamThe type of the resolver - online, offline, or stream.
cron_run_requestCounterProvides information about the number of times a cron run was executed. This metric is useful for monitoring the status of resolver runs that are scheduled or triggered via API to load data into the online and/or offline store.
idStringThe name of the resolver executed by the cron run, for example, my.company.get_user
statussuccess | failureThe status of the cron run (success or failure)
cron_feature_writesCounterProvides information about the number of features computed by cron written to online / offline store. This metric is useful for monitoring resolver runs that are scheduled or triggered via API to load data into the online and/or offline store.
idStringThe name of the resolver executed by the cron run, for example, my.company.get_user
contextStringWhether the features were written to online or offline store.
feature_valueSummaryProvides statistical information about the value of features.
idStringThe name of the feature, for example, user.age
quantile0.5 | 0.75 | 0.95 | 0.99Whether this value represents the median, 75th percentile, 95th percentile, or 99th percentile of the feature value
query_requestCounterProvides information about the number of times an online query was executed.
idStringThe name of the query, for example, eligibility_query_v2. Queries without names are
labeled “Unnamed”
statussuccess | failureThe status of the query (success or failure)
deploymentGaugeThe active deployment version. This gauge will always have a value of 1 for active deployments.
A gauge of this kind is sometimes called an Info metric.
idStringThe ID of the deployment.
query_http_responseGaugeThe response counts by HTTP response code.
environmentStringThe ID of the environment.
environment_nameStringThe human-readable name of the environment (for example, prod or staging).
stream_offset_lag_messagesSummaryProvides information about how far streaming message processors are behind, measured in messages.
idStringThe name of the stream resolver.
quantile0.5 | 0.75 | 0.95 | 0.99Whether this lag represents the median, 75th percentile, 95th percentile, or 99th percentile of the offset lag.
environmentStringThe ID of the environment.
environment_nameStringThe human-readable name of the environment (for example, prod or staging).
cpu_utilization_percentageGaugeThe maximum container CPU utilization percentage observed over the recent lookback window, grouped by Chalk service.
environmentStringThe ID of the environment.
environment_nameStringThe human-readable name of the environment (for example, prod or staging).
service_kindhttp | grpc | branchThe Chalk service reporting the metric. http is the HTTP query server,
grpc is the gRPC query server, and branch is the branch server.
memory_utilization_bytesGaugeThe maximum container memory usage in bytes observed over the recent lookback window, grouped by Chalk service.
environmentStringThe ID of the environment.
environment_nameStringThe human-readable name of the environment (for example, prod or staging).
service_kindhttp | grpc | branchThe Chalk service reporting the metric. http is the HTTP query server,
grpc is the gRPC query server, and branch is the branch server.
replica_countGaugeThe maximum number of distinct replicas observed over the recent lookback window, grouped by Chalk service.
environmentStringThe ID of the environment.
environment_nameStringThe human-readable name of the environment (for example, prod or staging).
service_kindhttp | grpc | branchThe Chalk service reporting the metric. http is the HTTP query server,
grpc is the gRPC query server, and branch is the branch server.
resolver_high_water_marksGaugeThe current max_ingested_timestamp in UNIX epoch time for resolvers.
idStringThe ID of the resolver.