Infrastructure
Background Persistence Installation Via the UI
Chalk uses background writers hosted in the (“Customer Cloud”) Kubernetes cluster to write information about queries to various storage locations.
In order to install Chalk persistence writers, you need to have the following:
If using Kafka:
If using Pub/Sub:
Navigate to the Settings > Team > Shared Resources > Background Persistence page in the Chalk UI
to view the background persistence configuration. If no background persistence is configured,
you will see a message indicating that no background persistence is currently present, and the
first save and apply will create background persistence writers.
Chalk supports different types of background persistence writers, each designed for specific data flow and storage purposes:
COPY INTO operations.bigquery-streaming-write-loader is typically used instead.Each writer type requires specific subscription IDs and topics to be configured in the common persistence specifications.
When using Pub/Sub, topics and subscriptions are 2 separate entities, but for Kafka, we use the same topic for both publishing and subscribing. Additionally, we need to provide Kafka authentication credentials, whereas Pub/Sub uses its Google identity to authenticate.
In the JSON format, these fields are in the common_specs field.
bus_backendstringThe backend to use for the bus. “KAFKA” for AWS, “PUBSUB” for GCP.
namespacestringThe namespace to deploy the background persistence writers in.
service_account_namestringThe service account to use for the background persistence writers.
secret_clientstringThe client to use for secrets. “AWS” for AWS Secrets Manager, “GCP” for GCP Secrets Manager.
kafka_dlq_topicstringThe topic to use for Kafka dead letter queue.
api_server_hoststringThe hostname of the API server.
kafka_sasl_secretstringThe cloud secret to use for Kafka authentication.
kafka_bootstrap_serversstringThe Kafka bootstrap servers to use as a comma separated list.
kafka_security_protocolstringThe Kafka security protocol to use.
kafka_sasl_mechanismstringThe Kafka SASL mechanism to use.
redis_is_clusteredstringWhether Redis is clustered or not, if using a Redis online store.
snowflake_storage_integration_namestringThe name of the Snowflake storage integration.
metadata_providerstringThe metadata provider to use (“GRPC_SERVER”)
namestringThe name of the writer.
bus_subscriber_typestringThe type of bus subscriber to use.
requestobjectThe resource requests for the writer.
limitobjectThe resource limits for the writer.
versionstringThe version of the writer.
default_replica_countintThe default number of replicas to create for the writer.
In the JSON format, these fields are in the common_specs field but are not necessarily required.
Writers will each require an image and some, but not all, of the subscription and topic IDs.
In each writer’s specification form, a writer will ask for its required fields and images.
bus_writer_image_gostringThe docker image to use for Go bus writers.
bus_writer_image_pythonstringThe docker image to use for Python bus writers.
bus_writer_image_bswlstringThe docker image to use for BigQuery streaming bus writers.
bigquery_parquet_upload_subscription_idstringThe subscription ID used to read parquet files into the offline store.
bigquery_streaming_write_subscription_idstringThe subscription ID to use for streaming writes to the offline store.
bigquery_streaming_write_topicstringThe topic to use for streaming writes to the offline store.
bigquery_upload_bucketstringThe S3 bucket to use for uploading files to the offline store.
bigquery_upload_topicstringThe topic to use for uploading files to the offline store.
metrics_bus_subscription_idstringThe subscription ID to use for metrics bus.
metrics_bus_topic_idstringThe topic ID to use for metrics bus.
result_bus_metrics_subscription_idstringThe subscription ID to use for result bus metrics.
result_bus_offline_store_subscription_idstringThe subscription ID to use for offline store result bus.
result_bus_online_store_subscription_idstringThe subscription ID to use for online store result bus.
The following is an example configuration for background persistence writers:
{
"common_persistence_specs": {
"bus_backend": "KAFKA",
"bus_writer_image_go": "<go bus writer image>",
"bus_writer_image_python": "<python bus writer image>",
"bus_writer_image_bswl": "<bswl bus writer image>",
"namespace": "background-persistence",
"service_account_name": "background-persistence-sa",
"secret_client": "AWS",
"bigquery_parquet_upload_subscription_id": "offline-store-bulk-insert-bus-1",
"bigquery_streaming_write_subscription_id": "offline-store-streaming-insert-bus-1",
"bigquery_streaming_write_topic": "offline-store-streaming-insert-bus-1",
"bigquery_upload_bucket": "s3://<your data bucket>",
"bigquery_upload_topic": "offline-store-bulk-insert-bus-1",
"metrics_bus_subscription_id": "metrics-bus-1",
"metrics_bus_topic_id": "metrics-bus-1",
"result_bus_metrics_subscription_id": "result-bus-1",
"result_bus_offline_store_subscription_id": "result-bus-1",
"result_bus_online_store_subscription_id": "result-bus-1",
"kafka_dlq_topic": "dlq-1",
"operation_subscription_id": "operation-bus-1"
},
"api_server_host": "<your api server here>",
"kafka_sasl_secret": "<your aws kafka auth secret here>",
"kafka_bootstrap_servers": "<bootstrap server1>:<port>, <bootstrap server2>:<port>, ...",
"kafka_security_protocol": "SASL_SSL",
"kafka_sasl_mechanism": "SCRAM-SHA-512",
"redis_is_clustered": "1",
"snowflake_storage_integration_name": "<snowflak integration name>",
"metadata_provider": "GRPC_SERVER",
"writers": [
{
"name": "go-metrics-bus-writer",
"bus_subscriber_type": "GO_METRICS_BUS_WRITER",
"request": {
"cpu": "200m",
"memory": "512Mi"
},
"limit": {
"cpu": "1",
"memory": "512Mi"
},
"version": "1.0",
"default_replica_count": 1
},
{
"name": "go-result-bus-metrics-writer",
"bus_subscriber_type": "GO_RESULT_BUS_METRICS_WRITER",
"request": {
"cpu": "400m",
"memory": "1024Mi"
},
"limit": {
"cpu": "1",
"memory": "1024Mi"
},
"version": "1.0",
"default_replica_count": 1
}
]
}
If you are experiencing issues with background persistence writers, such as data not appearing in your online or offline stores, writers running out of memory, or pods stuck in a pending state, see the Debugging Persistence guide.