Chalk’s feature platform enables machine learning teams to focus on building the unique products and models that make their business stand out. Chalk provides a feature store so that you can deploy production machine learning pipelines for real-time data in minutes.

Chalk is both a framework and a platform — developers can write code using familiar Python packages, and deploy their feature and data pipeline definitions to Chalk’s platform. In the Customer Cloud deployment, Chalk runs and administers its platform on the customer’s cloud account. Chalk’s managed infrastructure then executes the customer-defined pipelines to compute feature data for machine learning applications. Chalk then serves this data back to customer applications for online inference and to customer data teams for training set generation.


​
Configuring an Iceberg offline store

The Iceberg offline store uses AWS Glue Catalog, a fully managed metadata catalog for data discovery and schema management in data lakes. This option gives you direct control over your AWS infrastructure and data storage. Managed solutions like Snowflake and Databricks abstract away compute and storage. Your data remains in your S3 buckets.

​
Required information

To set up the Iceberg offline store, you need:

  • S3 bucket name: The S3 bucket where your offline store data will be written. You may reuse an existing bucket or create a new one specifically for offline store usage. This bucket should be in the same AWS region in which Chalk is deployed.

​
Step 1: AWS permissions

Grant the following IAM permissions to the Chalk execution role. Many of these are base permissions that you may have already configured. The key addition is Glue access for catalog operations, including table creation, schema evolution, and metadata management:

{
    "Statement": [
        {
            "Action": [
                "s3:*",
                "dynamodb:*",
                "secretsmanager:*",
                "sqs:*",
                "sts:AssumeRole",
                "glue:*"
            ],
            "Effect": "Allow",
            "Resource": "*"
        }
    ]
}

​
Step 2: Creating the database in Glue Catalog

Before Chalk can use Iceberg for your offline store, you must create an AWS Glue database for the offline store that will serve as the metadata store for your Iceberg tables. You can create the database in your default catalog for the region in which Chalk is deployed.

You may name the database whatever you like, although Chalk recommends a name such as “offline_store”. Note the name, because you need it along with the bucket name above when configuring the offline store.


​
Step 3: Creating and activating the offline store connection

After you set up your S3 bucket and Glue database, configure the connection in the Chalk dashboard:

  1. In the dashboard, go to Integrations > Offline store.
  2. Select New Connection.
  3. For Provider, select Iceberg (Glue + S3).
  4. Fill in the following fields:
    • Connection name: A name to identify this connection (for example, Iceberg Offline Store).
    • S3 bucket name: The name of the S3 bucket from the Required information section above.
    • Glue database name: The name of the Glue database you created in Step 2.
    • AWS account ID (optional): The ID of the AWS account that holds the Glue catalog, if it is not the account that Chalk runs in.
    • Role ARN (optional): A role in that account for Chalk to assume for cross-account access.
  5. Optionally, select Test connection to verify that Chalk can reach your S3 bucket and Glue catalog.
  6. Select Create connection.

After you create the connection, Chalk returns you to the offline store connections list. Select Activate on your new connection to make it the active offline store for this environment.


​
Step 4: Configure AWS Glue optimization policies

AWS Glue provides managed optimizers for Iceberg:

  • Compaction combines small data files into larger files to reduce metadata overhead and improve read and write performance.
  • Snapshot retention expires old snapshots and can remove their associated data and metadata files.
  • Orphan file deletion removes files that are no longer referenced by Iceberg table metadata.

These can be turned on at the catalog level or at the table level from within your AWS console.

After you turn on a catalog-level policy, all tables created after that point inherit the policy, and all existing tables inherit the policy as soon as they are updated. All of Chalk’s Iceberg write operations count as updates. However, any table-level setting (including disabling optimizations) always overrides catalog-wide policies.

We recommend that you enable compaction at a minimum.

  • Iceberg tables accumulate metadata with every write, and this will degrade read and write performance over time.
  • Compaction is particularly helpful if your workloads result in many small writes — which correspond to more datafiles and faster metadata file accumulation.

Many customers also find it valuable to enable a snapshot retention policy and orphan file deletion, but this decision should be based on your own data retention needs.

  • You can set snapshot retention to preserve the recovery and time-travel window your organization requires. Expired snapshots can no longer be queried, and enabling cleanup of expired files makes the associated file deletion permanent unless separately protected by storage recovery controls.
  • Enabling orphan file deletion means that failed or interrupted writes do not leave unused files indefinitely. Use a retention period longer than the longest expected write or maintenance operation so an in-progress file is not mistaken for an orphan. Each table must have a unique, non-overlapping S3 location; otherwise, cleanup for one table can delete files belonging to another table or data source.

Review AWS’s optimizer considerations and limitations before enabling file deletion. The Apache Iceberg maintenance guidance provides additional detail about safe orphan-file retention for concurrent writers.

​
Viewing your current optimization policies within Chalk

You configure all policies in AWS Glue, and Chalk shows your current settings in the Offline store connection page.

The page displays the catalog-level settings automatically. To export a summary of any tables that override the catalog defaults, select Export table overrides, then choose to CSV or to Notebook.

This summary checks the active deployment’s feature tables and the query_log table. It exports only policies whose AWS Glue configuration source is the table itself; tables that inherit every catalog default are omitted. Large catalogs can take several minutes to inspect. The CSV contains:

feature_fqn,table_name,compaction_settings,retention_settings,orphan_file_deletion_settings

​
Querying against the Iceberg offline store

The Iceberg offline store can be accessed using the Chalk SQL Interface. In the dashboard, you can navigate to the SQL Explorer and run SQL queries directly against the historical feature value tables and the query values tables that Chalk maintains in the offline store. In the SDK, you can use the ChalkGRPCClient.run_sql() method to execute the same SQL queries programmatically.