​
Overview

Chalk Machine Types provide a simple and robust interface for configuring cloud resource consumption for workloads running in your Chalk environment. They are a replacement for the legacy configuration system that required more direct configuration of the underlying Kubernetes API and provide improvements in performance, simplicity, and durability of your workloads, as well as enhanced usage reporting.


​
What is a Chalk Machine Type?

Chalk Machine Types are a simple way of categorizing cloud instances of different sizes, based on the amount of CPU and memory they have available. For instance, small corresponds to a cloud instance with 4 vCPU and 16GiB of memory. Each Chalk Machine Type corresponds to a single concrete instance type in the appropriate cloud provider backend, along with an ordered list of fallback instance types that will automatically be used if the primary instance type is not available due to a cloud provider stockout or other availability issue (like regional availability restrictions). For an online workload, small in an AWS region may map to m8a.xlarge, then fall back to m8i.xlarge, m7a.xlarge, and m7i.xlarge in that order when an instance type is unavailable. This list of fallbacks is configurable per-cluster and per-environment.

When you pick a Chalk Machine Type, your workload is guaranteed to run isolated from other Chalk workloads on its own cloud instance, with appropriate requests and limits automatically configured to ensure that it will schedule and that it can make full use of the resources on the machine.

​
Usage Labeling for Chalk Machine Types

Workloads configured with a Chalk Machine Type will run on Karpenter nodepools created and configured by Chalk in AWS and Azure, and will run on equivalent ComputeClasses in GCP. These nodepools and ComputeClasses will provide appropriate labeling for the Chalk credits consumed by these workloads, depending on the type of workload: online, offline, compute, infrastructure, etc. These nodepools and ComputeClasses are not directly configurable and are automatically created and configured by Chalk.

​
Disk Spilling for Offline Workloads

For offline workloads, using a Chalk Machine Type will automatically enable disk spilling via local SSDs. Offline workloads will be automatically configured to take advantage of this feature, which significantly reduces memory usage for large workloads, reducing their cost and improving their efficiency. No manual configuration will be required to set up disk spilling when using a Chalk Machine Type.

​
GPU Support for Online Workloads

Some inference and training workloads require access to a GPU on the selected instance type. GPU instance support for Chalk Machine Types is not yet enabled but will be available soon.

​
Chalk Machine Types and Resource Groups

Chalk Machine Types integrate nicely with the high-level categorization and isolation provided by Resource Groups. Using a Chalk Machine Type already provides inter-service isolation, and moving workloads with different logical isolation requirements (often useful for separating different business use-cases) to different resource groups allows both for isolation between use-cases and also allows you to select different machine sizes according to the compute needs of each different use-case.

For offline workloads like the Job Queue Consumer, it can be convenient to make several different Resource Groups, each configured to run the consumer against a different Chalk Machine Type, and potentially each with different scaling limits. Then, workloads like offline queries can be directed at the appropriate Resource Group depending on the resources they need, and resource consumption can be controlled via the different scaling configurations.


​
Customizing Chalk Machine Types

Chalk Machine Types are mapped to a primary instance type along with an ordered list of fallback instance types, depending on regional availability. The default list of instance types can be customized in the organization-level settings page for viewing and configuring attached clusters. You can change the order in which different instance types will be chosen as fallbacks, and you can change which instance types are available at all for a given Chalk Machine Type. This can be useful, for example, if you want to ensure that your environment always tries to use the fastest instance types available for maximum performance, or less expensive instance types for cost savings.

Customizing Chalk Machine Type definitions requires organization-level (team-level) permissions.

Browse the default instance and fallback catalog to see the built-in lists for each cloud provider, workload, machine size, disk spilling setting, and GPU setting. Regional availability and your cluster or environment overrides can change the list used by your workloads.


​
Migrating to Chalk Machine Types

Migrating to Chalk Machine Types from a custom configuration is a fairly simple process in most cases. Select the machine type that corresponds most closely to your currently-configured CPU and memory requests, and click Save and apply. Note that the minimum size of machine type that is suitable for production workloads is small-highcpu with 4 CPU and 8Gi of memory; the micro machine type is not recommended for production workloads.

There are a few edge cases and caveats that you may run into:

​
Account for Kubernetes overhead

Because Kubernetes nodes have inherent overhead, pod requests never perfectly match up one-to-one with the node size that a pod is eventually scheduled on. For instance, a request of 4 CPU and 16Gi cannot schedule on a machine that only has 4 CPU and 16Gi available, and will always schedule on a larger machine, often a machine that is twice as large (8 CPU and 32Gi). When migrating a service from custom config to a Chalk Machine Type, you can choose a machine type that corresponds either to the existing request, or to the size of the machine that the workload was actually scheduled on. We recommend picking the size that corresponds to the existing request, but you may find that the workload’s performance implicitly depended on the extra resources from the actual scheduled node. If the workload’s performance degrades because of the smaller node size, we recommend picking the larger size.

​
Translate custom instance type configuration

Each Chalk Machine Type corresponds to a “primary” concrete instance type in the underlying cloud provider, and then falls back to an ordered series of fallback instance types if the primary is out of stock in your region. If you configured your workloads to use specific instance types from the underlying cloud provider (EC2 Instance Type in AWS, Nodepool / Instance Type in GCP), then you may find that switching to a machine type causes the system to choose a different underlying instance type than your custom config, which may mean a different CPU generation or CPU manufacturer. Typically, the system will pick the fastest and most modern available processor type, so you are likely to see a performance improvement rather than a degradation, but if you have a specific instance type configured it may be helpful to check the default instance and fallback catalog to see what underlying instance types your workload will select when migrated.

If you want to change which concrete instance types correspond to your selected machine type in order to balance performance and cost considerations or guarantee a specific instance type, then it’s possible to override the defaults in the organization-level cluster settings page.

​
Remove disk spilling configuration (offline only)

In the legacy custom resource configuration system, special settings were required in order to enable disk spilling for offline workloads like the job queue consumer or async offline query. With Chalk Machine Types, disk spilling is always automatically enabled for offline workloads, and those custom values may conflict and cause spilling to fail or cause the workload to fail to boot. Remove these custom values from your offline workload configuration before migrating to a machine type:

  • Engine Configuration Variables CHALK_VELOX_SPILL_DIRECTORY and CHALK_VELOX_QUERY_SPILLING_MODE
  • Custom Kubernetes Volume Mounts that point to a spill directory

​
List of Available Chalk Machine Types

Machine types larger than xlarge are only available for offline workloads. “High-CPU” machine types are only available for online workloads. “High-memory” and “high-cpu” indicate a larger and smaller ratio of memory to vCPU respectively on the instances compared to the “standard” machine type sizes, approximately:

  • High-CPU: 2GiB per vCPU
  • Standard: 4GiB per vCPU
  • High-memory: 8GiB per vCPU

This is the full list of machine type sizes:

Standard

AWS Machine typevCPUMemory
small416Gi
medium832Gi
large1664Gi
xlarge32128Gi
2xlarge48192Gi
3xlarge64256Gi
4xlarge96384Gi
5xlarge128512Gi

High-CPU

AWS Machine typevCPUMemory
small-highcpu48Gi
medium-highcpu816Gi
large-highcpu1632Gi
xlarge-highcpu3264Gi

High-memory

AWS Machine typevCPUMemory
small-highmem432Gi
medium-highmem864Gi
large-highmem16128Gi
xlarge-highmem32256Gi
2xlarge-highmem48384Gi
3xlarge-highmem64512Gi
4xlarge-highmem96768Gi
5xlarge-highmem1281024Gi

Micro

AWS Machine typevCPUMemory
micro24Gi

The “micro” machine type is not suitable for production workloads and is only recommended for scenarios where a minimal deployment is required for testing.