Tableflow Quick Start with Iceberg Tables Using Your Storage on Google Cloud in Confluent Cloud

Confluent Tableflow exposes Apache Kafka® topics as Apache Iceberg™ tables. Iceberg is an open table format for large analytic datasets in object storage.

Complete these steps to materialize a Kafka topic as an Iceberg table using your own Google Cloud Storage bucket and a catalog integration:

Prerequisites

Before you begin, make sure that you meet the following requirements.

Your account needs the following role-based access control (RBAC) roles in Confluent Cloud:

  • The DeveloperRead role on all schema subjects.

  • The CloudClusterAdmin role on your Kafka cluster.

  • The Assigner role on all provider integrations.

  • The EnvironmentAdmin role on your environment, or the OrganizationAdmin role, so that you can create the provider integration in Step 2.

For more information, see Grant Role-Based Access for Tableflow in Confluent Cloud.

You also need the following resources:

  • A Confluent Cloud cluster on Google Cloud, in a region where Tableflow is available. For supported regions, see Google Cloud region availability.

  • A Google Cloud project where you have permission to create a Google Cloud Storage bucket, an Identity and Access Management (IAM) role, a service account, and IAM role bindings.

Step 1: Create a topic and publish data

In this step, you create a stock-trades topic and publish sample stock trade data to it with the Datagen Source connector.

Go to the Connectors page

Follow these steps to open the Connectors page for your Google Cloud cluster:

  1. Sign in to the Confluent Cloud Console at https://confluent.cloud/.

  2. In the navigation menu, click Environments.

  3. Click the environment that contains your Google Cloud cluster.

  4. Click Clusters.

  5. Click your cluster.

  6. In the cluster navigation menu, click Connectors.

Create the topic and launch the Datagen Source connector

Follow these steps to create the stock-trades topic and launch a Datagen Source connector that publishes data to it:

  1. On the Connectors page, click Add Connector.

  2. On the Connector Plugins page, click Sample Data.

    The Launch Sample Data dialog opens.

  3. Click Additional configuration.

    Don’t click Launch. It creates a topic with a different name and data format.

  4. On the Topic selection step, click Add new topic.

  5. For Topic name, enter stock-trades.

  6. Click Create with defaults.

  7. Click Continue.

  8. On the Kafka access step, keep My account selected.

    For production workloads, use a service account instead.

  9. Click Generate API key and download.

    The wizard creates an API key for your user account and downloads the key and secret as a text file.

  10. Click Continue.

  11. On the Configuration step, for Select output record value format, select AVRO.

  12. Under Select a schema, click Stock trades.

  13. Click Continue.

  14. On the Sizing step, click Continue.

  15. On the Review and launch step, click Continue.

When the connector’s status on the Connectors page changes to Running, the connector is publishing data to the stock-trades topic. To view the data, open the stock-trades topic and click Messages.

For more information, see Datagen Source Connector Quick Start.

Step 2: Configure your Google Cloud Storage bucket and provider integration

Configure the storage bucket that holds the table data before you materialize your Kafka topic as an Iceberg table.

Create a Confluent Cloud provider integration so that Tableflow can access your Google Cloud Storage bucket and write materialized data into it. The provider integration lets Tableflow impersonate a Google Cloud service account that you own.

Complete the following steps, switching between the Google Cloud console and Confluent Cloud Console:

  1. In the Google Cloud console, create the bucket, IAM role, and service account that Tableflow uses by doing the following:

    1. Create an empty bucket. For Location type, select Region, and then select the region of your Kafka cluster. You can also use a dual-region bucket that includes your cluster’s region. Tableflow doesn’t support multi-region buckets. Bucket names must be unique across Google Cloud, so choose a unique name, for example, tableflow-quickstart-<unique_suffix>. Leave uniform bucket-level access enabled. For more information, see Create a bucket and Bucket locations in the Google Cloud documentation.

    2. Create a custom IAM role that includes the following IAM permissions:

      • storage.objects.get

      • storage.objects.list

      • storage.objects.create

      • storage.objects.delete

      • storage.folders.create

      You can also use the predefined Storage Object User role (roles/storage.objectUser), which includes these permissions and others. For more information, see Create a custom role in the Google Cloud documentation.

    3. Create a service account for Tableflow to impersonate, for example, tableflow-quickstart-sa. For more information, see Create service accounts in the Google Cloud documentation.

    4. On the bucket that you created, grant the role to your service account.

  2. In Confluent Cloud Console, start creating the provider integration and get the Confluent service account by doing the following:

    1. On the Environments page, click the environment that contains your Google Cloud cluster.

    2. In the environment navigation menu, click Integrations.

    3. On the Provider integrations tab, click Add integration.

    4. Click Google service account.

    5. For Provider integration name, enter a name for the integration, for example, tableflow-quickstart-integration.

    6. Click Continue.

    7. Under Step 1, click Create service account.

      Confluent Cloud creates a service account in the Confluent Google Cloud project.

    8. Copy the email of the Confluent service account.

  3. In the Google Cloud console, grant the Confluent service account the Service Account Token Creator role (roles/iam.serviceAccountTokenCreator) on your service account.

    This role lets Confluent Cloud impersonate your service account.

  4. In Confluent Cloud Console, connect the provider integration to your service account by doing the following:

    1. Click Continue.

    2. For Google Cloud service account, enter the email of your service account.

    3. Click Validate.

      If validation fails right after you grant the role, wait a few minutes for Google Cloud to apply the change, and then try again.

    4. Click Continue.

Step 3: Enable Tableflow on your topic

You can now enable Tableflow on your Kafka topic to materialize it as an Iceberg table in the storage bucket that you created in Step 2.

  1. In Confluent Cloud Console, on the Environments page, click the environment that contains your Google Cloud cluster, and then click Clusters.

  2. Click your cluster.

  3. In the cluster navigation menu, click Topics.

  4. In the stock-trades row, under Tableflow sync status, click Enable Tableflow.

  5. In the Enable Tableflow dialog, keep Iceberg selected.

  6. Click Configure custom storage.

  7. On the Configure storage step, make sure that Store in your own storage is selected.

  8. For Provider integration, select the provider integration that you created in Step 2.

  9. Click Continue.

  10. On the Configure storage details step, for GCS bucket name, enter the name of the bucket that you created in Step 2.

    This step also shows how to grant your service account access to the bucket. You already granted this access in Step 2.

  11. Click Continue.

  12. On the Review and launch step, click Launch.

Materializing a newly created topic as an Iceberg table can take a few minutes.

Step 4: Configure a catalog integration

Configure a catalog integration so that Tableflow publishes your Iceberg tables to a catalog service, where query engines can discover them. Tableflow bases the namespace name on your Kafka cluster ID.

Follow the steps in Integrate Tableflow with Snowflake Open Catalog to configure Snowflake Open Catalog as a catalog integration.

It can take a few minutes for Tableflow to publish Iceberg tables to your catalog.

Tip

You can monitor each topic’s catalog sync status from the Tableflow page in Cloud Console after you configure the catalog integration. For more information, see Topic catalog sync status.

Step 5: Query Iceberg tables

You can query the stock-trades table from Snowflake. For the steps, see Use Snowflake with Snowflake Open Catalog. When you create the Snowflake external volume, use Configure an external volume for Google Cloud Storage instead of the Amazon S3 instructions.

Step 6: Query data with other analytics engines (optional)

You can also read the tables through the Tableflow Iceberg REST Catalog from other Iceberg-compatible engines: