Monitor Kafka Streams Applications in Confluent Cloud

Confluent Cloud provides tools to monitor and manage your Kafka Streams applications. Access the Kafka Streams monitoring page by navigating to your cluster’s overview page in Confluent Cloud Console and clicking Kafka Streams.

Note

The Kafka Streams monitoring features in Confluent Cloud Console require your applications to be built with Kafka Streams and Apache Kafka® client libraries version 4.0.0 or later. For more information, see Kafka Streams Upgrade Guide.

For this guide, you create a Kafka Streams application by using Confluent for VS Code, or you can run an existing Kafka Streams application that uses Kafka topics in Confluent Cloud.

If you’re using an existing Kafka Streams application, you can skip to Step 7: Monitor the application. Ensure that the application is built with the latest version of the Kafka Streams and Kafka client libraries. To use all of the monitoring features, Kafka version 4.0.0 or later is required.

For a complete list of metrics available for Kafka Streams applications, see Kafka Streams Metrics.

Prerequisites

  • Confluent for VS Code.

  • Docker installed and running in your development environment.

  • A Kafka cluster running in Confluent Cloud.

    • Kafka bootstrap server host:port, for example, pkc-abc123.<cloud-provider-region>.<cloud-provider-region>.confluent.cloud:9092, which you can get from the Cluster Settings page in Cloud Console. For more information, see How do I view cluster details with Cloud Console?.

    • Kafka cluster API key and secret, which you can get from the Cluster Overview > API Keys page in Cloud Console.

Step 1: Create the Kafka Streams project

Create the Kafka Streams project by using the Kafka Streams Application template and filling in a form with the required parameters.

Open the template in VS Code directly

To go directly to the Kafka Streams Application template in VS Code, click this button:

Open template in VS Code

The Kafka Streams Application form opens.

Skip the manual steps and proceed to Step 2: Fill in the template form.

Open the template in VS Code manually

Follow these steps to open the Kafka Streams Application template manually.

  1. Open VS Code.

  2. In the Activity Bar, click the Confluent icon. If you have many extensions installed, you might need to click … to access Additional Views and select Confluent from the context menu.

  3. In the extension’s Side Bar, locate the Support section and click Generate Project from Template.

    The palette opens and shows a list of available project templates.

  4. Click Kafka Streams Application.

    The Kafka Streams Application template opens.

Step 2: Fill in the template form

The project needs a few parameters to connect with your Kafka cluster.

  1. In the Kafka Streams Application form, provide the following values.

    • Kafka Bootstrap Server: Enter the host:port string from the Cluster Settings page in Cloud Console. If you’re logged in with Confluent for VS Code, you can right-click on the Kafka cluster in the Resources pane and select Copy bootstrap server.

    • Kafka Cluster API Key: Enter the Kafka cluster API key.

    • Kafka Cluster API Secret: Enter the Kafka cluster API secret.

    • Input Topic: The name of a topic that the Kafka Streams application consumes messages from. Enter input_topic. You create this topic in a later step.

    • Output Topic: The name of a topic that the Kafka Streams application produces messages to. Enter output_topic. You create this topic in a later step.

  2. Click Generate & Save, and in the save dialog, navigate to the directory in your development environment where you want to save the project files and click Save to directory.

    Confluent for VS Code generates the following project files.

    • The Kafka Streams code, in the src/main/java/examples directory, in a file named KafkaStreamsApplication.java.

    • A docker-compose.yml file that declares how to build the Kafka Streams code.

    • A config.properties file that holds configuration settings, like bootstrap.servers.

    • A .env file that holds secrets, like the Kafka cluster API key.

    • A README.md file with instructions for compiling and running the project.

  3. Open the generated build.gradle file and update the version of the Kafka Streams and Kafka client libraries to the latest version. Kafka version 4.0.0 or later is required to use the monitoring features.

    dependencies {
      implementation 'org.apache.kafka:kafka-streams:4.0.0'
      implementation 'org.apache.kafka:kafka-clients:4.0.0'
    }
    

Step 3: Connect to Confluent Cloud

  1. In the extension’s side bar, click Sign in to Confluent Cloud.

  2. In the dialog that appears, click Allow.

    A browser window opens to the Confluent Cloud login page.

  3. Enter your Confluent Cloud credentials, and click Log in.

    After you authenticate, VS Code displays your Confluent Cloud resources in the extension’s Side Bar.

Step 4: Create topics

Confluent for VS Code enables creating Kafka topics within VS Code.

  1. In the extension’s Side Bar, open Local in the Resources section and click cluster-local.

    The Topics section refreshes, and the cluster’s topics are listed.

  2. In the Topics section, click + to create a new topic.

    The palette opens with a text box for entering the topic name.

  3. In the palette, enter input_topic. Press ENTER to confirm the default settings for the partition count and replication factor properties.

    The new topic appears in the Topics section.

  4. Repeat the previous steps for another new topic named output_topic.

Step 5: Compile and run the project

Your Kafka Streams project is ready to build and run in a Docker container.

  1. In your terminal, navigate to the directory where you saved the project.

  2. The Confluent for VS Code extension saves the project files in a subdirectory named kafka-streams-simple-example. Run the following command to navigate to this directory.

    cd kafka-streams-simple-example
    
  3. Run the following command to build and run the Kafka Streams application.

    docker compose up --build
    

    Docker downloads the required images and starts a container that compiles the project.

Step 6: Produce messages to the input topic

The Kafka Streams application you created in the previous step consumes messages from input_topic and produces messages to an output_topic. For convenience, this guide uses a Datagen Source connector to produce messages to input_topic.

  1. In your browser, log in to Confluent Cloud Console and navigate to your Kafka cluster.

  2. In the navigation menu, click Connectors.

  3. Click Add connector.

  4. On the Search box, type datagen.

  5. Select the Datagen Source connector, and in the Launch Sample Data dialog, click Additional configuration.

  6. In the Choose the topics you want to send data to section, select input_topic, and click Continue.

  7. In the API key section, click Generate API key & download and click Continue.

  8. In the Configuration page, select JSON for the output record value format, and Orders for the schema, and click Continue.

  9. For Connector sizing, leave the slider at the default of 1 task and click Continue.

  10. Name the connector Kafka_Streams_data_source and click Launch connector.

    Confluent Cloud provisions the connector. After a short time, the connector starts producing messages to input_topic.

Step 7: Monitor the application

After your Kafka Streams application is running, you can monitor it by using the Confluent Cloud Console.

  1. In Cloud Console, navigate to your Kafka cluster’s overview page, and in the navigation menu, click Kafka Streams.

    The Kafka Streams page displays a list of all the Kafka Streams applications running in your Kafka cluster, along with metrics aggregated across all the Kafka Streams applications in your Kafka cluster.

    The page shows the following metrics:

    • Application Name: The name of the Kafka Streams application.

    • Client version: The version of the client used by the application.

    • Status: The status of the application. For more information, see Application status.

    • Running threads: The number of threads running in the application.

    • Total production: The total number of messages produced by the application in the last minute.

    • Total consumption: The total number of messages consumed by the application in the last minute.

    • Total lag: The total lag of the application.

  2. In the list, find the application you want to monitor. You can search for the application by name. If you created it with Confluent for VS Code, the application name starts with vscode-kafka-streams-simple-example-.

  3. Click the link in the Application column to open the application’s overview page.

    The overview page displays the following metrics:

    • Running threads: The number of stream threads running in the application.

    • Total production (last min): The total number of messages produced by the application in the last minute.

    • Total consumption (last min): The total number of messages consumed by the application in the last minute.

    • Size of memtables: The approximate size of the active, unflushed immutable, and pinned immutable memtables.

    • Estimated number of keys: The estimated number of keys in the active and unflushed immutable memtables and storage.

    • Block cache usage: The memory size of the entries residing in the block cache.

    The overview page also displays the following graphs:

    • End-to-end latency: The maximum, average, and minimum end-to-end latency of a record, measured by comparing the record timestamp with the system time when the record has been fully processed.

    • CPU ratio: The percentage of time the application’s stream threads spend on Commit, Poll, Process, and Punctuate operations. A high Process ratio indicates a performance issue in your application code. A high Commit ratio indicates an issue with the cluster or network.

    Below the graphs, a table lists each stream thread running in the application:

    • Thread: The name of the stream thread.

    • Process ID: The identifier of the client process that the thread belongs to. Use this ID to correlate the thread with logs or metrics in your own observability tools.

    • Status: The current state of the thread, for example, Running.

    Confluent Cloud takes a few minutes to collect and display the metrics.

Application status

Cloud Console derives the application Status shown on the Kafka Streams page and the application overview page from the state of the application’s threads:

  • Running: All threads are actively processing records.

  • Rebalancing: The application is reassigning tasks among its threads, for example, because an instance started or stopped.

  • Restoring: One or more threads are restoring state from a changelog topic, for example, after a new instance joins the application or after local state is lost.

Rebalancing and restoring are normal, typically short-lived phases of a Kafka Streams application’s lifecycle. If an application shows a Rebalancing or Restoring status for an extended period, check the application logs for errors.

View thread and task details

To see per-thread detail for a running application, click a thread name in the threads table on the application overview page. The thread detail panel shows the Client ID for the thread and a Tasks table listing every task assigned to the thread, with the following columns:

  • Task ID: The task identifier, in the form <subtopology-id>_<partition>, for example, 0_5.

  • Task type: The task’s assignment type: Active, Standby, or Warmup.

  • Topic: The topic associated with the task. Click the topic name to go to the topic’s detail page. Confluent Cloud prefers the task’s input topic. If the task has no input topic, Confluent Cloud shows its repartition topic. If it has neither, Confluent Cloud shows its changelog topic.

  • Partition: The partition, within the topic, that the task processes.

  • Lag: The consumer lag for the task’s topic-partition.

Note

For tasks associated with an internal repartition topic, the Topic and Lag columns might be blank.

Sort or scan the Lag column to find tasks with disproportionately high lag compared to other tasks in the same application. A task with high lag relative to its peers points to a hot partition. Use the task’s Topic and Partition values to investigate further, for example, by checking whether your keying strategy sends a disproportionate share of traffic to that partition.

Get deeper diagnostics with the Kafka Streams Add-On

The Kafka Streams Add-On unlocks deeper diagnostics for applications that use the streams group protocol (group.protocol=streams), introduced by KIP-1071. With both in place, the application overview page and thread detail panel show more diagnostics while an application is rebalancing or restoring, including:

  • The progress of task reassignment for each thread during a rebalance, comparing the thread’s current task assignment against its target assignment.

  • State restoration progress during a restore, including the proportion of each task’s state that has been restored from its changelog topic.

If your organization doesn’t have the Kafka Streams Add-On, or if your application uses the classic consumer group protocol, Cloud Console shows a banner on the application overview page. The banner prompts you to upgrade to the streams group protocol and get the Kafka Streams Add-On. Without the Add-On, the application list, overview page, and thread detail panel described earlier in this topic are still available.

For upgrade steps, see Enable the streams group protocol. To request the Kafka Streams Add-On, click Get Add-On on the application overview page, or contact your Confluent account team.

Performance considerations

Network traffic and metadata fetching

Kafka Streams applications have different metadata requirements than standard Kafka producers and consumers. Understanding these differences is important for monitoring network usage and application performance.

Metadata fetching behavior

Kafka Streams applications fetch more comprehensive metadata than plain Kafka clients. Specifically, Kafka Streams applications pull metadata for all topics in the cluster, not only the topics they directly produce to or consume from. The Kafka Streams framework requires this behavior for its internal operations.

Impact on service accounts with DataDiscovery or DataSteward roles

When you use Kafka Streams applications with service accounts that have the DataDiscovery or DataSteward role, you might see increased network traffic compared to applications that use service accounts with more restrictive permissions.

This increase occurs because:

  • The DataDiscovery and DataSteward roles provide read access to topic metadata across all topics in an environment.

  • Kafka Streams applications fetch comprehensive topic metadata as part of their normal operation.

  • The combination results in additional network requests for metadata retrieval.

Monitoring recommendations

When you monitor Kafka Streams applications that use service accounts with the DataDiscovery or DataSteward role, follow these recommendations:

  • Monitor network utilization metrics to understand the baseline metadata traffic.

  • Consider this additional metadata fetching when planning network capacity.

  • Use the Metrics API to track request_count and other network-related metrics for your applications.

  • Evaluate whether the DataDiscovery or DataSteward role is necessary for your specific Kafka Streams application use case.

For more information about role-based access control and service account configuration, see Role-based Access Control (RBAC) on Confluent Cloud and Service Accounts on Confluent Cloud.