<a id="cloud-cluster-monitor-performance"></a>

# Dedicated Cluster Performance and Expansion in Confluent Cloud

A number of factors can affect the overall performance of applications running
on Dedicated Apache Kafka® clusters. Monitor your clusters for clues to how
you can improve performance in your applications, and to determine whether
cluster expansion is the right solution.

Key cluster performance metrics:

- [Cluster load metrics](#cloud-cluster-load-expansion)
- [CKU count metric](#cloud-cku-count-metric)
- [Hot-partition metrics](#ccloud-hot-partitions)
- [Consumer lag](#cloud-consumer-lag-performance)
- [Unsupported client versions](#cloud-unsupported-client-versions)
- [Client throttling](#cloud-client-apps-throttled)
- [Producer latency](#cloud-producer-latency-monitoring)

For more about monitoring your applications, see
[Monitoring Your Event Streams: Tutorial for Observability Into Apache Kafka Clients (blog)](https://www.confluent.io/blog/monitoring-event-streams-visualize-kafka-clients-in-confluent-cloud/).

<a id="cloud-cluster-load-expansion"></a>

## Interpret a high cluster load value

Cluster load is a percentage value between 0 and 100, with 0% indicating no
load on the cluster and 100% representing a fully saturated cluster. Expect
higher latency and some degree of throttling if the average cluster load is
greater than 80%.

You can measure cluster load in these ways:

- Average value: The average load across all brokers in the cluster.
- Maximum value: The highest load among all brokers in the cluster.

A high average cluster load, which typically also means a high maximum
cluster load, usually indicates a fully saturated cluster, commonly
resulting in higher latencies, client throttling, or both for your
application.

If the maximum cluster load is high but the average cluster load is much
lower, this is a sign of a skewed workload, where a small number of
partitions handle disproportionately more traffic than the rest of the
cluster. In this case, investigate the [hot-partition metric](#ccloud-hot-partitions) to identify where the skew is occurring.

You could expand your cluster in an attempt to lower the load, but before you do
so you should look at the time-series graph for cluster load to obtain a
historical perspective of the cluster load variation.

![Cluster load graph](images/_monitoring/cluster-load-chart.png)

When viewing this graph consider that an average cluster load of 70% might be
acceptable if it is an occasional spike or a normal high point for a
workload. However, a load of 70% might be too high if the cluster needs
more capacity to accommodate load spikes due to variations in
application workload patterns, or if you add new workloads to the cluster.
In this case, expanding the cluster is probably the right solution.

Generally, expanding your Dedicated cluster provides more capacity for
your workloads, and often helps improve the performance of your Kafka
applications. In addition, a lower cluster load can help improve latency for
your applications.

If you expand a Dedicated cluster, and the expansion does not resolve
the performance issues, you can shrink the cluster back to its original size.
For more information, see
[Resize a Dedicated Kafka Cluster](../clusters/resize.md#cloud-cluster-resize).

To interpret a high average cluster load value, consider the following
guidelines:

- Values of 70 to 80: Unless these are occasional peaks, consider adding CKU.
- Values of 80 or more: Expect throttling and degraded performance if the
  workload pattern changes and introduces additional load on the overall system.

<a id="cloud-cku-count-metric"></a>

## Determine CKU count

The Confluent Unit for Kafka (CKU) count determines the capacity of your cluster. Use the CKU
count metric to monitor the capacity of your Dedicated cluster. While
some performance dimensions for Dedicated clusters are fixed, others
have a recommended guideline that allows you greater use of one dimension at
the expense of another. For more information, see [Fixed limits and recommended guidelines](../clusters/cluster-types.md#cku-details),
[Dimensions with fixed limits](../clusters/cluster-types.md#fixed-limit) and [Dimensions with recommended guidelines](../clusters/cluster-types.md#guideline).

<a id="ccloud-hot-partitions"></a>

## Identify hot partitions

A hot partition is a partition where the maximum cluster load is much higher
than the average cluster load. Use the hot-partition metric together with
the cluster load metrics to identify workload skew. If a hot-partition
metric value of `1` appears on any topic or partition, and there is a large
gap between the maximum and average cluster load, the workload is skewed.
Self-balancing clusters (SBC) might not be able to balance these partitions,
so to avoid throttling or availability issues, you must manually address hot
partitions.

**Considerations**

- If `hot_partition_ingress` or `hot_partition_egress` has a value of `1`,
  then the cluster load for that partition is too high for SBC to balance the
  partition.
- To recover a hot partition, first reassess your keying strategy to ensure
  traffic is distributed more evenly across partitions. If reassessing your
  keying strategy doesn’t sufficiently resolve the skew, consider adding
  partitions and redistributing client traffic. The goal is to have enough
  partitions with traffic distributed evenly across the partitions.
- Ensure unsupported clients are not accessing the hot partition. Unsupported
  clients can cause high cluster load and issues with balancing clusters. For
  more information, see [What client and protocol versions are supported?](../faq.md#cloud-faq-supported-clients) and
  [Confluent Platform and Apache Kafka
  compatibility](/platform/current/installation/versions-interoperability.html#cp-ak-compatibility).
- To use the hot-partition metrics with an application performance monitoring
  (APM) application, your APM must use a default zero function to fill empty
  time intervals with a zero value or an interpolation value, if your APM uses
  interpolation.

<a id="cloud-consumer-lag-performance"></a>

## Review high consumer lag

The [Consumer lag](monitor-lag.md#cloud-monitoring-lag) metric indicates the number of
records for any partition that the consumer is behind in the log. If the rate of
production of data exceeds the rate at which it is getting consumed, consumer
groups measure lag. An increase in consumer lag can indicate a client-side
issue, a Kafka server-side issue, or both.

![Consumer lag detail graph](images/_monitoring/consumer-group-lag-detail.png)

Use the following checks, in order, to identify what is causing consumer lag:

1. Check the [cluster load metric](#cloud-cluster-load-expansion). If
   the cluster is heavily loaded at greater than 70%, expand your cluster.
   This is likely to help with consumer lag.
2. If the cluster is not heavily loaded, check the partition count. If the
   cluster has fewer than the recommended 6 to 10 partitions per CKU
   (partitions parallelize the workload across your cluster), add
   partitions to improve parallelization. See [Optimize and Tune Confluent Cloud Clients](../client-apps/optimizing/overview.md#ccloud-optimizing).
   Your specific workload might warrant a higher or lower number of
   partitions.
3. If the cluster is not heavily loaded and has the recommended number of
   partitions or more, check consumer application parallelism. Add
   consumers to your application. This might resolve the lag by increasing
   parallelism.

<a id="cloud-unsupported-client-versions"></a>

## Monitor for unsupported client versions

Use only
[supported clients](/platform/current/installation/versions-interoperability.html#cp-ak-compatibility)
with Confluent Cloud. Update client versions regularly to avoid undesired behavior in
your clusters. Some issues that can arise from running unsupported client
versions:

- **Compatibility issues**: Confluent Cloud is designed to work with specific
  versions of client libraries such as Java, .NET, and Python. If you use an
  unsupported client, it might not have access to necessary libraries, which
  can lead to unexpected behavior, errors, or failures.
- **Security vulnerabilities**: Confluent Cloud is regularly updated with security
  patches and bug fixes. Unsupported clients might be missing critical
  security updates, leaving your applications and data vulnerable to
  potential security threats or exploits.
- **Lack of support**: Confluent provides support and maintenance only for
  supported client versions. If you encounter issues or bugs while using an
  unsupported client version, Confluent might not be able to provide
  help or troubleshooting, leaving you to resolve the issues on your
  own.
- **Missing features and improvements**: Newer versions of client libraries
  often introduce new features, performance improvements, and bug fixes. By
  using an unsupported client version, you might miss out on these
  enhancements, which could impact the capabilities, performance, and
  reliability of your applications.
- **Potential service disruptions**: Confluent Cloud might introduce changes or updates
  that are designed to work only with supported client versions. Using an
  unsupported client version could lead to service disruptions or unexpected
  behavior when such changes are made.

To ensure a stable, secure, and well-supported experience with Confluent Cloud, use
only supported client versions for your specific programming language or
framework. Confluent provides documentation and guidance on the supported
client versions and their compatibility with Confluent Cloud.

For more information, see
[Supported Versions and Interoperability for Confluent Platform](/platform/current/installation/versions-interoperability.html).

<a id="cloud-client-apps-throttled"></a>

## Review client application throttles

Confluent Cloud clusters throttle client applications if they exceed the rate the
cluster is configured to handle based on its allocated capacity. Throttles
are a normal part of working with cloud services. This throttling prevents
excess usage that could cause a cluster outage, which might be
catastrophic. Throttles are negotiated between the Kafka server and Kafka
consumers and producers, ensuring the clients wait long enough for the
server to handle the request without compromising uptime.

Beyond client-side metrics, Confluent Cloud provides a managed throttling
metric through the Metrics API. The
`io.confluent.kafka.server/client_limit_milliseconds` metric reports the
average throttle time applied to a principal, which is a user or service
account, when a quota is violated. It requires no client-side instrumentation.
For more information, see [Throttled clients metric](metrics-api.md#throttled-clients-metric).

For more information about client-side, producer, and consumer metrics, which
provide visibility into whether your producers, consumers, or both are being
throttled, see [Client Monitoring](../client-apps/monitoring.md#ccloud-monitoring) and the discussion
of `produce-throttle-time-avg` and `produce-throttle-time-max` in the
[Producer Metrics](https://docs.confluent.io/platform/current/kafka/monitoring.html#producer-metrics)
section of the Confluent Platform documentation.

### Determine if throttling is caused by server or client issues

Use the following checks, in order, to determine whether throttling is caused
by server-side or client-side issues.

1. Check the [Throttled clients metric](metrics-api.md#throttled-clients-metric), which includes a
   `metric.reason` field that identifies the underlying cause of a
   throttle, such as a cluster-level quota violation, a principal-level
   quota violation, or skewed traffic. Use this field as a starting point
   to determine whether the throttling is a server-side or client-side
   issue.
2. Check the [cluster load metric](#cloud-cluster-load-expansion). If
   it indicates a high or increasing load, expanding the cluster likely
   mitigates the throttling. For more information, see
   [Resize a Dedicated Kafka Cluster](../clusters/resize.md#cloud-cluster-resize).
3. Evaluate ingress, egress, and request rate. The throughput and request
   rate for a Dedicated cluster are limited by the number of CKUs
   allocated to it. If your applications are consuming more throughput or
   making more requests than the cluster can currently handle, you could
   expand your cluster to likely resolve the issue.

If you cannot identify a server-side cause, the throttling might be due to
a client-side access pattern. For example, your workload might be
unbalanced, with one or a few partitions sustaining most of the traffic
while the rest of the cluster is underutilized. Confluent Cloud constantly
monitors cluster balance to optimize workload distribution automatically,
but it might not find an optimal balance if client-side access patterns,
such as unbalanced partition assignment strategies, are the cause.

To get the best throughput from Kafka, architect your client-side
applications to use balanced access patterns.

<a id="cloud-producer-latency-monitoring"></a>

## Monitor latency in producer applications

For the latest Kafka clients, Confluent Cloud provides a native metric to help
monitor producer latency without extra client-side instrumentation.
The `io.confluent.kafka.server/producer_latency_avg_milliseconds` metric
is available only for clients using Java 3.8 or later with the
`enable.metrics.push` client configuration set to `true` (the default).
For more information, see [KIP-714 client metrics](metrics-api.md#kip-714-client-metrics). KIP stands for Kafka Improvement Proposal.

Certain metrics can also indicate that the producers are experiencing latency.
Specifically, the `buffer-available-bytes(=0)`, or increases in
`bufferpool-wait-time` could be an indication of producers experiencing
latency.

Monitor `active_connection_count`. Benchmarking shows that exceeding the
number of total client connections per-CKU often leads to an exponential
increase in produce latency. For more information, see the total client
connections dimension in [Dimensions with recommended guidelines](../clusters/cluster-types.md#guideline) table.

If the [cluster load metric](#cloud-cluster-load-expansion) is high and the
[producer buffer](https://kafka.apache.org/documentation/#producer_monitoring)
is high, it is likely that expanding the cluster by adding CKUs improves
producer latency. If latency is not improved after cluster expansion, see
[Client application throttles](#cloud-client-apps-throttled).

## Related content

- [Consumer Lag](monitor-lag.md#cloud-monitoring-lag)
- [Resize a Dedicated Kafka Cluster](../clusters/resize.md#cloud-cluster-resize)
- [Monitor Clients](../client-apps/monitoring.md#ccloud-monitoring)
- [Monitoring Your Event Streams: Tutorial for Observability Into Apache Kafka Clients (blog)](https://www.confluent.io/blog/monitoring-event-streams-visualize-kafka-clients-in-confluent-cloud/)
