<a id="cc-s3-connect-source"></a>

# Amazon S3 Source Connector for Confluent Cloud

The fully managed Amazon S3 Source connector for Confluent Cloud reads data from files
in an S3 bucket. The file names don’t have to be in a specific format. The file
format has to be supported (for example, Avro, Bytes, CSV, JSON, or Parquet)
for the connector to read from.

#### NOTE
* This Quick Start is for the fully managed Confluent Cloud connector. If you are
  installing the connector locally for Confluent Platform, see [Generalized Amazon S3
  Source Connector for Confluent Platform](https://docs.confluent.io/kafka-connectors/s3-source/current/generalized/overview.html).
* If you require private networking for fully managed connectors, make sure to set up the proper
  networking beforehand. For more information, see [Manage Networking for Confluent Cloud Connectors](networking/internet-resource.md#clusters-connect-cloud).

## Features

The Amazon S3 Source connector provides the following features:

* **At least once delivery**: The connector guarantees that records are delivered at least once.
* **Supports multiple tasks**: The connector supports running one or more tasks.
* **Client-side encryption (CSFLE and CSPE) support**: The connector supports CSFLE and CSPE for sensitive data.
  For more information about CSFLE or CSPE setup, see the [connector configuration](#cc-s3-source-setup-connection).
* **Offset management capabilities**: The connector supports offset management. For more information, see [Manage custom offsets](#cc-s3-source-custom-offsets).
* **Supported input data formats**: The connector supports Avro, Bytes, CSV, JSON, and Parquet input formats. The supported compression types for Parquet formats are `snappy`, `gzip`, and `none`. Note that the connector can support Parquet input files up to 2GB in size.
* **Provider integration support**: The connector supports IAM role-based authorization using Confluent Provider Integration. For more information about provider integration setup, see the [IAM roles authentication](#cc-s3-source-setup-connection).

For more information and examples to use with the Confluent Cloud API for Connect,
see the [Confluent Cloud API for Connect Usage Examples](connect-api-section.md#ccloud-connect-api) section.

Refer to Confluent Cloud [connector limitations](limits.md#s3-source-limits) for additional information.

## IAM Policy for S3

The AWS user account accessing the S3 bucket must have the following permissions:

* ListBucket
* GetObject
* ListAllMyBuckets

#### NOTE
- This is the IAM policy for the user account and not a bucket policy.
- If you’re running the connector on a Google Cloud or Azure cluster, you need the
  `s3:ListAllMyBuckets` permission. It’s optional if you’re running the
  connector on an AWS cluster.
- The connector uses `GetBucketAcl` to verify bucket existence, but this permission is optional. To avoid AccessDenied errors in your CloudTrail logs, you can add the equivalent permission to the connector.

For more information, see [Create and attach a policy to an IAM user](https://docs.aws.amazon.com/apigateway/latest/developerguide/api-gateway-create-and-attach-iam-policy.html).

<a id="cc-s3-source-custom-offsets"></a>

## Manage custom offsets

You can manage the offsets for this connector. Offsets provide information on the
point in the system from which the connector is accessing data. For more
information, see [Manage Offsets for Fully Managed Connectors in Confluent Cloud](offsets.md#connect-custom-offsets).

**To manage offsets**:

- Manage offsets using Confluent Cloud APIs. For more information, see [Connect offsets API reference](https://docs.confluent.io/cloud/current/ccloud/offsets-connect-v-1/).

### Get the current offset

To get the current offset, make a `GET` request that specifies the environment, Kafka cluster, and connector name.

```bash
GET /connect/v1/environments/{environment_id}/clusters/{kafka_cluster_id}/connectors/{connector_name}/offsets
Host: https://api.confluent.cloud
```

**Response:**

Successful calls return HTTP `200` with a JSON payload that describes the offset.

```bash
{
    "id": "lcc-example123",
    "name": "{connector_name}",
    "offsets": [
        {
          "partition": {
            "taskId": "lcc-example123-0-in_progress"
          },
          "offset": {
            "earliestIncomplete": "2023-08-03T10:24:25Z",
            "completedFiles": "[{\"filePath\":\"topics/abc_0/partition=0/abc_0+0+00000.json\",\"creationTime\":\"2023-08-03T10:24:25Z\"},{\"filePath\":\"topics/abc_1/partition=0/abc_1+0+00000.json\",\"creationTime\":\"2023-08-03T10:34:56Z\"},{\"filePath\":\"topics/abc_3/partition=0/abc_3+0+00000.json\",\"creationTime\":\"2023-08-03T10:48:28Z\"}]",
            "recordNum": "98"
          }
        }
        {
          "partition": {
            "taskId": "lcc-example123-1"
          },
          "offset": {
            "earliestIncomplete": "2023-08-03T10:24:25Z",
            "completedFiles": "[{\"filePath\":\"topics/babc_4/partition=0/babc_4+0+00000.json\",\"creationTime\":\"2023-08-03T10:33:04Z\"},{\"filePath\":\"topics/abc_2/partition=0/abc_2+0+00000.json\",\"creationTime\":\"2023-08-03T10:46:06Z\"},{\"filePath\":\"topics/weird/partition=0/weird+0+00000 copy.json\",\"creationTime\":\"2023-08-03T10:51:09Z\"}]",
            "recordNum": "99"
          }
        }
        {
          "partition": {
            "taskId": "lcc-example123-0"
          },
          "offset": {
            "earliestIncomplete": "2023-08-03T10:24:25Z",
            "completedFiles": "[{\"filePath\":\"topics/abc_0/partition=0/abc_0+0+00000.json\",\"creationTime\":\"2023-08-03T10:24:25Z\"},{\"filePath\":\"topics/abc_1/partition=0/abc_1+0+00000.json\",\"creationTime\":\"2023-08-03T10:34:56Z\"},{\"filePath\":\"topics/abc_3/partition=0/abc_3+0+00000.json\",\"creationTime\":\"2023-08-03T10:48:28Z\"},{\"filePath\":\"topics/abc_5/partition=0/abc_5+0+00000.json\",\"creationTime\":\"2023-08-03T10:59:06Z\"}]",
            "recordNum": "99"
          }
        }
        {
          "partition": {
            "taskId": "lcc-example123-1-in_progress"
          },
          "offset": {
            "earliestIncomplete": "2023-08-03T10:24:25Z",
            "completedFiles": "[{\"filePath\":\"topics/babc_4/partition=0/babc_4+0+00000.json\",\"creationTime\":\"2023-08-03T10:33:04Z\"},{\"filePath\":\"topics/abc_2/partition=0/abc_2+0+00000.json\",\"creationTime\":\"2023-08-03T10:46:06Z\"}]",
            "recordNum": "98"
          }
        }
    ],
    "metadata": {
        "observed_at": "2024-03-28T17:57:48.139635200Z"
    }
}
```

Responses include the following information:

- The position of latest offset.
- The observed time of the offset in the metadata portion of the payload. The `observed_at` time
  indicates a snapshot in time for when the API retrieved the offset. A running connector is always updating
  its offsets. Use `observed_at` to get a sense for the gap between real time and the time at which the request
  was made. By default, offsets are observed every minute. Calling `GET` repeatedly will fetch more recently
  observed offsets.
- Information about the connector.

### Update the offset

You can approach offset updates in two ways:

- Modify the `earliestIncomplete` time to reset the offsets so that next scan will source the files with `creationTime` equal to or
  after the new `earliestIncomplete`.

  If you use this approach, consider this:
  - If `earliestIncomplete` is set to a later time,  the connector starts sourcing the files with `creationTime` equal
    to or after the `earliestIncomplete` and skips records.
  - If `earliestIncomplete` is set to an earlier time, the connector might produce duplicate records because it starts
    sourcing every record from files with a `creationTime` equal to or after the earlier time.
- If you want to skip processing a file or files, add the files to `completedFiles`.

To update the offset, make a `POST` request that specifies the environment, Kafka cluster, and connector
name. Include a JSON payload that specifies new offset and a patch type.

```bash
POST /connect/v1/environments/{environment_id}/clusters/{kafka_cluster_id}/connectors/{connector_name}/offsets/request
Host: https://api.confluent.cloud

 {
     "type": "PATCH",
     "offsets": [
         {
             "partition": {
                 "taskId": "lcc-devc3m1zkj-0"
             },
             "offset": {
                 "completedFiles": "[{\"filePath\":\"source/file_0\",\"creationTime\":\"2024-03-06T17:30:28.391Z\"},{\"filePath\":\"source/file_7\",\"creationTime\":\"2024-03-06T17:30:28.395Z\"},{\"filePath\":\"source/file_9\",\"creationTime\":\"2024-03-06T17:30:28.409Z\"},{\"filePath\":\"source/file_1\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_8\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_6\",\"creationTime\":\"2024-03-06T17:30:28.715Z\"},{\"filePath\":\"source/file_30\",\"creationTime\":\"2024-03-06T17:30:28.969Z\"},{\"filePath\":\"source/file_39\",\"creationTime\":\"2024-03-06T17:30:28.970Z\"},{\"filePath\":\"source/file_37\",\"creationTime\":\"2024-03-06T17:30:28.993Z\"},{\"filePath\":\"source/file_36\",\"creationTime\":\"2024-03-06T17:30:29.265Z\"},{\"filePath\":\"source/file_31\",\"creationTime\":\"2024-03-06T17:30:29.268Z\"},{\"filePath\":\"source/file_38\",\"creationTime\":\"2024-03-06T17:30:29.278Z\"},{\"filePath\":\"source/file_25\",\"creationTime\":\"2024-03-06T17:30:29.549Z\"},{\"filePath\":\"source/file_22\",\"creationTime\":\"2024-03-06T17:30:29.551Z\"},{\"filePath\":\"source/file_13\",\"creationTime\":\"2024-03-06T17:30:29.552Z\"},{\"filePath\":\"source/file_47\",\"creationTime\":\"2024-03-06T17:30:30.015Z\"},{\"filePath\":\"source/file_14\",\"creationTime\":\"2024-03-06T17:30:30.020Z\"},{\"filePath\":\"source/file_40\",\"creationTime\":\"2024-03-06T17:30:30.028Z\"},{\"filePath\":\"source/file_15\",\"creationTime\":\"2024-03-06T17:30:30.305Z\"}]",
                 "earliestIncomplete": "2024-03-06T17:30:28.391Z",
                 "recordNum": "0"
             }
         }
     ]
 }
```

Considerations:

- You can only make one offset change at a time for a given connector.
- This is an asynchronous request. To check the status of this request, you must use the check offset status API. For more information,
  see **Get the status of an offset request**.
- For source connectors, the connector attempts to read from the position defined by the requested offsets.

**Response:**

Successful calls return HTTP `202 Accepted` with a JSON payload that describes the offset.

```bash
{
    "id": "lcc-example123",
    "name": "{connector_name}",
    "offsets": [
        {
            "partition": {
                "taskId": "lcc-example123-0"
            },
            "offset": {
                "completedFiles": "[{\"filePath\":\"source/file_0\",\"creationTime\":\"2024-03-06T17:30:28.391Z\"},{\"filePath\":\"source/file_7\",\"creationTime\":\"2024-03-06T17:30:28.395Z\"},{\"filePath\":\"source/file_9\",\"creationTime\":\"2024-03-06T17:30:28.409Z\"},{\"filePath\":\"source/file_1\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_8\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_6\",\"creationTime\":\"2024-03-06T17:30:28.715Z\"},{\"filePath\":\"source/file_30\",\"creationTime\":\"2024-03-06T17:30:28.969Z\"},{\"filePath\":\"source/file_39\",\"creationTime\":\"2024-03-06T17:30:28.970Z\"},{\"filePath\":\"source/file_37\",\"creationTime\":\"2024-03-06T17:30:28.993Z\"},{\"filePath\":\"source/file_36\",\"creationTime\":\"2024-03-06T17:30:29.265Z\"},{\"filePath\":\"source/file_31\",\"creationTime\":\"2024-03-06T17:30:29.268Z\"},{\"filePath\":\"source/file_38\",\"creationTime\":\"2024-03-06T17:30:29.278Z\"},{\"filePath\":\"source/file_25\",\"creationTime\":\"2024-03-06T17:30:29.549Z\"},{\"filePath\":\"source/file_22\",\"creationTime\":\"2024-03-06T17:30:29.551Z\"},{\"filePath\":\"source/file_13\",\"creationTime\":\"2024-03-06T17:30:29.552Z\"},{\"filePath\":\"source/file_47\",\"creationTime\":\"2024-03-06T17:30:30.015Z\"},{\"filePath\":\"source/file_14\",\"creationTime\":\"2024-03-06T17:30:30.020Z\"},{\"filePath\":\"source/file_40\",\"creationTime\":\"2024-03-06T17:30:30.028Z\"},{\"filePath\":\"source/file_15\",\"creationTime\":\"2024-03-06T17:30:30.305Z\"}]",
                "earliestIncomplete": "2024-03-06T17:30:28.391Z",
                "recordNum": "0"
            }
        }
    ],
    "requested_at": "2024-03-28T17:58:45.606796307Z",
    "type": "PATCH"
}
```

Responses include the following information:

- The requested position of the offsets in the source.
- The time of the request to update the offset.
- Information about the connector.

### Delete the offset

To delete the offset, make a `POST` request that specifies the environment, Kafka cluster, and connector
name. Include a JSON payload that specifies the delete type.

```bash
 POST /connect/v1/environments/{environment_id}/clusters/{kafka_cluster_id}/connectors/{connector_name}/offsets/request
 Host: https://api.confluent.cloud

{
  "type": "DELETE"
}
```

Considerations:

- Delete requests delete the offset for the provided partition and reset to the base state. A
  delete request is as if you created a fresh new connector.
- This is an asynchronous request. To check the status of this request, you must use the check offset status API. For more information,
  see **Get the status of an offset request**.
- Do not issue delete and patch requests at the same time.
- For source connectors, the connector attempts to read from the position defined in the base state.

**Response**:

Successful calls return HTTP `202 Accepted` with a JSON payload that describes the result.

```bash
{
  "id": "lcc-example123",
  "name": "{connector_name}",
  "offsets": [],
  "requested_at": "2024-03-28T17:59:45.606796307Z",
  "type": "DELETE"
}
```

Responses include the following information:

- Empty offsets.
- The time of the request to delete the offset.
- Information about Kafka cluster and connector.
- The type of request.

### Get the status of an offset request

To get the status of a previous offset request, make a `GET` request that specifies the environment, Kafka cluster, and connector
name.

```bash
GET /connect/v1/environments/{environment_id}/clusters/{kafka_cluster_id}/connectors/{connector_name}/offsets/request/status
Host: https://api.confluent.cloud
```

Considerations:

- The status endpoint always shows the status of the most recent PATCH/DELETE operation.

**Response**:

Successful calls return HTTP `200` with a JSON payload that describes the result. The following is an example
of an applied patch.

```bash
{
   "request": {
      "id": "lcc-example123",
      "name": "{connector_name}",
      "offsets": [
          {
              "partition": {
                  "taskId": "lcc-example123-0"
              },
              "offset": {
                  "completedFiles": "[{\"filePath\":\"source/file_0\",\"creationTime\":\"2024-03-06T17:30:28.391Z\"},{\"filePath\":\"source/file_7\",\"creationTime\":\"2024-03-06T17:30:28.395Z\"},{\"filePath\":\"source/file_9\",\"creationTime\":\"2024-03-06T17:30:28.409Z\"},{\"filePath\":\"source/file_1\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_8\",\"creationTime\":\"2024-03-06T17:30:28.681Z\"},{\"filePath\":\"source/file_6\",\"creationTime\":\"2024-03-06T17:30:28.715Z\"},{\"filePath\":\"source/file_30\",\"creationTime\":\"2024-03-06T17:30:28.969Z\"},{\"filePath\":\"source/file_39\",\"creationTime\":\"2024-03-06T17:30:28.970Z\"},{\"filePath\":\"source/file_37\",\"creationTime\":\"2024-03-06T17:30:28.993Z\"},{\"filePath\":\"source/file_36\",\"creationTime\":\"2024-03-06T17:30:29.265Z\"},{\"filePath\":\"source/file_31\",\"creationTime\":\"2024-03-06T17:30:29.268Z\"},{\"filePath\":\"source/file_38\",\"creationTime\":\"2024-03-06T17:30:29.278Z\"},{\"filePath\":\"source/file_25\",\"creationTime\":\"2024-03-06T17:30:29.549Z\"},{\"filePath\":\"source/file_22\",\"creationTime\":\"2024-03-06T17:30:29.551Z\"},{\"filePath\":\"source/file_13\",\"creationTime\":\"2024-03-06T17:30:29.552Z\"},{\"filePath\":\"source/file_47\",\"creationTime\":\"2024-03-06T17:30:30.015Z\"},{\"filePath\":\"source/file_14\",\"creationTime\":\"2024-03-06T17:30:30.020Z\"},{\"filePath\":\"source/file_40\",\"creationTime\":\"2024-03-06T17:30:30.028Z\"},{\"filePath\":\"source/file_15\",\"creationTime\":\"2024-03-06T17:30:30.305Z\"}]",
                  "earliestIncomplete": "2024-03-06T17:30:28.391Z",
                  "recordNum": "0"
              }
          }
      ],
      "requested_at": "2024-03-28T17:58:45.606796307Z",
      "type": "PATCH"
   },
   "status": {
      "phase": "APPLIED",
      "message": "The Connect framework-managed offsets for this connector have been altered successfully. However, if this connector manages offsets externally, they will need to be manually altered in the system that the connector uses."
   },
   "previous_offsets": [
       {
           "partition": {
               "taskId": "lcc-example123-0"
           },
           "offset": {
               "completedFiles": "[{\"filePath\":\"source/file_31\",\"creationTime\":\"2024-03-06T17:30:29.268Z\"},{\"filePath\":\"source/file_38\",\"creationTime\":\"2024-03-06T17:30:29.278Z\"},{\"filePath\":\"source/file_25\",\"creationTime\":\"2024-03-06T17:30:29.549Z\"},{\"filePath\":\"source/file_22\",\"creationTime\":\"2024-03-06T17:30:29.551Z\"},{\"filePath\":\"source/file_13\",\"creationTime\":\"2024-03-06T17:30:29.552Z\"},{\"filePath\":\"source/file_47\",\"creationTime\":\"2024-03-06T17:30:30.015Z\"},{\"filePath\":\"source/file_14\",\"creationTime\":\"2024-03-06T17:30:30.020Z\"},{\"filePath\":\"source/file_40\",\"creationTime\":\"2024-03-06T17:30:30.028Z\"},{\"filePath\":\"source/file_15\",\"creationTime\":\"2024-03-06T17:30:30.305Z\"},{\"filePath\":\"source/file_49\",\"creationTime\":\"2024-03-06T17:30:30.313Z\"},{\"filePath\":\"source/file_12\",\"creationTime\":\"2024-03-06T17:30:30.326Z\"},{\"filePath\":\"source/file_23\",\"creationTime\":\"2024-03-06T17:30:30.600Z\"},{\"filePath\":\"source/file_24\",\"creationTime\":\"2024-03-06T17:30:30.613Z\"},{\"filePath\":\"source/file_48\",\"creationTime\":\"2024-03-06T17:30:30.639Z\"},{\"filePath\":\"source/file_46\",\"creationTime\":\"2024-03-06T17:30:30.899Z\"},{\"filePath\":\"source/file_41\",\"creationTime\":\"2024-03-06T17:30:30.926Z\"},{\"filePath\":\"source/file_3\",\"creationTime\":\"2024-03-06T17:30:30.927Z\"},{\"filePath\":\"source/file_4\",\"creationTime\":\"2024-03-06T17:30:31.198Z\"},{\"filePath\":\"source/file_5\",\"creationTime\":\"2024-03-06T17:30:31.220Z\"},{\"filePath\":\"source/file_2\",\"creationTime\":\"2024-03-06T17:30:31.225Z\"}]",
               "earliestIncomplete": "2024-03-06T17:30:29.268Z",
               "recordNum": "0"
           }
       }
   ],
   "applied_at": "2024-03-28T17:58:48.079141883Z"
}
```

Responses include the following information:

- The original request, including the time it was made.
- The status of the request: applied, pending, or failed.
- The time you issued the status request.
- The previous offsets. These are the offsets that the connector last updated
  prior to updating the offsets. Use these to try to restore the state of your connector
  if a patch update causes your connector to fail or to return a connector to its
  previous state after rolling back.

### JSON payload

The table below offers a description of the unique fields in the JSON payload for managing offsets of the object store connectors, including
the following connectors:

- Amazon S3 Source connector
- Azure Blob Storage Source connector
- Google Cloud Storage (GCS) Source connector

| Field                | Definition                                                                                                                                                                                                                                                                                                                                                 | Required/Optional   |
|----------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------------|
| `taskId`             | Represents the partition in the following format: `connector-name`-<`taskid`>[-`in-progress`]<br/><br/>- `connector-name` is the name of the connector.<br/>- `taskid` is the task id.<br/>- `in-progress` is conditional and only appears if a file is currently being sourced. After the file is processed, the file appears listed in `completedFiles`. | Required            |
| `earliestIncomplete` | The position of the latest offset. When a connectors starts or restarts, the connector reads the files<br/>with a creation time equal to or after `earliestIncomplete` offset. These files are sorted by creation time then filename.                                                                                                                      | Required            |
| `completedFiles`     | List of sourced files.                                                                                                                                                                                                                                                                                                                                     | Required            |
| `recordNum`          | Number of records sourced.                                                                                                                                                                                                                                                                                                                                 | Required            |

## Quick Start

Use this quick start to get up and running with the Confluent Cloud Amazon S3 Source
connector. The quick start provides the basics of selecting the connector and
configuring it to get files from an Amazon S3 bucket.

<a id="cc-s3-connect-source-prereqs"></a>

Prerequisites
: - Authorized access to a [Confluent Cloud](https://www.confluent.io/confluent-cloud/) cluster on Amazon Web Services (AWS), Microsoft Azure (Azure), or Google Cloud.
  - The Confluent CLI installed and configured for the cluster. See [Install the Confluent CLI](https://docs.confluent.io/confluent-cli/current/install.html).
  - [Schema Registry](../get-started/schema-registry.md#cloud-sr-config) must be enabled to use a Schema Registry-based format (for example, Avro, JSON_SR (JSON Schema), or Protobuf).
  - For networking considerations, see [Networking and DNS](overview.md#connect-internet-access-resources). To use a set of public egress IP addresses, see [Public Egress IP Addresses for Confluent Cloud Connectors](static-egress-ip.md#cc-static-egress-ips).
  - An AWS [User Account IAM Policy](cc-s3-sink/cc-s3-sink.md#cc-s3-bucket-policy) configured for bucket access.
  - An AWS account configured with [Access Keys](https://docs.aws.amazon.com/general/latest/gr/aws-sec-cred-types.html#access-keys-and-secret-access-keys). You use these access keys when setting up the connector.
  <br/>
  - Kafka cluster credentials. The following lists the different ways you can provide credentials.
    - Enter an existing [service account](service-account.md#s3-cloud-service-account) resource ID.
    - Create a Confluent Cloud [service account](service-account.md#s3-cloud-service-account) for the connector. Make sure to review the ACL entries required in the [service account documentation](service-account.md#s3-cloud-service-account). Some connectors have specific ACL requirements.
    - Create a Confluent Cloud API key and secret. To create a key and secret, you can use [confluent api-key create](https://docs.confluent.io/confluent-cli/current/command-reference/api-key/confluent_api-key_create.html) *or* you can autogenerate the API key and secret directly in the Cloud Console when setting up the connector.
  <br/>
  - Confluent Cloud Schema Registry must be enabled for your cluster, if you are using a messaging  schema (like [Apache Avro](https://avro.apache.org/docs/current/)). See [Work with schemas](../sr/schemas-manage.md#cloud-schemas-manage).

### Using the Confluent Cloud Console

#### Step 1: Launch your Confluent Cloud cluster

To create and launch a Kafka cluster in Confluent Cloud, see [Create a kafka cluster in Confluent Cloud](../get-started/index.md#cloud-create-kafka-cluster).

#### Step 2: Add a connector

In the left navigation menu, click **Connectors**. If you already have connectors in your cluster, click **+ Add
connector**.

#### Step 3: Select your connector

Click the **Amazon S3 Source** connector card.

![Amazon S3 Source connector card](images/ccloud-s3-source-icon.png)

<a id="cc-s3-source-setup-connection"></a>

#### Step 4: Enter the connector details

#### NOTE
* Make sure you have all your [prerequisites](#cc-s3-connect-source-prereqs) completed.
* An asterisk ( \* ) designates a required entry.

At the **Add Amazon S3 Source Connector** screen, complete the following:

### Kafka access

1. Select the way you want to provide **Kafka Cluster credentials**. You can
   choose one of the following options:
   - **My account**: This setting allows your connector to globally access everything
     that you have access to. With a user account, the connector uses an API key and
     secret to access the Kafka cluster. This option is not recommended for production.
   - **Service account**: This setting limits the access for your connector by using a
     [service account](service-account.md#s3-cloud-service-account). This option is recommended for
     production.
   - **Use an existing API key**: This setting allows you to specify an API key and a
     secret pair. You can use an existing pair or create a new one. This method is not
     recommended for production environments.

   #### NOTE
   Freight clusters support only service accounts for Kafka authentication.
2. Click **Continue**.

### Authentication

1. Configure the authentication properties:

   **AWS credentials**
   - **Authentication method**: Under **AWS credentials**, select how you want to authenticate with AWS:
     - If you select **Access Keys**, enter your AWS credentials in the **Amazon Access Key ID** and **Amazon Secret Access Key** fields to connect to Amazon S3. For information about how to set these up, see [Access Keys](https://docs.aws.amazon.com/general/latest/gr/aws-sec-cred-types.html#access-keys-and-secret-access-keys).
     - If you select **IAM Roles**, choose an existing integration name under **Provider integration name** dropdown that has access to your resource. For more information, see [Manage Provider Integration for Fully Managed Connectors in Confluent Cloud](provider-integration.md#cloud-pi-quickstart).
   - **AWS access key ID**: Enter your AWS Access Key that lets the connector access your Amazon S3 resources if you select **Access Keys** as your authentication method.
   - **Provider Integration**: Select an existing integration that has access to your resource if you select **IAM Roles** as your authentication method.
   - **AWS secret access key**: Enter your AWS Secret Key that lets the connector access your Amazon S3 resources if you select **Access Keys** as your authentication method.

   **How should we connect to your S3 bucket?**
   - **S3 bucket name**: S3 bucket name.
   - **AWS Region**: Specify the AWS region where your S3 bucket resides.
   - **S3 Path-style Access**: Whether to use S3 path-style access. For
     more information, see the AWS [Path-style access](https://docs.aws.amazon.com/AmazonS3/latest/userguide/access-bucket-intro.html#path-style-url-ex)
     documentation.
2. Click **Continue**.

### Configuration

**Input and output messages**

- **Input Message Format**: Select an **Input Message Format**. Supports Avro, Bytes, CSV, JSON,
  and Parquet format. A valid schema must be available in [Schema
  Registry](../get-started/schema-registry.md#cloud-sr-config) to use a schema-based message format,
  like Avro.
- **Output Message Format**: Select an **Output Message Format**. Defaults to the file format
  selected for the input message format. Supports Avro, Bytes, JSON,
  JSON Schema, Protobuf, and String. A valid schema must be available in
  [Schema Registry](../get-started/schema-registry.md#cloud-sr-config) if using a schema-based
  format.

**Which topic(s) do you want to send data to?**

- **Topic Name Regex Patterns**: Enter the **Topic Name Regex Patterns**: A comma-separated list of
  pairs in the format `<kafka topic>:<regex>`. The connector uses
  this list to map file paths to Kafka topics. For example, the
  property `topic1:.*\.json` sources all files ending in `.json`
  to a Kafka topic named `topic1`. You can specify multiple of these
  `<kafka topic>:<regex>` mappings to send different sets of files
  to different topics. Any files that aren’t mapped by a regex are
  ignored. The connector sends files that match multiple mappings to
  the first topic in the list that maps the file.

**Storage**

- **Topics directory**: Top-level directory name where data to be
  ingested is stored. Defaults to `topics`.

  #### NOTE
  If you enter a blank space instead of accepting the default option
  `topics`, the connector reads all the data specified under the
  Amazon S3 bucket.

**Data encryption**

- Enable **Client-Side Field Level Encryption**
  for data encryption. Specify a **Service Account** to
  access the Schema Registry and associated encryption rules or keys with that schema. For more
  information on CSFLE or CSPE setup,
  see [Manage encryption for connectors](csfle.md#connect-csfle).

### **Show advanced configurations**

- **Schema context**: Select a schema context to use for this connector, if using
  a schema-based data format. This property defaults to the **Default** context,
  which configures the connector to use the default schema set up for Schema Registry in your
  Confluent Cloud environment. A schema context allows you to use separate schemas (like
  schema sub-registries) tied to topics in different Kafka clusters that share the
  same Schema Registry environment. For example, if you select a non-default context, a
  **Source** connector uses only that schema context to register a schema and a
  **Sink** connector uses only that schema context to read from. For more
  information about setting up a schema context, see [What are schema contexts and when should you use them?](../sr/faqs-cc.md#faq-schema-contexts).

**Additional Configs**

- **Value Converter Replace Null With Default**: Specifies whether to replace fields that have a default value and that are null to the default value. When set to `true`, the connector uses the default value; otherwise, it uses `null`. Applies to the `JSON` converter.
- **Value Converter Reference Subject Name Strategy**: Sets the subject reference name strategy for values. Valid entries are `DefaultReferenceSubjectNameStrategy` or `QualifiedReferenceSubjectNameStrategy`. You can use this strategy only with `PROTOBUF` format; the default strategy is `DefaultReferenceSubjectNameStrategy`.
- **Value Converter Schemas Enable**: Includes schema within each of the serialized values. Input messages must contain `schema` and `payload` fields and must not contain additional fields. For plain `JSON` data, set this to `false`. Applies to the `JSON` converter.
- **Errors Tolerance**: Use this property to configure the connector’s error handling behavior.

  #### WARNING
  Use this property with caution for sink connectors, as it can lead to data loss. If you set this property to `all`, the connector does not fail on errant records, but logs them (and sends to DLQ for sink connectors) and continues processing. If you set this property to `none`, the connector task fails on errant records.
- **Value Converter Ignore Default For Nullables**: When set to `true`, this property ensures that the corresponding record in Kafka is `null`, instead of showing the default column value. Applies to the `AVRO`, `PROTOBUF`, and `JSON_SR` converters.
- **Value Converter Decimal Format**: Specifies the `JSON` or `JSON_SR` serialization format for Connect `DECIMAL` logical type values with two allowed literals:
  `BASE64` to serialize `DECIMAL` logical types as base64 encoded binary data, and
  `NUMERIC` to serialize `DECIMAL` logical type values in `JSON` or `JSON_SR` as a number representing the decimal value.
- **Key Converter Schema ID Serializer**: The class name of the schema ID serializer for keys. This is used to serialize schema IDs in the message headers.
- **Value Converter Connect Meta Data**: Enables the Connect converter to add its metadata to the output schema. Applies to Avro converters.
- **Value Converter Value Subject Name Strategy**: Determines how to construct the subject name under which the value schema is registered with Schema Registry.
- **Key Converter Key Subject Name Strategy**: Determines how to construct the subject name for key schema registration.
- **Value Converter Schema ID Serializer**: The class name of the schema ID serializer for values. This is used to serialize schema IDs in the message headers.

**Auto-restart policy**

- **Enable Connector Auto-restart**: Enables the auto-restart behavior of the connector and its
  task in the event of user-actionable errors. Defaults to `true`, enabling the connector to
  automatically restart in case of user-actionable errors. Set this property to `false` to
  disable auto-restart for failed connectors. If disabled, you must manually restart the connector.

**Data polling policy**

- **S3 poll interval (ms)**: Frequency in milliseconds to poll for new or removed folders. This may result in updated task configurations starting to poll for data in added folders or stopping polling for data in removed folders. Defaults to `60000` ms (one minute).
- **Max records per poll**: The maximum amount of records to return
  each time the connector polls storage. Defaults to `200`. The
  maximum value supported is `10000` and the minimum value is
  `1`.

**How should we connect to your S3 bucket?**

- **Number of Retries on S3 Errors**: The number of times a single
  S3 API call should be retried in the case that it fails with a
  retriable error (such as a throttling exception). Once this limit
  is exceeded, the Kafka Connect poll itself may retry (based upon
  the Kafka Connect-based retry configuration).
- **Retry Backoff on S3 Errors (ms)**: How long to wait in
  milliseconds before attempting the first retry of a failed S3
  request. Upon a failure, this connector  may wait up to twice as
  long as the previous wait, up to the maximum number of retries.
  This avoids retrying in a tight loop under failure scenarios.
- **S3 Accelerated Endpoint**: Use an S3 accelerated endpoint.
- **Send S3 Expect Continue Request**: Enable/disable use of the
  HTTP/1.1 handshake using `EXPECT: 100-CONTINUE` during a multi-
  part upload. If `true`, the client waits for a 100
  (CONTINUE) response before sending the request body. If `false`, the
  client uploads the entire request body without checking if the
  server is willing to accept the request.
- **S3 Server Side Encryption Algorithm**: The S3 server-side
  encryption algorithm.
- **S3 Server Side Encryption Customer-Provided Key (SSE-C)**: The
  S3 Server-Side Encryption customer-provided key (SSE-C).

**Storage**

- **Task Batch Size**: The number of files assigned to each task at
  a time. Defaults to `10`. The maximum value supported is
  `2000` and the  minimum value is `1`.
- **File Discovery Starting Timestamp**: A Unix timestamp–that is,
  seconds since Jan 1, 1970 UTC–in epoch milliseconds that denotes
  where to start processing files. Any file encountered with a
  creation time earlier than this will be ignored. Note that this
  configuration property should only be used when there are no
  stored offsets for a connector–that is, this parameter is intended
  for new connectors to start from a specific timestamp rather than
  reading all the files in a bucket.
- **Directory Delimiter Character**: The pattern to use as the
  delimiter character for directories. Defaults to `/`.
- **Behavior on Errors**: Error handling behavior setting for storage connectors. Must be configured to one of the following: `IGNORE` or `FAIL`.
- **Byte Array Line Separator**: String inserted between records for ByteArrayFormat. Defaults to `\n` and may contain escape sequences like `\n`.  An input record that contains the line separator looks like multiple records in the storage object input.
- **Enable Embedded JSON Schema**: Enable reading of JSON messages with schema embedded.
- **CSV - Separator character**: The character that separates each field in the form of an integer. Typically in a CSV file, this is a `,` (`44`) character. A TSV file would use a tab (`9`) character. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Treat first row as header**: Flag to indicate if the fist row of data contains the header of the file. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Null field indicator**: Indicator to let the CSV Reader determine if a field is null. For more information, see CSVReaderNullFIeldIndicator <http://opencsv.sourceforge.net/apidocs/com/opencsv/enums/CSVReaderNullFieldIndicator.html>_\_. Applicable only if \`\`input.data.format\` is set to `CSV`.
- **CSV - Value schema**: The schema for the value written to Kafka. A default schema will be auto-generated if no value schema is provided. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - File character set**: Character set to read file with. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Skip lines**: The number of lines to skip in the beginning of the file. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Escape character**: The character as an integer to use when a special character is encountered. The default escape character is typically a `\` (`92`). Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Quote character**: The character that is used to quote a field. Typically in a CSV file,  this is a `"` (`34`) character. This happens when the `csv.separator.char` is within the data. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Ignore leading whitespace**: Sets the ignore leading whitespace setting. If `true`, the white space in front of a quote in a field is ignored. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Ignore quotations**: Sets the ignore quotations mode. If `true`, quotations are ignored. Applicable only if `input.data.format` is set to `CSV`.
- **CSV - Use strict quotes**: Sets the strict quotes setting. If `true`, characters outside the quotes are ignored. Applicable only if `input.data.format` is set to `CSV`.

**Headers**

- **Include file metadata in record headers**: When enabled, each produced Kafka record carries headers
  describing the source file: `file.name`, `file.path`,
  `file.last.modified` (TIMESTAMP), and `file.size` (LONG).

  Consumers that don’t read these headers are unaffected.

  Headers whose underlying metadata is unavailable from the AWS
  S3 bucket are omitted. Supported `header.converter` values
  are `SimpleHeaderConverter` (default), `StringConverter`,
  and `JsonConverter`.

**Transforms**

- **Single Message Transformations**: To add a new SMT, see [Add transforms](single-message-transforms.md#cc-single-message-transforms-ui).
  For more information about unsupported SMTs, see
  [Unsupported transformations](single-message-transforms.md#cc-single-message-transforms-unsupported-transforms).

For all property values and definitions, see
[Configuration Properties](#cc-s3-source-config-properties).

- Click **Continue**.

### Sizing

Based on the number of topic partitions you select, you will be provided
with a recommended number of tasks.

1. To change the number of tasks, use the Range Slider to select the
   desired number of tasks.
2. Click **Continue**.

### Review and Launch

1. Verify the connection details by previewing the running configuration.
2. Once you’ve validated that the properties are configured to your
   satisfaction, click **Launch**.

   The status for the connector should go from **Provisioning** to
   **Running**.

#### Step 5. Check the Kafka topic

After the connector is running, verify that records are populating the Kafka topic.

#### NOTE
The S3 Source connector loads and filters all object names in the bucket
before it starts sourcing records. When starting up, the connector may
display `RUNNING` but not show any throughput. This is because bucket
loading is not finished. For buckets with a large amount of objects, bucket
loading can take several minutes to complete.

Creating **more top-level folders** will help process files with less delays
and scale better in the long term.

For more information and examples to use with the Confluent Cloud API for Connect,
see the [Confluent Cloud API for Connect Usage Examples](connect-api-section.md#ccloud-connect-api) section.

### Using the Confluent CLI

Complete the following steps to set up and run the connector using the Confluent CLI.

#### NOTE
Make sure you have all your [prerequisites](#cc-s3-connect-source-prereqs) completed.

#### Step 1: List the available connectors

Enter the following command to list available connectors:

```none
confluent connect plugin list
```

#### Step 2: List the connector configuration properties

Enter the following command to show the connector configuration properties:

```none
confluent connect plugin describe <connector-plugin-name>
```

The command output shows the required and optional configuration properties.

#### Step 3: Create the connector configuration file

Create a JSON file that contains the connector configuration properties. The following example shows the required connector properties.

```json
{
  "connector.class": "S3Source",
  "name": "S3SourceConnector_0",
  "topic.regex.list": "topic1:.*\.json",
  "topics.dir": " ",
  "kafka.auth.mode": "SERVICE_ACCOUNT",
  "kafka.service.account.id": "<service-account-resource-ID>",
  "input.data.format": "JSON",
  "output.data.format": "BYTES",
  "aws.access.key.id": "<access-key>",
  "aws.secret.access.id": "<secret-access-id>",
  "s3.bucket.name": "<bucket-name>",
  "tasks.max": "1",
}
```

Note the following required property definitions:

* `"connector.class"`: Identifies the connector plugin name.
* `"name"`: Sets a name for your new connector.
* `"topic.regex.list"`: A comma-separated list of pairs in the format `<kafka topic>:<regex>`. The connector uses this list to map file paths to Kafka topics. For example, the property `topic1:.*\.json` sources all files ending in `.json` to a Kafka topic named `topic1`. You can specify multiple of these `<kafka topic>:<regex>` mappings to send different sets of files to different topics. Any files that aren’t mapped by a regex are ignored. The connector sends files that match multiple mappings to the first topic in the list that maps the file.

  #### NOTE
  For more information about accepted regular expressions, see [Google RE2 syntax](https://github.com/google/re2/wiki/Syntax/).
* `"topics.dir"`: (Optional) If this property is not used, the default folder where the connector reads data from is `topics`. If you set this property to a blank space (as shown in the example configuration), the connector reads all data in the S3 bucket.

* `"kafka.auth.mode"`: Identifies the connector authentication mode you want to use. There are two options: `SERVICE_ACCOUNT` or `KAFKA_API_KEY` (the default). To use an API key and secret, specify the configuration properties `kafka.api.key` and `kafka.api.secret`, as shown in the example configuration (above).  To use a [service account](service-account.md#s3-cloud-service-account), specify the **Resource ID** in the property `kafka.service.account.id=<service-account-resource-ID>`. To list the available service account resource IDs, use the following command:
  ```bash
  confluent iam service-account list
  ```

  For example:
  ```bash
  confluent iam service-account list

     Id     | Resource ID |       Name        |    Description
  +---------+-------------+-------------------+-------------------
     123456 | sa-l1r23m   | sa-1              | Service account 1
     789101 | sa-l4d56p   | sa-2              | Service account 2
  ```

* `"input.data.format"`: Supports Avro, Bytes, CSV, JSON, and Parquet format.
  A valid schema must be available in [Schema Registry](../get-started/schema-registry.md#cloud-sr-config)
  to use a schema-based message format, like Avro.
* `"output.data.format"`: Sets the output Kafka record value format. Options
  are Avro, Bytes, JSON, JSON Schema, Protobuf, and String. A valid schema must
  be available in [Schema Registry](../get-started/schema-registry.md#cloud-sr-config) if using a
  schema-based format.
* `"value.schema"`: The schema for the value written to Kafka. A default schema will be auto-generated if
  no value schema is provided. Applicable only if `input.data.format` is set to `CSV`.

  For example:
  ```json
  {
     "type": "STRUCT",
     "fieldSchemas": {
       "age": {"type": "INT64"},
       "firstname": {"type": "STRING"},
       "lastname": {"type": "STRING"}
     }
   }
  ```
* `"tasks.max"`: The total number of tasks to run in parallel. More tasks may improve performance.
* Transforms and Predicates: See the [Single Message Transformation (SMT)](single-message-transforms.md#cc-single-message-transforms) documentation for details.

#### NOTE
To enable CSFLE or CSPE for data encryption, specify the following properties:

* `csfle.enabled`: Flag to indicate whether the connector honors CSFLE or CSPE rules.
* `sr.service.account.id`: A Service Account to access the Schema Registry and associated encryption rules or keys with that schema.

For more information on CSFLE or CSPE setup, see [Manage encryption for connectors](csfle.md#connect-csfle).

For configuration property values and descriptions, see [Configuration Properties](#cc-s3-source-config-properties).

#### Step 4: Load the properties file and create the connector

Enter the following command to load the configuration and start the connector:

```none
confluent connect cluster create --config-file <file-name>.json
```

For example:

```none
confluent connect cluster create --config-file s3-source-config.json
```

Example output:

```none
Created connector S3SourceConnector_0 lcc-ix4dl
```

#### Step 5: Check the connector status

Enter the following command to check the connector status:

```none
confluent connect cluster list
```

Example output:

```none
ID          |       Name            | Status  | Type
+-----------+-----------------------+---------+------+
lcc-ix4dl   | S3SourceConnector_0   | RUNNING | source
```

#### Step 6. Check the Kafka topic

After the connector is running, verify that records are populating the Kafka topic.

#### NOTE
The S3 Source connector loads and filters all object names in the bucket
before it starts sourcing records. When starting up, the connector may
display `RUNNING` but not show any throughput. This is because bucket
loading is not finished. For buckets with a large amount of objects, bucket
loading can take several minutes to complete.

For more information and examples to use with the Confluent Cloud API for Connect,
see the [Confluent Cloud API for Connect Usage Examples](connect-api-section.md#ccloud-connect-api) section.

<a id="cc-s3-source-config-properties"></a>

## Configuration Properties

Use the following configuration properties with the fully managed connector. For
self-managed connector property definitions and other details, see the connector
docs in [Self-managed connectors for Confluent Platform](/platform/current/connect/kafka_connectors.html).

### How should we connect to your data?

`name`
: Sets a name for your connector.
  <br/>
  * Type: string
  * Valid Values: A string at most 64 characters long
  * Importance: high

### Which topic(s) do you want to send data to?

`topic.regex.list`
: A list of topics along with a regex expression of the files which are to be sent to that topic.  For example: “my-topic:.\*” will send all files to “my-topic”, while a list containing only the expression “special-topic:.\*.json” will send only files starting with “.json” to “special-topic”, and all other files not matching any patterns will be ignored and not sourced. Files that match multiple mappings will be sent to the first topic in the list that maps the file. The `topic.regex.list` property matches the full path (for example, `folder/file.txt`), not just the filename.
  <br/>
  * Type: list
  * Importance: high

### Kafka Cluster credentials

`kafka.auth.mode`
: Kafka Authentication mode. It can be one of KAFKA_API_KEY or SERVICE_ACCOUNT. It defaults to KAFKA_API_KEY mode, whenever possible.
  <br/>
  * Type: string
  * Valid Values: SERVICE_ACCOUNT, KAFKA_API_KEY
  * Importance: high

`kafka.api.key`
: Kafka API Key. Required when kafka.auth.mode==KAFKA_API_KEY.
  <br/>
  * Type: password
  * Importance: high

`kafka.service.account.id`
: The Service Account that will be used to generate the API keys to communicate with Kafka Cluster.
  <br/>
  * Type: string
  * Importance: high

`kafka.api.secret`
: Secret associated with Kafka API key. Required when kafka.auth.mode==KAFKA_API_KEY.
  <br/>
  * Type: password
  * Importance: high

### Schema Config

`schema.context.name`
: Add a schema context name. A schema context represents an independent scope in Schema Registry. It is a separate sub-schema tied to topics in different Kafka clusters that share the same Schema Registry instance. If not used, the connector uses the default schema configured for Schema Registry in your Confluent Cloud environment.
  <br/>
  * Type: string
  * Default: default
  * Importance: medium

### Input and output messages

`input.data.format`
: Sets the input message format. Valid entries are AVRO, JSON, or BYTES. Note that you need to have Confluent Cloud Schema Registry configured if using a schema-based message format like AVRO.
  <br/>
  * Type: string
  * Valid Values: AVRO, BYTES, CSV, JSON, PARQUET
  * Importance: high

`output.data.format`
: Set the output message format for values. Valid entries are AVRO, JSON, JSON_SR, PROTOBUF, STRING, or BYTES. Note that you need to have Confluent Cloud Schema Registry configured if using a schema-based message format like AVRO, JSON_SR and PROTOBUF. If no value for this property is provided, the value specified for the ‘input.data.format’ property is used.
  <br/>
  * Type: string
  * Valid Values: AVRO, BYTES, JSON, JSON_SR, PROTOBUF, STRING
  * Importance: high

### AWS credentials

`authentication.method`
: Select how you want to authenticate with AWS.
  <br/>
  * Type: string
  * Default: Access Keys
  * Importance: high

`aws.access.key.id`
: The AWS Access Key used to connect to Amazon S3.
  <br/>
  * Type: password
  * Importance: high

`provider.integration.id`
: Select an existing integration that has access to your resource. In case you need to integrate a new IAM role, use provider integration
  <br/>
  * Type: string
  * Importance: high

`aws.secret.access.key`
: The AWS Secret Key used to connect to Amazon S3.
  <br/>
  * Type: password
  * Importance: high

### How should we connect to your S3 bucket?

`s3.bucket.name`
: * Type: string
  * Importance: high

`s3.region`
: Set to the AWS region where your S3 bucket resides.
  <br/>
  * Type: string
  * Importance: high

`s3.part.retries`
: The number of times a single S3 API call should be retried in the case that it fails with a “retriable” error (such as a throttling exception). Once this limit is exceeded, the Kafka Connect poll itself may retry (based upon the Kafka Connect-based retry configuration).
  <br/>
  * Type: int
  * Default: 3
  * Importance: medium

`s3.retry.backoff.ms`
: How long to wait in milliseconds before attempting the first retry of a failed S3 request. Upon a failure, this connector  may wait up to twice as long as the previous wait, up to the maximum number of retries. This avoids retrying in a tight loop under failure scenarios.
  <br/>
  * Type: int
  * Default: 200
  * Importance: medium

`ui.s3.wan.mode`
: Use an S3 accelerated endpoint.
  <br/>
  * Type: string
  * Default: NO
  * Valid Values: NO, YES
  * Importance: medium

`ui.s3.path.style.access`
: Whether to use s3 path-style access.
  <br/>
  * Type: string
  * Default: NO
  * Valid Values: NO, YES
  * Importance: medium

`s3.http.send.expect.continue`
: Enable/disable use of the HTTP/1.1 handshake using EXPECT: 100-CONTINUE during multi-part upload. If true, the client waits for a 100 (CONTINUE) response before sending the request body. If false, the client uploads the entire request body without checking if the server is willing to accept the request.
  <br/>
  * Type: string
  * Default: YES
  * Valid Values: NO, YES
  * Importance: medium

`ui.s3.ssea.name`
: The S3 server-side encryption algorithm.
  <br/>
  * Type: string
  * Default: NONE
  * Valid Values: AES256, AWS:KMS, NONE
  * Importance: medium

`s3.sse.customer.key`
: The S3 Server-Side Encryption customer-provided key (SSE-C).
  <br/>
  * Type: password
  * Importance: medium

### Storage

`topics.dir`
: Top-level directory (in the S3 bucket) where data to be ingested is stored.
  <br/>
  * Type: string
  * Default: topics
  * Importance: high

`task.batch.size`
: The number of files assigned to each task at a time
  <br/>
  * Type: int
  * Default: 10
  * Valid Values: [1,…,2000]
  * Importance: high

`file.discovery.starting.timestamp`
: A Unix timestamp (in epoch milliseconds since Jan 1, 1970 UTC) that denotes where to start processing files. The connector ignores any file with a creation time earlier than this timestamp. Note that the connector only uses this configuration property when no offsets are stored for a connector. This parameter allows new connectors to start from a specific timestamp instead of reading all files in a bucket.
  <br/>
  * Type: long
  * Default: 0
  * Importance: high

`directory.delim`
: Directory delimiter pattern.
  <br/>
  * Type: string
  * Default: /
  * Importance: medium

`ui.behavior.on.error`
: Error handling behavior setting for storage connectors. Must be configured to one of the following: IGNORE, FAIL
  <br/>
  * Type: string
  * Default: FAIL
  * Valid Values: FAIL, IGNORE
  * Importance: medium

`format.bytearray.separator`
: String inserted between records for ByteArrayFormat. Defaults to n and may contain escape sequences like n.  An input record that contains the line separator looks like multiple records in the storage object input.
  <br/>
  * Type: string
  * Importance: medium

`format.json.schema.enable`
: Enable reading of JSON messages with schema embedded
  <br/>
  * Type: boolean
  * Default: false
  * Importance: medium

`csv.separator.char`
: The character that separates each field in the form of an integer. Typically in a CSV file, this is a `,` (`44`) character. A TSV file would use a tab (`9`) character. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: int
  * Default: 44
  * Importance: low

`csv.first.row.as.header`
: Flag to indicate if the fist row of data contains the header of the file. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: boolean
  * Default: true
  * Importance: medium

`csv.null.field.indicator`
: Indicator to determine how the CSV Reader can determine if a field is null. For more information, see [http://opencsv.sourceforge.net/apidocs/com/opencsv/enums/CSVReaderNullFieldIndicator.html](http://opencsv.sourceforge.net/apidocs/com/opencsv/enums/CSVReaderNullFieldIndicator.html). Applicable only if `input.data.format` is set to `CSV`  .
  <br/>
  * Type: string
  * Default: NEITHER
  * Importance: low

`value.schema`
: The schema for the value written to Kafka. A default schema will be auto-generated if no value schema is provided. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: string
  * Importance: high

`csv.file.charset`
: Character set to read file with. Applicable only if `input.data.format` is set to `CSV`
  <br/>
  * Type: string
  * Default: UTF-8
  * Importance: low

`csv.skip.lines`
: The number of lines to skip in the beginning of the file. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: int
  * Default: 0
  * Importance: low

`csv.escape.char`
: The character as an integer to use when a special character is encountered. The default escape character is typically a `\` (`92`). Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: int
  * Default: 92
  * Importance: low

`csv.quote.char`
: The character that is used to quote a field. Typically in a CSV file,  this is a `"` (`34`) character. This happens when the `csv.separator.char` is within the data. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: int
  * Default: 34
  * Importance: low

`csv.ignore.leading.whitespace`
: Sets the ignore leading whitespace setting. If `true`, the white space in front of a quote in a field is ignored. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: boolean
  * Default: true
  * Importance: low

`csv.ignore.quotations`
: Sets the ignore quotations mode. If `true`, quotations are ignored. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

`csv.strict.quotes`
: Sets the strict quotes setting. If `true`, characters outside the quotes are ignored. Applicable only if `input.data.format` is set to `CSV`.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

### Data polling policy

`s3.poll.interval.ms`
: Frequency in milliseconds to poll for new or removed folders. This may result in updated task configurations starting to poll for data in added folders or stopping polling for data in removed folders
  <br/>
  * Type: long
  * Default: 60000 (1 minute)
  * Valid Values: [1000,…]
  * Importance: medium

`record.batch.max.size`
: The maximum amount of records to return each time storage is polled.
  <br/>
  * Type: int
  * Default: 200
  * Valid Values: [1,…,10000]
  * Importance: medium

### Number of tasks for this connector

`tasks.max`
: The total number of tasks to run in parallel.
  <br/>
  * Type: int
  * Valid Values: [1,…,1000]
  * Importance: high

### Headers

`file.metadata.headers.enable`
: When enabled, each produced Kafka record carries headers describing the source file: `file.name`, `file.path`, `file.last.modified` (TIMESTAMP), and `file.size` (LONG). Consumers that don’t read these headers are unaffected. Headers whose underlying metadata is unavailable from the AWS S3 bucket are omitted. Supported `header.converter` values are `SimpleHeaderConverter` (default), `StringConverter`, and `JsonConverter`.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

### Additional Configs

`header.converter`
: The converter class for the headers. This is used to serialize and deserialize the headers of the messages.
  <br/>
  * Type: string
  * Importance: low

`producer.override.compression.type`
: The compression type for all data generated by the producer. Valid values are none, gzip, snappy, lz4, and zstd.
  <br/>
  * Type: string
  * Importance: low

`producer.override.linger.ms`
: The producer groups together any records that arrive in between request transmissions into a single batched request. More details can be found in the documentation: [https://docs.confluent.io/platform/current/installation/configuration/producer-configs.html#linger-ms](https://docs.confluent.io/platform/current/installation/configuration/producer-configs.html#linger-ms).
  <br/>
  * Type: long
  * Valid Values: [100,…,1000]
  * Importance: low

`value.converter.allow.optional.map.keys`
: Allow optional string map key when converting from Connect Schema to Avro Schema. Applicable for Avro Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.auto.register.schemas`
: Specify if the Serializer should attempt to register the Schema.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.connect.meta.data`
: Allow the Connect converter to add its metadata to the output schema. Applicable for Avro Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.enhanced.avro.schema.support`
: Enable enhanced schema support to preserve package information and Enums. Applicable for Avro Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.enhanced.protobuf.schema.support`
: Enable enhanced schema support to preserve package information. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.flatten.unions`
: Whether to flatten unions (oneofs). Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.generate.index.for.unions`
: Whether to generate an index suffix for unions. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.generate.struct.for.nulls`
: Whether to generate a struct variable for null values. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.int.for.enums`
: Whether to represent enums as integers. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.latest.compatibility.strict`
: Verify latest subject version is backward compatible when use.latest.version is true.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.object.additional.properties`
: Whether to allow additional properties for object schemas. Applicable for JSON_SR Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.optional.for.nullables`
: Whether nullable fields should be specified with an optional label. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.optional.for.proto2`
: Whether proto2 optionals are supported. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.scrub.invalid.names`
: Whether to scrub invalid names by replacing invalid characters with valid characters. Applicable for Avro and Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.use.latest.version`
: Use latest version of schema in subject for serialization when auto.register.schemas is false.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.use.optional.for.nonrequired`
: Whether to set non-required properties to be optional. Applicable for JSON_SR Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.wrapper.for.nullables`
: Whether nullable fields should use primitive wrapper messages. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`value.converter.wrapper.for.raw.primitives`
: Whether a wrapper message should be interpreted as a raw primitive at root level. Applicable for Protobuf Converters.
  <br/>
  * Type: boolean
  * Importance: low

`errors.tolerance`
: Use this property if you would like to configure the connector’s error handling behavior. WARNING: This property should be used with CAUTION for SOURCE CONNECTORS as it may lead to dataloss. If you set this property to ‘all’, the connector will not fail on errant records, but will instead log them (and send to DLQ for Sink Connectors) and continue processing. If you set this property to ‘none’, the connector task will fail on errant records.
  <br/>
  * Type: string
  * Default: none
  * Importance: low

`key.converter.key.schema.id.serializer`
: The class name of the schema ID serializer for keys. This is used to serialize schema IDs in the message headers.
  <br/>
  * Type: string
  * Default: io.confluent.kafka.serializers.schema.id.PrefixSchemaIdSerializer
  * Importance: low

`key.converter.key.subject.name.strategy`
: How to construct the subject name for key schema registration.
  <br/>
  * Type: string
  * Default: TopicNameStrategy
  * Importance: low

`value.converter.decimal.format`
: Specify the JSON/JSON_SR serialization format for Connect DECIMAL logical type values with two allowed literals:
  <br/>
  BASE64 to serialize DECIMAL logical types as base64 encoded binary data and
  <br/>
  NUMERIC to serialize Connect DECIMAL logical type values in JSON/JSON_SR as a number representing the decimal value.
  <br/>
  * Type: string
  * Default: BASE64
  * Importance: low

`value.converter.flatten.singleton.unions`
: Whether to flatten singleton unions. Applicable for Avro and JSON_SR Converters.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

`value.converter.ignore.default.for.nullables`
: When set to true, this property ensures that the corresponding record in Kafka is NULL, instead of showing the default column value. Applicable for AVRO,PROTOBUF and JSON_SR Converters.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

`value.converter.reference.subject.name.strategy`
: Set the subject reference name strategy for value. Valid entries are DefaultReferenceSubjectNameStrategy or QualifiedReferenceSubjectNameStrategy. Note that the subject reference name strategy can be selected only for PROTOBUF format with the default strategy being DefaultReferenceSubjectNameStrategy.
  <br/>
  * Type: string
  * Default: DefaultReferenceSubjectNameStrategy
  * Importance: low

`value.converter.replace.null.with.default`
: Whether to replace fields that have a default value and that are null to the default value. When set to true, the default value is used, otherwise null is used. Applicable for JSON Converter.
  <br/>
  * Type: boolean
  * Default: true
  * Importance: low

`value.converter.schemas.enable`
: Include schemas within each of the serialized values. Input messages must contain schema and payload fields and may not contain additional fields. For plain JSON data, set this to false. Applicable for JSON Converter.
  <br/>
  * Type: boolean
  * Default: false
  * Importance: low

`value.converter.value.schema.id.serializer`
: The class name of the schema ID serializer for values. This is used to serialize schema IDs in the message headers.
  <br/>
  * Type: string
  * Default: io.confluent.kafka.serializers.schema.id.PrefixSchemaIdSerializer
  * Importance: low

`value.converter.value.subject.name.strategy`
: Determines how to construct the subject name under which the value schema is registered with Schema Registry.
  <br/>
  * Type: string
  * Default: TopicNameStrategy
  * Importance: low

### Auto-restart policy

`auto.restart.on.user.error`
: Enable connector to automatically restart on user-actionable errors.
  <br/>
  * Type: boolean
  * Default: true
  * Importance: medium

## FAQs

Find answers to frequently asked questions about the Amazon S3 Source connector.

### How can I improve the connector performance for buckets with many objects?

To improve the connector performance and reduce latency when working with S3 buckets
that contain a large number of objects:

* **Create more top-level folders**: Organize your files into multiple top-level folders. This
  helps the connector process files more efficiently and scale better.
* **Use appropriate regex patterns**: Configure the `topic.regex.list` property with specific patterns
  to filter only the required files, reducing the number of objects the connector must process.
* **Increase the number of tasks**: Use the `tasks.max` property to run more tasks in parallel.

### Why isn’t the connector reading files from a new bucket?

For a new bucket, you must create a new connector with a unique name. If you reconfigure an existing connector
to source from a new bucket, or create a connector with a name already used by another connector in the cluster,
the connector might not source from the beginning of the data.

This behavior occurs because the connector maintains offsets tied to the connector name.
Each connector instance must have a unique name when reading from a new bucket to ensure a fresh offset start.

### Why is the connector slow when reading many small files?

Reading a large number of small objects creates high overhead. To improve throughput, tune the following areas:

* **Increase parallelism**: Set a higher `tasks.max` value to process multiple files concurrently.
* **Batching**:
  * Increase `task.batch.size` to process more records per batch.
  * Adjust `record.batch.max.size` to read more records from each file at once.
  * Decrease `s3.poll.interval.ms` to scan for new and removed folders more frequently.

### The connector is running, but the Kafka topic has no data. What should I check?

If the connector status is `RUNNING` but no records appear in your Kafka topic, check the following:

* **Bucket loading phase**: At startup, the connector loads and filters all object names in the bucket before it
  sources records. Wait a few minutes, especially if the bucket has many objects. Creating more top-level folders
  reduces listing overhead and improves performance.
* **topic.regex.list pattern mismatch**: Verify the regex matches the entire object key, not just the
  filename. Include folder prefixes if necessary. If a file matches multiple mappings, the file is routed to the first
  matching topic in the list.
* **topics.dir or bucket root mismatch**: The `topics.dir` property defines the starting path the connector uses to scan for data.
  Verify the following:
  - Set `topics.dir` to `topics`. The connector only looks for data stored under the `topics/` prefix.
  - To read from the root of the bucket rather than a specific folder, set `topics.dir` to an empty value.
  - Ensure your S3 object keys begin with the string defined in `topics.dir`. If the paths don’t match,
    the connector skips the objects.
* **Input format and file content mismatch**: Ensure the `input.data.format` matches the actual file encoding of the S3 objects.

### Why does the connector stop processing files after switching to CSV format?

If the connector stopped processing files after you changed the input message format to CSV,
there is likely a structural mismatch between the S3 files and the connector configuration.

Verify your Confluent Cloud settings using the following configuration checklist:

* Confirm the input message format is set to CSV.
* Ensure the value in `csv.separator.char` matches the actual delimiter used in your source files.
* Set `csv.first.row.as.header` to correctly reflect whether the first row is a header row.
* Review settings for the character set, the number of lines to skip, and the specific quote or escape characters used in your files.

### How do I reprocess or skip specific S3 files?

Amazon S3 Source connector supports manual offset management through Confluent Cloud APIs. This allows you to control which
files to ingest or ignore. For more information about offsets, see [Connect offsets API reference](https://docs.confluent.io/cloud/current/ccloud/offsets-connect-v-1/).

* **Get current offsets**: Retrieve the current offset for each task. This includes `earliestIncomplete`,
  `completedFiles`, and the `recordNum`.
* **Update offsets**: Change the ingestion state to control which files the connector processes.
  * **Reprocess data**: Move `earliestIncomplete` to an earlier point. This may cause duplicate records.
  * **Skip older files**: Move `earliestIncomplete` to a later point.
  * **Mark files as processed**: Manually add specific file paths to the `completedFiles` list.
* **Delete offsets**: Reset the connector to its base state. This makes the connector behave as if it were newly created.

**Key considerations**

* Offset changes are not instantaneous. Track the progress using the offset request status API.
* Submit only one `PATCH` or `DELETE` request at a time for a single connector.
* Rolling offsets backward can re-emit records, while jumping forward skips data.

## Next Steps

* For an example that shows fully managed Confluent Cloud connectors in action with
  Confluent Cloud for Apache Flink, see the [Cloud ETL Demo](/platform/current/tutorials/examples/cloud-etl/docs/index.html).
  This example also shows how to use Confluent CLI to manage your resources in
  Confluent Cloud.
  [![image](images/topology.png)](https://docs.confluent.io/platform/current/tutorials/examples/cloud-etl/docs/index.html)
* Try [Confluent Cloud on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-g5ujul6iovvcy?trk=14575e70-1766-4f20-8083-0c2757a1ec75&sc_channel=el)
  with $1000 of free usage for 30 days, and pay as you go. No credit card is
  required.
