Connecting via Amazon S3
The Altrata Data Feed files can be delivered to you through Amazon S3. Your feed's files are made available in a dedicated S3 bucket, from which you can pull files directly into your own environment using standard AWS tooling.
This documentation walks through the onboarding information required to set up S3 access and how to connect to the bucket and download files.
If you have any questions, please contact [email protected].
Onboarding information
To provision your S3 access, your account representative will ask you to provide two values:
Value | Description |
|---|---|
Customer access principal ARN (customerAccessPrincipalArn) | The Amazon Resource Name (ARN) of the IAM principal in your AWS account - typically an IAM role, or an IAM user - that will be used to access the Altrata S3 bucket. Example format: arn:aws:iam::123456789012:role/altrata-datafeed-reader |
Customer external ID (customerExternalId) | A unique string of your choosing. Together with your principal ARN, this secures the trust relationship on the Altrata IAM role and allows only your principal to assume it. See "What is the external ID?" below. |
What is the external ID?
The external ID is a standard AWS mechanism used to secure cross-account access to S3, protecting against what AWS calls the "confused deputy" problem. It is included as a condition in the IAM role's trust policy, so the role can only be assumed when the correct external ID is supplied alongside your principal ARN. The external ID is not issued by AWS and is not retrieved from anywhere - it is simply a unique string chosen by the party that will access the data. Because you (the customer) are the accessing party, you choose the value. A randomly generated GUID (e.g. b8ee735e-0f96-4b44-96c2-166cdcce8dfc) is a good choice. Note that AWS does not treat the external ID as a secret; its purpose is uniqueness, not secrecy. For more detail, see AWS's documentation: Providing access to AWS accounts owned by third parties
Once provisioned, your account representative will provide you with the following:
- The ARN of a dedicated Altrata IAM role, created for your account, which your principal will assume to access the bucket. The role's trust policy is configured with your principal ARN and external ID, so only your principal - supplying your external ID - can assume it. You will need this role ARN for the final permissions configuration in the next step.
- The name of your dedicated S3 bucket
- The AWS Region of the bucket
Configuring permissions to assume the Altrata Access role
Once you have the ARN of the role you need to assume, a critical final configuration step is to set the permissions required for your Customer access principal to assume the Altrata role.
To allow this permission, add a policy to your Customer access principal allowing the AssumeRole action on the Altrata access role Resource. For example:
Connecting to the S3 bucket
You can access the bucket with any tooling that supports Amazon S3, including the AWS CLI, AWS SDKs, or data platform ingestion connectors. The examples below use the AWS CLI.
Step 1 - Assume the Altrata access role
Use your access principal to assume the role provided by Altrata, passing your external ID:
This returns temporary credentials (AccessKeyId, SecretAccessKey, SessionToken). You must use these credentials for the S3 commands in Steps 2 and 3 below - either export them as the AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and AWS_SESSION_TOKEN environment variables, or save them as a named profile and add --profile to each command. Running the S3 commands under your normal AWS identity, without assuming the role, will result in an Access Denied error.
Note: your access principal must be permitted, within your own AWS account, to call sts:AssumeRole on the role ARN provided by Altrata. Configuring this permission is managed by your organization.
Step 2 - View available directories and files
You should see one directory (prefix) per subscribed data feed package:
Step 3 - Copy the files
There are many options for copying the Altrata data feed files from the S3 bucket owned by Altrata to a destination location of your choice.
As a developer, if you are working locally, you could choose to copy the files to you location machine. See Examples 1 and 2 below.
Example 1 - Copy all files to a local directory using S3 Copy
Example 2 - Keep a local folder in sync with the S3 bucket using S3 Sync
However, in production, you would likely choose to copy the files from the Altrata-owned S3 bucket to another S3 bucket in your AWS account. See Example 3 and 4 below.
Example 3 - Copy all files to another S3 bucket using S3 Copy
Example 4 - Keep another S3 bucket in sync with the Altrata bucket using S3 Sync
Storage capacity and bandwidth considerations
Before downloading data from S3, you should ensure that you have sufficient disk space available on your system to store the delivery files. Disk space requirements will vary depending on the data packages you wish to download. Full data files can often exceed 1 GB; it is recommended to have sufficiently high-speed internet access bandwidth to download these files.