CloudFront origin failover

On this page 6

A site served from one S3 bucket goes down with that bucket's region. Origin failover gives CloudFront a second origin to try. On a cache miss CloudFront asks the primary; when the primary answers with one of the configured status codes, refuses the connection, or times out, CloudFront retries the same request against the secondary and serves its answer.

Stacks configures this through config/cloud.ts and deploys it with buddy deploy on the AWS target. It is off until you turn it on. The CloudFront side is an origin group, generated by ts-cloud.

A replica bucket in another region

Add failover to a website bucket:

// config/cloud.ts
export const tsCloud: TsCloudConfig = {
  project: { name: 'my-app', slug: 'my-app', region: 'us-east-1' },
  infrastructure: {
    dns: { domain: 'example.com', hostedZoneId: 'Z0123456789' },
    ssl: { enabled: true },
    storage: {
      public: {
        website: true,
        failover: {
          region: 'us-west-2',
        },
      },
    },
  },
}

On the next buddy deploy:

  1. Before the stack, the deploy creates the replica bucket in us-west-2: private, encrypted, versioned. A CloudFormation stack can only create buckets in its own region, so the replica cannot be part of it.
  2. The stack turns on versioning and S3 replication on the primary bucket, so every write is copied to the replica, and gives the distribution a second origin and an origin group. The default cache behavior serves from the group.
  3. After the stack, the deploy adds one statement to the replica's bucket policy letting that distribution, and only that distribution, read it through origin access control. Then it copies across any object the replica does not have yet, because S3 replication only copies writes made after it was turned on.
KeyDefaultMeaning
regionrequiredRegion of the replica. Must differ from project.region.
bucket<bucket>-<region>Replica bucket name. Set it when the default passes 63 characters or is taken.
statusCodes[403, 404, 500, 502, 503, 504]Primary responses that send the request to the replica.
connectionAttempts3 (CloudFront)Times CloudFront tries to connect to the primary, 1 to 3.
connectionTimeout10 (CloudFront)Seconds CloudFront waits for the primary to accept a connection, 1 to 10.
replicatetruefalse skips S3 replication, when you keep the replica in sync yourself.

The default status codes include 403 and 404 because a bucket behind origin access control answers a missing object with 403. With them, an object the primary has lost is served from the replica; the 5xx codes cover the bucket or region failing.

Requirements

Failover needs a CloudFront distribution in front of the bucket, so the bucket must be a website bucket (website: true), must not be mounted under another site's path, and the project needs dns.domain and SSL (ssl.certificateArn, or ssl.enabled with dns.hostedZoneId). Without them the template generation stops with an error that says which piece is missing.

Any second origin, for a standalone distribution

An entry in infrastructure.cdn takes any second origin, such as a bucket you replicate yourself or another host:

cdn: {
  frontend: {
    origin: 'my-app-frontend.s3.us-east-1.amazonaws.com',
    failoverOrigin: 'my-app-frontend-dr.s3.us-west-2.amazonaws.com',
    failoverStatusCodes: [500, 502, 503, 504], // the default here
    connectionAttempts: 1,
    connectionTimeout: 3,
  },
},

failoverOrigin is a domain, or { domain, originPath, connectionAttempts, connectionTimeout } to tune the secondary as well. An S3 REST endpoint becomes an S3 origin and anything else an HTTPS-only custom origin. Nothing is created or replicated for you on this path.

Failing over faster

By default CloudFront tries an unreachable primary for up to 30 seconds (3 connection attempts of 10 seconds) before it asks the secondary. connectionAttempts: 1 and connectionTimeout: 3 cut that to 3 seconds. The same settings decide how quickly a primary that is not in a group returns a 504, so lower them with care on an origin that is slow to accept connections.

What is checked before anything deploys

ts-cloud validates the generated distribution the way CloudFront would and fails the deploy with a message naming the problem:

  • each origin group has exactly two members, both existing origins of the distribution, and different from each other;
  • failover status codes are among the ones CloudFront accepts: 400, 403, 404, 416, 429, 500, 502, 503 and 504;
  • a cache behavior that targets an origin group allows only GET, HEAD and OPTIONS. CloudFront never fails over a write, so routes that accept POST, PUT, PATCH or DELETE (compute routes such as /api/*) stay on a single origin;
  • connectionAttempts is 1 to 3 and connectionTimeout is 1 to 10, as whole numbers;
  • the replica region differs from the primary's and the replica name is a valid bucket name.

Limits

  • Only GET, HEAD and OPTIONS fail over, and only on a cache miss. OPTIONS fails over only when it is a cached method, which it is on the default behavior Stacks generates.
  • CloudFront asks the primary first on every request, even right after a failover. Failover costs the time it takes the primary to fail, which is what the connection settings bound.
  • The replica outlives the stack. buddy cloud:remove leaves it, with its data, in place. Delete it yourself when you retire the site (ts-cloud exports deleteFailoverReplicaBuckets for this).
  • Replication is asynchronous. An object written a moment before the primary fails may not have reached the replica yet.