← Back to jobs

SRE - Database Reliability (All Levels)

Skills

auroraawsbashdata modelingdatadogdynamodbfintechinfrastructure as codenetworkingobservabilityopensearchpci dsspostgresqlpythonserverlesssoc 2terraformtimescaledbPythonPostgreSQLAWSTerraform

Description

Location(s)

This is a hybrid role based out of our San Francisco, CA or New York, NY office. Our team works in-office time on Mondays, Wednesdays, and Fridays.

About the Role

We're hiring an SRE to own our production databases end to end, weighted toward database reliability rather than general infrastructure.

Casap builds dispute-resolution software for banks and credit unions, and we onboard new financial-institution clients every month. Our data layer runs on Aurora PostgreSQL today, with OpenSearch and a TimescaleDB metrics store beside it, and part of it is moving to DynamoDB. You'll make sure that data is backed up, restorable, fast and upgraded on time, and you'll carry it safely through the migration and after it.

Platform engineering already covers infrastructure as code, CI and networking, so you'll spend most of your time on the data layer itself. You'll join the Foundations team and work in a regulated environment (PCI DSS and SOC 2).

What You'll Do

  • Aurora PostgreSQL in production: availability, performance, capacity planning and major-version upgrades.
  • Backups and restores: backup policy, point-in-time recovery, and scheduled restore tests with
  • measured recovery time and data-loss targets.
  • The DynamoDB migration, from the data side: access-pattern and data-model review, backfill
  • and dual-write plans, validation, cutover and rollback.
  • Database access and security: least-privilege roles, credential rotation, audit logging and
  • encryption, in line with PCI DSS and SOC 2.
  • Data-layer observability: slow queries, connection saturation, replication lag and storage
  • growth, with alerting in Datadog.
  • On-call for data-layer incidents: response, follow-through and written post-incident reviews.
  • Runbooks clear enough that the rest of engineering can operate the databases without you.

What You’ve Done

  • You've run PostgreSQL in production and been the person accountable when it broke.
  • You have deep PostgreSQL skills: query planning, indexing, vacuum and bloat, connection pooling, replication and major-version upgrades.
  • You've done real backup and restore work: point-in-time recovery, restore drills and measured recovery times.
  • You know data stores on AWS, including Aurora PostgreSQL (Serverless v2) and DynamoDB data modeling (access patterns, key design, capacity).
  • You're comfortable with Terraform or OpenTofu, CI pipelines and scripting in Python, Go or Bash.
  • You've worked under PCI DSS, SOC 2 or a similar regime, or you're ready to.
  • You write clearly: runbooks, migration plans and incident reviews that other people can act on.

Nice to Have

  • TimescaleDB in production
  • A completed migration from a relational database to DynamoDB
  • Datadog database monitoring
  • Experience in payments, banking, or another fintech setting

Get similar jobs in United States by email

We'll email you when new jobs similar to this one appear.

Similar jobs

Explore more DevOps Engineer jobs in United States.

No similar openings right now. Refine your search or create an alert above.