Site Reliability Engineer - Vice President
Is this the right opportunity for you?
Explore more jobs, compare advertised salaries and see which skills employers want.
Explore this careerSearch for a different role
More DevOps Engineer jobs →Skills
Description
About the Role
iCapital is looking to hire a Site Reliability Engineer to join the Site Reliability Engineering team, which plays a critical role in ensuring the platform delivers consistent, reliable service to clients. The services run on Kubernetes in AWS, and this role is focused on maintaining the health, stability, and performance of production environments.
This position is ideal for a hands-on engineer who thrives in diagnosing complex production issues where the root cause is not immediately apparent. This role transforms insights gained from production incidents into stronger reliability standards, safer deployment practices, more effective alerting, and automation that helps prevent recurring issues. While metrics, logs, and traces are leveraged daily, the core focus is identifying, resolving, and learning from the issues they reveal.
Responsibilities
- Act as a senior escalation point for complex production issues in Kubernetes-based services, leading hands-on investigation across application, container, Kubernetes, and AWS layers through to root cause.
- Partner with the teams that own our cluster infrastructure to resolve platform-level issues, and feed what you learn back into reliability standards and improvements.
- Define and implement reliability and operability standards for Kubernetes-based services, including resource sizing, health checks, scaling patterns, rollout safety, and baseline monitoring, as part of service onboarding.
- Serve as the incident commander or senior technical lead for high-severity incidents, lead postmortems, and drive systemic improvements through action items and measurable follow-through.
- Turn recurring investigations into runbooks, diagnostic tooling, and automated remediation that shorten time to resolution and eliminate toil.
- Coach engineers in systematic troubleshooting and Kubernetes best practices through reviews, game days, and knowledge sharing.
- Define, implement, and iterate service level objectives (SLOs) and service level indicators (SLIs) that reflect customer and business expectations.
- Standardize monitoring and alerting through “monitors as code” (Terraform preferred), including quality gates such as severity, ownership, and runbook links.
- Operate and tune OpenTelemetry pipelines, including sanitizing data, managing cardinality, and enriching telemetry with consistent fields.
- Maintain the infrastructure behind our monitoring stack, including Helm charts, ArgoCD ApplicationSets, Kinesis streams, Lambda functions, and Prometheus.
- Participate in on-call rotations with a focus on improving reliability, reducing alert noise, and increasing signal quality over time.
Qualifications
- 5+ years in SRE or related roles, with evidence of technical seniority across multiple services and teams
- 3+ years of hands-on experience running and troubleshooting production workloads on Kubernetes, with a deep understanding of how Kubernetes schedules, connects, scales, and rolls out workloads
- Proven ability to independently diagnose complex failures in distributed systems running on Kubernetes, spanning application, configuration, resource, networking, and scaling issues, including cases where the root cause wasn't obvious
- Strong Linux and networking fundamentals, with the ability to troubleshoot below the container layer
- Strong production experience with AWS, including how cloud networking and access controls affect workloads running in Kubernetes
- Strong incident response skills, including leading retrospectives and postmortems and following through on systemic fixes
- Strong IaC skills (Terraform preferred), with examples of reusable modules, monitors-as-code patterns, or similar standardization work, and the coding ability to build automation and diagnostic tooling
- Track record of defining SLOs/SLIs and using an observability stack (i.e. Prometheus and Grafana, New Relic, Splunk, or similar) to drive decisions and find root causes quickly
- Clear written and verbal communication skills with the ability to influence engineering teams through standards, tooling, and practical guidance
- Kubernetes certification (CKA or CKS) or equivalent depth
- Experience improving rollout safety and autoscaling behavior for Kubernetes workloads
- Familiar with common data stores (Postgres, MongoDB, DynamoDB) and how they fail in distributed systems
- OpenTelemetry instrumentation experience at scale is preferred
Benefits
The base salary range for this role is $130,000 to $160,000 depending on level. iCapital offers a compensation package which includes salary, equity for all full-time employees, and an annual performance bonus. Employees also receive a comprehensive benefits package that includes an employer matched retirement plan, generously subsidized healthcare with 100% employer paid dental, vision, telemedicine, and virtual mental health counseling, parental leave, and unlimited paid time off (PTO).
We believe the best ideas and innovation happen when we are together. Employees in this role will work in the office Monday-Thursday, with the flexibility to work remotely on Friday.
For additional information on iCapital, please visit Twitter: @icapitalnetwork | LinkedIn: | Awards Disclaimer: https://www.icapitalnetwork.com/about-us/recognition/
iCapital is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, gender, sexual orientation, gender identity, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.
Get similar jobs in United States by email
We'll email you when new jobs similar to this one appear.
Similar jobs
Explore more DevOps Engineer jobs in United States.
No similar openings right now. Refine your search or create an alert above.