Senior Site Reliability Engineer (AWS / EKS)
B2B Contract | Fully remote
Role Overview
We are looking for a Senior Site Reliability Engineer with deep, hands-on experience operating highly available production environments on AWS and Amazon EKS.
This is a true SRE position, not a cloud architecture, infrastructure design, or monitoring-focused role. You will take direct ownership of production reliability, participate in the on-call rotation, respond to critical incidents, troubleshoot complex Kubernetes and distributed-system failures, and drive permanent improvements following incidents.
The role also requires strong technical communication. You will interact directly with customers during technical discussions and production escalations, clearly explaining issues, making sound technical decisions, and driving problems through to resolution.
We are looking for someone who has spent significant time running production systems, not simply designing them.
Key Responsibilities:
Own the reliability, availability and operational health of production services running on AWS and Amazon EKS.
Operate and troubleshoot Kubernetes clusters in production, including cluster lifecycle, upgrades, node management, networking, scaling, capacity and workload reliability.
Participate actively in on-call and pager rotations and take ownership of production incidents.
Lead or play a key technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery and communication.
Coordinate technical incident bridges and communicate directly with customers during production escalations when required.
Perform root cause analysis and lead blameless postmortems, ensuring incidents result in concrete engineering improvements.
Define, monitor and improve SLIs, SLOs and error budgets for production services.
Develop and maintain actionable alerts, operational runbooks and automated remediation.
Build and improve infrastructure using Terraform/Terragrunt and Infrastructure as Code practices.
Operate GitOps-based delivery environments using tools such as Argo CD or FluxCD.
Improve Kubernetes scaling and efficiency using technologies such as Karpenter, KEDA and native Kubernetes autoscaling capabilities.
Build and improve observability using technologies such as Prometheus, Grafana, OpenTelemetry, Datadog and/or ELK.
Support highly available distributed and event-driven systems, including environments using technologies such as Kafka/MSK.
Design, implement and validate disaster recovery and business continuity mechanisms against measurable RTO and RPO objectives.
Identify recurring operational problems and eliminate toil through automation and engineering.
Improve AWS performance, scalability, security and cost efficiency across production environments.
Work closely with software, platform and engineering teams to build reliability into systems throughout the development lifecycle.
Contribute to continuous improvement of incident management, operational readiness and SRE engineering practices.
Requirements:
Significant professional experience as a hands-on Site Reliability Engineer, Production Engineer or senior Platform Engineer with direct production ownership.
Several years of recent, hands-on experience operating production environments on AWS.
Strong, demonstrable experience operating Amazon EKS in production.
Deep Kubernetes operational knowledge beyond application deployment, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting and production failure scenarios.
Proven participation in a production on-call/pager rotation.
Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, RCA and post-incident improvements.
Practical experience with SLIs, SLOs, error budgets, alerting and runbooks.
Strong Infrastructure as Code experience with Terraform and/or Terragrunt.
Production experience with Kubernetes delivery and GitOps practices; Argo CD or FluxCD strongly preferred.
Strong production observability experience with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog or ELK.
Experience operating highly available, distributed production systems.
Strong understanding of AWS networking, IAM, security, availability and resilience.
Experience implementing and testing disaster recovery strategies with measurable RTO/RPO objectives.
Proven external customer-facing technical experience, including technical discussions, production escalations, architecture/reliability conversations or incident communication.
Ability to explain complex technical problems clearly and make sound decisions during high-pressure production incidents.
Strong troubleshooting mindset and ability to work independently during complex production failures.
Strong professional English communication skills (min. C1) for regular interaction with clients,
A coherent track record demonstrating sustained hands-on production engineering ownership.
What's on Offer:
Full-time permanent B2B cooperation.
Fully remote working environment.
Senior hands-on engineering position with meaningful ownership of business-critical production systems.
Opportunity to work on complex AWS, Kubernetes and distributed-system environments at scale.
Direct influence over reliability engineering, operational practices and platform improvements.
Modern engineering environment with strong emphasis on automation, observability and continuous improvement.
Collaboration with experienced engineering, platform and product teams.
Opportunity to introduce and use modern approaches, including AI-assisted engineering and operational automation.
Long-term opportunity for engineers who want to remain deeply technical and close to production.
Diversity and Inclusion Commitment
We are dedicated to creating and sustaining an inclusive, respectful workplace for all -regardless of gender, ethnicity, or background. We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.
- Locations
- Seattle, WA, United States
- Remote status
- Fully Remote
- Employment type
- Full-time
About Salve.Inno Consulting
Bringing a personalized approach to connecting exceptional talent with unique opportunities. Specializing in recruitment for diverse roles, leveraging extensive experience and innovative strategies to find the perfect match for any business needs. Collaboration builds a stronger, more successful future – one strategic hire at a time.