The Senior DevOps Engineer is the architect of our cloud-native ecosystem. This role focuses on maximizing engineering velocity through sophisticated automation, impeccable system reliability, and advanced Kubernetes orchestration. As a senior member of the technical team, this individual takes full ownership of the infrastructure roadmap, ensuring that the platform is secure, scalable, and optimized for high-performance critical business applications.
What You’ll be Doing
Expert Kubernetes Orchestration & Management
-
Cluster Architecture: Designs, deploys, and maintains production-grade Kubernetes clusters (EKS, GKE, or self-managed) across multiple environments.
-
Workload Optimization: Manages complex scheduling, resource quotas, and horizontal/vertical scaling to ensure cost-efficiency and performance.
-
Networking & Connectivity: Configures and maintains advanced networking components, including Ingress Controllers, Service Meshes (e.g., Istio, Linkerd), and CNI plugins.
-
Storage & Persistence: Implements and manages persistent storage solutions (CSI) for stateful applications, ensuring data integrity and high availability.
-
Upgrades & Lifecycle: Executes seamless cluster upgrades and maintenance with zero downtime, utilizing blue/green or canary deployment strategies.
Infrastructure as Code (IaC) & Automation
-
Foundation as Code: Provisions and manages 100% of the cloud infrastructure through IaC frameworks like Terraform, OpenTofu, or Pulumi.
-
Module Development: Creates reusable, version-controlled infrastructure modules that allow developers to provision resources in a standardized, self-service manner.
-
Configuration Management: Implements automated configuration management to maintain consistency across distributed systems.
Advanced CI/CD & Developer Experience Ìý
-
Pipeline Engineering: Builds and optimizes sophisticated CI/CD pipelines that automate the entire path to production, integrating automated testing, security scanning, and deployment gates.
-
GitOps Implementation: Drives the adoption of GitOps workflows (e.g., ArgoCD, Flux) to ensure the cluster state is always synchronized with the source of truth in Git.
-
Local Development Flow: Optimizes the “inner loop” of development, providing tools and environments that allow engineers to test Kubernetes-native applications locally or in ephemeral cloud environments.
Security, Compliance & Observability Ìý
-
Cluster Hardening: Implements rigorous security protocols, including Role-Based Access Control (RBAC), Network Policies, and Pod Security Standards.
-
Compliance and Governance: Ensures all infrastructure meets applicable regulatory standards through automated audit logging and encryption at rest/in transit.
-
Full-Stack Visibility: Designs and maintains a comprehensive observability stack (Prometheus, Grafana, Jaeger, or similar) to provide deep insights into cluster health and application performance.
-
Incident Response: Serves as a primary technical lead for infrastructure-related incidents, performing deep-dive root cause analysis and implementing automated preventions.
Observability, SRE & Performance Analysis
-
Observability Architecture: Designs and scales enterprise observability solutions, leveraging metrics, logging, and tracing to provide comprehensive visibility into platform health, performance, and reliability. Ìý
-
SLO/SLI Engineering: Partners with product and engineering teams to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs), establishing error budgets that balance velocity with reliability.
-
Proactive Alerting & Self-Healing: Develops sophisticated, non-fatiguing alerting strategies and implements automated remediation scripts to resolve infrastructure anomalies.
-
Cost Observability: Implements FinOps tooling to provide granular visibility into cluster spend, attributing costs to specific teams, projects, or epics.
-
Chaos Engineering: Regularly conducts failure injection experiments to validate system resilience and ensure the infrastructure can handle regional outages or component failures.
What We’re Looking For
- 7+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering.
- CKA (Certified Kubernetes Administrator) is required. Additional CKS or CKAD is highly preferred.
- Deep expertise in at least one major cloud provider (AWS, GCP, or Azure) with specific focus on managed Kubernetes services.
- Strong ability to write production-quality code in Go, Python, or Ruby for automation and custom tooling.
- Advanced understanding of Linux systems, networking protocols (TCP/IP, DNS, TLS) and container runtimes
- Can proactively identifies and addresses technical debt, driving infrastructure initiatives aligned with long-term business objectives.
- Ability to act as a trusted technical mentor, sharing expertise and elevating the organization’s Kubernetes and cloud-native capabilities.
- Ability to communicate complex infrastructure concepts, risks, and constraints clearly to both technical and non-technical stakeholders.
Compensation & Perks
- $150,000 – $180,000 Annual Compensation.
- Annual Bonus.
- Health & Dental Benefits.
- Health & Wellness Spending Accounts.
- Perkopolis Staff Discounts.
- On-site Gym + Plant Fitness Membership Concession.
- Coffee Bar with Barista + Unlimited Snacks.
- Monthly Events, Friday Lunches & Daily Breakfast.
- Professional Growth & Development Opportunities.
Space Ops welcomes applications from people with disabilities. Accommodations are available on request for candidates taking part in all aspects of the selection process.
Ìý
AI Disclosure: As part of our recruitment process, we may use AI-enabled tools to support the screening and assessment of applications. All AI-generated recommendations and outputs are reviewed and validated by our recruitment team.