Description
Summary:
We are seeking a Senior DevOps Engineer to own the reliability, scalability, and performance of our infrastructure and CI/CD ecosystem, driving standards and improving delivery.
Highlights:
1. Globally distributed team with strong engineering culture
2. Pivotal role shaping systems, deployments, monitoring, and evolution
3. Fully remote setup with high autonomy and ownership
At Crunch, we’re not just building tech — we’re building a place where talented people grow, create, and feel like they belong. We’re a globally distributed team with a strong engineering culture, straightforward communication, and a genuine passion for what we do. You won’t find rigid hierarchies or endless bureaucracy here — just curious minds solving meaningful problems together.
We’re looking for a **Senior DevOps Engineer** who will own the reliability, scalability, and performance of our infrastructure and CI/CD ecosystem. You’ll play a critical role in shaping how our systems are built, deployed, monitored, and evolved — enabling engineering teams to ship confidently and efficiently.
This role goes far beyond maintaining pipelines. You’ll drive infrastructure standards, improve delivery performance, and lead initiatives around system stability, observability, and web performance.
This is a pivotal role for a technical leader who thrives at the intersection of **cloud infrastructure, CI/CD automation, system reliability, and cross\-team collaboration**.
**Responsibilities**
* Own the **design, provisioning, and lifecycle management** of production infrastructure across development, staging, and production environments
* Architect, maintain, and scale **CI/CD platforms and deployment automation** used by engineering and QA teams
* Ensure **system reliability, scalability, and availability**, including capacity planning, autoscaling, and failure recovery strategies
* Establish and enforce **infrastructure and deployment standards**, including blue/green and canary deployments, rollback mechanisms, and environment parity
* Lead **incident response, root cause analysis, and postmortems**, driving long\-term reliability improvements
* Design and maintain **observability systems** (monitoring, logging, alerting, metrics) aligned with SLI/SLO best practices
* Drive **infrastructure\-level performance optimization**, including improvements impacting Core Web Vitals and backend latency
* Optimize **cloud cost efficiency and resource utilization**, providing data\-driven recommendations and PoCs
**Requirements**
* 5\+ years of hands\-on experience in **DevOps, SRE, or Platform Engineering roles**
* Strong expertise with **Git** and modern version control workflows
* Proven experience managing **complex CI/CD pipelines** in multi\-repository environments
* Deep practical experience with **AWS**, including serverless architectures and managed services
* Solid understanding of **infrastructure as code** (Terraform, CloudFormation, or similar tools)
* Experience implementing and maintaining **monitoring, logging, and alerting systems**
* Strong understanding of **web performance fundamentals** and how infrastructure impacts Core Web Vitals
* Experience working in **cross\-functional engineering environments** and supporting multiple teams
* Strong analytical mindset with the ability to translate metrics into actionable infrastructure improvements
**Nice to Have**
* Experience with **Node.js, React, and modern web application stacks**
* Familiarity with **containerization and orchestration** (Docker, Kubernetes, ECS)
* Experience with **SonarQube**, quality gates, or CI\-level code quality tooling
* Background in **performance optimization, cost optimization, or high\-traffic systems**
* Experience supporting **visual testing or test automation infrastructure**
**Work Environment \& Schedule**
* Working hours aligned with **EST (Eastern Standard Time)** for effective collaboration with the core team
* High\-impact role with direct influence on **infrastructure scalability, system reliability, and developer productivity**
* Fully remote setup with a high level of autonomy and ownership
**For Team Members in Latin America**
* 18 paid leave days per year (available after 6 months)
* 15 unpaid personal days
* 10 paid sick days with a medical certificate
* An extra paid day off for marriage, childbirth, or bereavement
* English language courses
* Educational and Health budgets
Job Type: Full\-time
Application Question(s):
* Have you worked with the full end\-to\-end release lifecycle and on projects with complex release processes?
* What were your specific areas of responsibility during releases?
* How many years of commercial experience do you have with:
\- Release Management
\- AWS
\- SonarQube
\- Quality Gates
\- CI/CD pipelines
\- Git
\- Jenkins
* What is your level of spoken English?
* What are your monthly gross salary expectations in USD?
Work Location: Remote