SRE Tech Lead
Transform the future of cloud industry together with Scalr! ๐
Company Overview
Scalr is a SaaS product company that offers everything necessary to scale Terraform. We place a strong emphasis on Terraform / OpenTofu, DevOps, GitOps, and the "everything as code" philosophy, prioritizing consistency and simplicity. Scalr builds a management layer atop Terraform / OpenTofu, which helps DevOps teams scale across their entire organization. As an engineering organization, we also embrace a DevOps approach, researching cloud services, adopting best practices, and utilizing Terraform / OpenTofu throughout. This enables us to better understand our customers' challenges and use cases.
As we expand our offerings, we are seeking an SRE Tech Lead with a passion for pushing the boundaries of technology to solve complex problems.
Position Overview
As an SRE Tech Lead, you will own the Scalr platform's reliability, scalability, and observability - proactively identifying and eliminating risks before they become incidents and impact customers - while designing new architecture components, promoting and enforcing SRE and DevOps best practices and driving strategic technical improvements.
The main infrastructure technology stack includes GCP, GitHub (including GA for CI/CD), Terraform, Datadog, Sentry, and Grafana.
At Scalr, we believe that the best software is produced when engineers take pride and ownership in the work they accomplish. Consequently, engineers are expected to provide customer support. We value troubleshooting skills and customer empathy because, ultimately, writing good code and helping customers succeed lay the foundation for building great companies.
Qualifications:
๐ธ Python (experience in Python scripting is enough)
๐ธ Terraform/OpenTofu
๐ธ Strong knowledge of Linux (RHEL/Debian, bash scripting)
๐ธ Docker
๐ธ Kubernetes
๐ธ Google Cloud Platform
๐ธ Leading SRE teams or initiatives
๐ธ Experience with monitoring and logging tools such as Grafana, Prometheus, Datadog, New Relic, etc.
๐ธ Experience with CI platforms such as GitHub Actions, Drone, CircleCI, etc.
๐ธ Strong written and verbal communication skills
Would be a plus:
๐ธ Experience with GitOps, Argo CD, Flux CD or similar
๐ธ Chef, Omnibus, Ruby
๐ธ JavaScript for GitHub Actions
As part of our team, you will work on:
๐ธ Create and maintain the SRE roadmap and backlog, keeping them continuously up to date from an infrastructure perspective (observability, performance, stability, reliability) and advocate for their prioritization based on business and product necessity
๐ธ Own the observability strategy end-to-end (monitoring, alerting, dashboards): ensure monitoring coverage of all critical paths and keep alerting actionable.
๐ธ Architect and evolve the platform infrastructure, ensuring its scalability, technical efficiency and resilience - systematically eliminating single points of failure across critical components, including disaster recovery and failover strategy.
๐ธ Define and own customer-centric SLIs, SLOs and error budgets for application and infrastructure reliability.
๐ธ Manage the infrastructure technology stack, including evaluating and integrating new CI/CD and monitoring tools, libraries etc.
๐ธ Cooperate with other Tech Leads and coordinate interaction with other departments.
๐ธ Lead the resolution of technical challenges, serve as an arbitrator in decision-making and take a hands-on role in system maintenance - staying directly involved in the SRE process and guiding the team toward high-quality software development.
๐ธ Foster a culture of ownership, autonomy, and continuous growth within the team.
๐ธ Teach, mentor, and ensure employees have the appropriate level of knowledge to work with different system components; initiate mentoring, Q&A sessions, and knowledge reviews.
What does Scalr offer:
๐ Work with an exciting engineering product in an enjoyable environment
๐ค The opportunity to see how your ideas and visions are realized
๐ฐ Attractive compensation and benefits package
๐
Long-term contract and tax compensations
๐ Flexible schedule and possibility to work entirely remotely
๐ฉบ Medical insurance
๐๏ธ 20 working days of paid vacation and 2 weeks of paid sick leaves
- Department
- Engineering
- Locations
- Lviv
- Remote status
- Fully Remote
About Scalr
Scalr is a remote operations backend for Terraform and OpenTofu. Scalr executes runs and stores state centrally allowing for easy collaboration across your organization. You can continue to use existing workflows that use the native Terraform or OpenTofu CLI, implement a GitOps workflow, or use No Code provisioning.