Site Reliability Engineer
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
As a Site Reliability Engineer at Kong, you will help build and operate the core cloud infrastructure that powers our API and AI platform. You will work with cross-functional teams to maintain high reliability, performance, and fast feature delivery. You will implement automation, observability, and capacity planning to scale our services. This role offers impact through shaping reliability at scale and contributing to blameless incident resolution and continuous improvement.
Responsabilità- Develop and maintain infrastructure as code using tools like Terraform and Ansible
- Build robust monitoring, logging, and alerting to achieve 99.99% uptime
- Investigate and resolve production incidents with blameless post-mortems
- Automate operations to reduce toil and enable self-service for engineers
- Collaborate with developers to embed reliability and scalability into the app lifecycle
- Contribute to capacity planning, DR drills, and security hardening
- Participate in a fair, sustainable on-call rotation
- Experience operating production workloads on AWS, GCP, or Azure
- Proficiency in at least one language (Golang, Python, or Bash)
- Hands-on with Docker and Kubernetes
- Knowledge of Infrastructure as Code (Terraform a plus)
- Familiarity with CI/CD concepts and tools (GitLab CI, Jenkins)
- Understanding of modern observability stacks (Prometheus, Grafana, ELK)
- Collaborative mindset
- Problem-solving under pressure
- Blameless communication and post-mortem participation
- Terraform
- Ansible
- Docker