Senior CloudOps Engineer
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
Senior CloudOps Engineer
This Full time on site position offers great opportunities for career growth.
We run a genuinely hybrid platform, combining AWS cloud environments with our own on-premise infrastructure and custom GPU clusters built for AI training and inference.
In this role, the core focus will be driving Cloud Governance, scaling our Kubernetes footprint (both cloud and on-prem), and enabling our AI development through modern MLOps tooling.
You will own the infrastructure lifecycle end-to-end, acting as an enabler for our engineering teams by embedding modern Infrastructure-as-Code (IaC) and cost-optimization practices across the entire stack.
About Translated
Translated is a leading provider of AI‑powered language solutions. Founded in 1999 by a linguist and a computer scientist, we are on a mission to allow everyone to understand and be understood in their own language. We envision a world where people from different cultures can communicate seamlessly, gaining unprecedented access to knowledge, cultural exchange, and opportunity. To make this possible, we combine proprietary translation AI (Lara Translate) with advanced text, audio, and video translation technologies (Matecat, Matesub, and Matedub) and the world’s largest network of vetted, native‑speaking language professionals. We welcome complex technical challenges from our customers and engineer tailored solutions that often become part of our core products, whether integrated into our enterprise localization platform, the tools designed to support translators, or LaraTranslate, our online AI translator for teams and individuals. Everything we build reflects a simple principle: technology should amplify human potential, not replace it. We believe in humans.
Responsibilities
- Cloud Governance: Lead cloud architecture governance on AWS. Own cost optimization, capacity planning, access management, and ensure 100% of infrastructure is declared via clean, modular Infrastructure-as-Code (Terraform / Ansible).
- Hybrid Kubernetes Architecture: Design, build, and operate enterprise-grade Kubernetes clusters across both cloud and on‑premise environments. Standardize deployments using GitOps (Argo CD), Helm, and modern networking (Cilium).
- MLOps Infrastructure: Build and maintain high-performance infrastructure pipelines for AI model deployment and inference. Optimize GPU usage and monitoring for AI/ML teams.
- Team Enabling & Developer Experience: Champion the "you build it, you run it" culture. Provide guidance, templates, and self‑service tooling to help dev teams deploy, monitor, and scale their services safely and autonomously.
- Infrastructure Operations (Nice‑to‑Have Support): Continuously optimize platform reliability, observability (Datadog/ELK/CloudWatch), CI/CD pipelines (GitHub Actions), and hybrid connectivity alongside our core infrastructure team.
Requirements
- 5+ years of hands‑on experience in a CloudOps / DevOps / SRE role managing production infrastructure.
- Cloud Governance & IaC: Strong background in AWS infrastructure management, cost optimization, and deep expertise with Terraform and Ansible.
- Kubernetes Expertise: Proven production experience building and operating Kubernetes clusters in both cloud (AWSEKS) and bare‑metal/on‑prem environments. Experience with Argo CD (GitOps) and Helm.
- MLOps Foundations: Experience serving, scheduling, and scaling machine learning inference/training workloads (NVIDIA GPU drivers, device plugins, or MLOps frameworks).
- Engineering Mindset: Strong scripting skills (Python, Bash, or Go) to build automation, CLI tools, and platform integrations.
- Team Enablement: Excellent communication skills with a passion for mentoring developers, driving best practices, and improving developer experience.
- Nice to Have –Production experience with bare‑metal virtualization (Proxmox VE, Ceph).
- Advanced Kubernetes networking using Cilium or eBPF.
- Experience with Docker Swarm to Kubernetes migrations.
- Relational database administration (MySQL, replication, tuning).
- Hands‑on networking skills (VLANs, BGP/routing, high‑availability firewalls).
Headquarters
Translated is hosted at Pi Campus, a working environment immersed in nature where 5 luxury villas in Rome (Italy) have been converted into functional offices to foster talent growth. Pi Campus is also a venture firm created by Translated to reinvest part of its profits into promising AI startups.
Our Offer
Based on a standard level of experience, the offered salary typically ranges between €38,000.00 and €58,000.00. The actual package depends on factors such as location, scope of responsibilities, and performance. Compensation generally grows quickly as experience leads to greater co