Site Reliability Engineer (Cloud & AI Platforms)
Job Description:
About us
Komodo Consulting is a technology and strategy firm specializing in Digital Transformation. Operating in Portugal and Poland, we provide IT Consulting & Nearshore services. We support both public and private sector organizations through two main areas:
Consulting — with a focus on strategy, investment analysis, and digital process improvement;
IT Team Augmentation — helping clients scale and strengthen their tech teams.
The project
We are seeking a Site Reliability Engineer (Cloud & AI Platforms) to work on a project for a Technology Company.
You will have the following responsibilities:
- Ensure the reliability, availability, and operational health of cloud and data platforms within Azure environments;
- Define and implement SLAs, SLOs, operational models, and disaster recovery strategies;
- Manage incidents, alerting systems, and continuously improve operational processes;
- Implement and maintain observability solutions using Azure Monitor, Application Insights, and telemetry-driven insights;
- Monitor platform performance, availability, and operational consumption, including cost efficiency;
- Design and implement automation aligned with existing architecture, promoting engineering over manual operations;
- Collaborate closely with data, architecture, and development teams to support cloud-native platforms and workloads;
- Provide operational support for data and AI platforms, including Microsoft Fabric and Azure Machine Learning.
You need to have the following skills/experience:
- Strong experience as an SRE or in a similar reliability engineering role with a software engineering mindset
- Solid background in cloud environments, particularly Microsoft Azure;
- Proven experience in automation, observability, and cloud platform operations;
- Experience implementing monitoring, alerting, and disaster recovery strategies from scratch;
- Familiarity with Azure Monitor, Application Insights, and telemetry-based observability practices;
- Experience working with data platforms and/or AI/ML environments (e.g., Azure Machine Learning, Microsoft Fabric);
- Ability to work cross-functionally with architecture, data, and development teams;
- Hands-on, autonomous profile with the ability to structure and improve operational models;
- Strong technical communication skills.
Location
Hybrid (3 days per week at the office in Lisbon)