Senior Site Reliability Engineer

Senior Site Reliability Engineer

Full-Time 50000 - 70000 € / anno (stimato) Smart working (parziale)
iGenius

In sintesi

  • Mansioni: Progetta e implementa sistemi di osservabilità per ottimizzare le prestazioni del supercomputer Colosseum.
  • Azienda: Domyn, leader nello sviluppo di AI responsabile per settori regolamentati.
  • Benefit: Salario competitivo, smart working, budget per formazione e opportunità di stock options.
  • Altre informazioni: Ambiente dinamico con opportunità di crescita professionale e collaborazione internazionale.
  • Perché questo lavoro: Lavora su tecnologie all'avanguardia e contribuisci a progetti innovativi nel campo dell'AI.
  • Qualifiche: Esperienza come Site Reliability Engineer e competenze in sistemi di monitoraggio.

La retribuzione prevista è compresa tra 50000 - 70000 € per anno.

We are looking for an experienced Site Reliability Engineer to join our growing team in Milan and help shape the future of our flagship project, Colosseum, one of Europe’s most powerful AI supercomputers, currently in development.

Designed to run our proprietary AI models at scale, it forms the compute backbone behind the intelligence we deliver to the world’s most demanding industries.

In this role, you will design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems to enable continuous optimization, including key system budgets such as power, cooling and service level, security-level objectives.

You will be responsible for actively guarding and maintaining these operational budgets as part of day-to-day system reliability and performance management.

You will also contribute to operational excellence through blameless post-mortem analysis and structured incident learning, ensuring continuous improvement of system behavior and resilience.

As a part of the team, you will work closely with Platform Engineering in a shared cybersecurity model, where SRE focuses on detection and monitoring, while Platform Engineering ensures the secure design and operation of the underlying infrastructure.

  • What You Have
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • At least 6 years of experience as a Site Reliability Engineer or in similar roles.
  • Strong experience with observability and monitoring systems such as Prometheus, Thanos, Grafana, and Open Telemetry
  • Experience with low-level system instrumentation and performance visibility using technologies such as e BPF
  • Experience with security monitoring and threat detection tools such as Zeek, Wazuh, or equivalent SIEM / security observability platforms
  • Strong experience with containerized and cloud-native environments, particularly Kubernetes
  • Strong software development skills, particularly in Python, with the ability to build automation, integrations, and custom tooling
  • Experience integrating heterogeneous infrastructure systems across multiple vendors, APIs, and evolving tool ecosystems
  • Familiarity with modern infrastructure automation and emerging agent-based frameworks such as MCP / A2A (or equivalent technologies)
  • Exposure to digital twin technologies and simulation platforms such as NVIDIA Omniverse or equivalent
  • Strong ability to design, build, and maintain software-driven infrastructure solutions in complex, large-scale environments

Who You Are

  • A versatile engineer, comfortable operating in complex and fast-paced environments.
  • Driven and fearless, you proactively tackle challenges and overcome obstacles with determination.
  • A systems thinker, capable of understanding the broader architecture and identifying dependencies across platforms and technologies.
  • A collaborative team player who is enthusiastic, curious, and passionate about problem-solving, thriving both independently and within cross-functional teams.
  • An effective communicator with strong interpersonal skills, able to engage with diverse stakeholders and foster collaboration.
  • Fluent in English and eager to contribute in a multicultural and international environment.

Benefits

Perks

  • Learning Friday. If our team members know more, so do we. That’s why we give everyone a training budget that they can spend on books, online courses or other training materials.
  • Smart Working. Trains can be a drag, you can save some commuting time by working from home.
  • Salary is based on experience and topped up with other bonuses.

We offer a competitive salary, as well as an opportunity to receive company equity.

The typical salary for this role ranges between € 50.000 and € 70.000.

As you gain experience and make more significant contributions to the business, your compensation will be reviewed to match your impact.

Additionally, depending on your seniority and your performance, you’ll have the opportunity to receive stock options, with a variable value calculated from your base salary, giving you the chance to directly participate in the company’s success.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry.

It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world’s largest supercomputers.

At the core of Domyn’s product offer is a chip-to-frontend architecture that allows organizations to control the entire AI stack — from hardware to application — ensuring isolation, security, and governance throughout the AI lifecycle.

Its foundational LLMs, Domyn Large and Domyn Small, are designed for advanced reasoning and optimized to understand each business’s specific language, logic, and context.

Provided under an open-enterprise license, these models can be fully transferred and owned by clients.

Once deployed, they enable customizable agents that operate on proprietary data to solve complex, domain-specific problems.

All solutions are managed via a unified platform with native tools for access management, traceability, and security.

Powering it all, Colosseum — a supercomputer in development using NVIDIA Grace Blackwell Superchips — will train next-gen models exceeding 1T parameters.

Domyn partners with Microsoft, NVIDIA, and G42.

Clients include Allianz, Intesa Sanpaolo, and Fincantieri.

Please review our Privacy Policy here https://bit. ly/4tndsz N .

Senior Site Reliability Engineer datore di lavoro: iGenius

Domyn è un datore di lavoro eccezionale, offrendo un ambiente di lavoro stimolante e innovativo a Milano, dove i dipendenti possono contribuire allo sviluppo di Colosseum, uno dei supercomputer AI più potenti d'Europa. Con opportunità di crescita professionale attraverso un budget per la formazione e un modello di lavoro flessibile, i membri del team possono migliorare le proprie competenze mentre partecipano attivamente al successo dell'azienda. La cultura aziendale promuove la collaborazione e l'apprendimento continuo, rendendo Domyn un luogo ideale per chi cerca un impiego significativo e gratificante.

iGenius

Dettagli di contatto:

Team di recruiting di iGenius

Consigli degli esperti StudySmarter🤫

Ecco come pensiamo che potresti ottenere Senior Site Reliability Engineer

Partecipa ai Meetup Locali

Immergiti nella community tech partecipando a meetup locali di ingegneria e sviluppo software. Questo non solo ti permetterà di imparare delle nuove tendenze, ma anche di incontrare potenziali datori di lavoro e colleghi. Non c'è nulla di meglio che un contatto personale per lasciare il segno!

Contribuisci a Progetti Open-Source

Un ottimo modo per farti notare è contribuire a progetti open-source. Questo non solo arricchisce il tuo portfolio, ma ti mette in contatto con professionisti del settore che possono facilmente raccomandarti per posizioni full-time. Inoltre, dimostrerai competenze pratiche che le aziende adorano!

Esplora le Piattaforme di Recruitment Tecnico

Ci sono piattaforme di recruitment specializzate per il settore tech dove le aziende cercano costantemente profili come il tuo. Assicurati di registrarti su siti specifici per ingegneri e sviluppatori per ricevere piuttosto offerte di lavoro pertinenti e opportunità interessanti. Fai in modo che il tuo profilo risalti!

Applica Direttamente a iGenius

Non dimenticare di controllare il sito web di iGenius per eventuali opportunità di lavoro nel settore ingegneristico. Applicare direttamente può aumentare le tue chance di essere notato, specialmente se segui i loro canali social per aggiornamenti su assunzioni e eventi informativi.

Pensiamo che ti servano queste competenze per eccellere come Senior Site Reliability Engineer

Esperienza con sistemi di osservabilità e monitoraggio (Prometheus, Thanos, Grafana, OpenTelemetry)
Strumenti di monitoraggio della sicurezza e rilevamento delle minacce (Zeek, Wazuh, SIEM)
Competenze in ambienti containerizzati e cloud-native (Kubernetes)
Sviluppo software (Python)
Integrazione di sistemi infrastrutturali eterogenei
Automazione dell'infrastruttura e framework agent-based
Tecnologie di digital twin e piattaforme di simulazione (NVIDIA Omniverse)

Alcuni consigli per la tua candidatura 🫡

Mostra i tuoi progetti!:Nel tuo CV, assicurati di includere una sezione dedicata ai progetti di sviluppo software a cui hai lavorato. Questi possono essere progetti universitari, lavori freelance o anche side projects. Se hai un GitHub, non dimenticare di linkarlo: i recruiter adorano vedere il codice e come affronti le sfide.

Competenze tecniche in evidenza:Fai un elenco chiaro delle tue competenze tecniche nel tuo CV, come linguaggi di programmazione, framework e strumenti che conosci. Assicurati che siano rilevanti per il ruolo di Senior Site Reliability Engineer in iGenius. Questo aiuterà a dimostrare che sei il candidato ideale per il lavoro!

Scrivi una lettera di motivazione mirata:Quando scrivi la tua lettera di motivazione, evidenzia perché sei appassionato di ingegneria e sviluppo software. Parla dei tuoi obiettivi professionali e di come pensi di crescere in iGenius. Ricorda, vogliamo vedere il tuo entusiasmo e il tuo desiderio di imparare!

Attenzione ai dettagli:Nell'ambito dell'ingegneria e dello sviluppo software, i dettagli contano. Fai attenzione alla formattazione del tuo CV e della lettera di motivazione. Un CV ben strutturato e privo di errori mostra che sei meticoloso e professionale, qualità fondamentali per un ruolo a tempo pieno come Senior Site Reliability Engineer.

Come prepararti a un colloquio di lavoro presso iGenius

Preparati con le tue abilità tecniche

Per un colloquio in ingegneria e sviluppo software, è fondamentale essere in grado di dimostrare le tue abilità tecniche. Preparati a rispondere a domande di programmazione e a risolvere problemi dal vivo. Puoi anche praticare con piattaforme come LeetCode o HackerRank per affrontare esempi di codice che potresti incontrare.

Mostra il tuo portfolio di progetti

Essendo un candidato full-time, è importante avere un portfolio ben curato che mostri il tuo lavoro. Porta con te esempi di progetti passati, sia personali che professionali. Spiega il tuo ruolo in ciascun progetto e i risultati ottenuti. Questo non solo dimostra le tue capacità, ma anche la tua passione per il settore.

Preparati a domande sul lavoro di squadra

Nel campo dell'ingegneria e sviluppo software, il lavoro di squadra è cruciale. Aspettati di ricevere domande su come hai collaborato in precedenti progetti o come affronti i conflitti con i membri del team. Pratica le tue risposte usando esempi concreti che mettano in luce le tue abilità relazionali.

Conosci il processo di sviluppo che usano

Ogni azienda ha il proprio modo di fare le cose. Prima del colloquio in iGenius, informati sul loro processo di sviluppo software, come Agile o Scrum. Essere in grado di discutere come ti adatteresti a questi metodi o come hai già lavorato con essi può farti risaltare come candidato ideale.