$ igor.silva

Senior Platform / SRE Engineer

Igor Silva

Designing and operating Kubernetes platforms at the K8s + AI infrastructure intersection. AWS, GitOps, observability.

Madrid, Spain · Remote-ready
01 /

About

I'm a Platform / SRE Engineer with 18+ years in software engineering and about five years focused on DevOps and SRE in cloud-native environments. My focus is Kubernetes (EKS), AWS, GitOps and observability, and the platforms I build are ones engineering teams can rely on without thinking about them.

Most recently at Cvent, I played a key role in bringing Splash's infrastructure into Cvent's broader platform after the acquisition, consolidating AWS accounts and services over roughly six months with zero critical incidents, while running EKS clusters and Datadog observability across both environments.

Before that, at Splash and Ryanair Labs, I designed and ran EKS platforms, drove GitOps with Argo CD that cut out manual deployment steps and out-of-band changes, and kept tightening CI/CD and observability so teams caught problems sooner. Earlier in my career I built payment integrations for Ryanair's customer-facing websites across European markets.

These days I'm going deeper at the Kubernetes + AI infrastructure intersection: the Kubestronaut path, AWS AI certifications, and a public portfolio of Civo-based clusters running Kubernetes-native AI agents.

02 /

Expertise

Cloud & Kubernetes

Designing and operating production EKS platforms on AWS that engineering teams across several products build on every day.

Platforms

AWSKubernetesAmazon EKSAmazon ECS (Fargate)Docker

IaC & GitOps

TerraformAWS CloudFormationArgo CDHelmAnsible

CI/CD & Observability

Reliable pipelines and the telemetry to know when they break, so deploys get faster and problems surface sooner.

Pipelines

GitHub ActionsJenkinsAWS CodePipeline

Observability

DatadogGrafanaDynatraceNew RelicPrometheus

SRE practices

SLIs/SLOsIncident responseOn-callBlameless reviews

Languages & Engineering

18+ years across application and platform layers. Honest practitioner of the languages the work demands.

Languages

PythonBashC# / .NETJavaScriptSQL

Practices

MicroservicesDDDTDDCI/CDAgile (Scrum / Kanban / XP)

Currently going deeper: Kubestronaut path, AWS AI certs, kagent / MCP / Anthropic SDK

03 /

Experience

  1. Sep 2024 to Present

    Senior Site Reliability Engineer

    Cvent · Spain (Remote)

    Global event-technology platform (~4.8K employees); acquired Splash in 2024.

    • Played a key role in bringing Splash's infrastructure and workflows into Cvent's broader platform, consolidating AWS accounts and services over roughly six months with zero critical incidents.
    • Operated AWS infrastructure (EKS, RDS, EC2, S3, VPC) across both the Splash and Cvent environments after the acquisition, running several Kubernetes clusters and RDS instances and supporting engineering teams on both sides.
    • Drove monitoring, alerting and observability improvements with Datadog, which brought the mean time to detection down across the platform.
    • Improved the CI/CD pipelines across both environments, cutting deploy time and making releases more reliable.
    • Coordinated incident response and on-call rotations, and ran blameless post-incident reviews so that follow-up actions actually got done.
  2. Apr 2023 to Sep 2024

    Senior DevOps Engineer

    Splash (splashthat.com) · Spain (Remote)

    Event marketing platform used by global brands (acquired by Cvent in 2024).

    • Built and operated the AWS infrastructure (EKS, RDS) behind a global event-marketing platform.
    • Introduced GitOps with Argo CD to keep infrastructure consistent, which cut manual deployment steps and reduced out-of-band changes.
    • Used Grafana and Dynatrace to watch cloud resources and applications across metrics, logs and traces, improving coverage and how quickly issues were caught.
    • Automated operational tasks and tuned CI/CD pipelines, reducing deploy time and taking recurring toil off the platform team.
  3. Sep 2020 to Apr 2023

    Senior DevOps Engineer

    Ryanair Labs · Madrid, Spain

    Internal tech arm of Ryanair, Europe's largest low-cost airline.

    • Designed and built EKS clusters on AWS for Ryanair's airline operations and customer-facing platforms.
    • Built and maintained CI/CD pipelines across Jenkins, AWS CodePipeline and Argo CD, bringing release time down across the services the team supported.
    • Ran ECS (Fargate) infrastructure for containerised workloads alongside the EKS platform.
    • Moved provisioning to infrastructure as code with AWS CloudFormation and Terraform, replacing manual setup and improving the developer experience.
    • Automated operational tasks with Ansible and Bash, cutting recurring manual toil across infrastructure operations.
    • Monitored cloud applications and services with New Relic, defined alerting rules, and took part in incident response.
  4. Sep 2017 to Sep 2020

    Senior Software Engineer

    Ryanair Labs · Madrid, Spain

    • Built payment solutions and integrations with external providers (PayPal, iDEAL, SEPA and others) for Ryanair's customer-facing websites across European markets.
    • Implemented PSD2-compliant strong customer authentication across Ryanair's payment journeys.
    • Engineered microservices and message-based systems (.NET Core, NServiceBus) on Microsoft Azure and AWS, applying DDD, TDD and CI/CD.
    • Worked with distributed teams across three countries using Scrum, XP and Kanban.
  5. Mar 2013 to Aug 2017

    Software Developer

    Siteware · Belo Horizonte, Brazil

    • Built strategy and corporate-performance applications with ASP.NET MVC, C#, NHibernate and JavaScript, shipping continuously under Scrum, XP and Kanban.
    • Worked with MS SQL Server and Oracle in multi-tenant product scenarios.
  6. Jul 2007 to Mar 2013

    Earlier experience · Brazil

    Software developer roles across MM Informática, Assedef (freelance) and the Audit Office of Minas Gerais, building financial and public-sector applications with .NET, C#, Java/Spring, Delphi, MS SQL Server and Oracle.

04 /

Portfolio

A production-grade cloud-native platform built in the open, at the Kubernetes + AI infrastructure intersection. Terraform provisions a Civo K3s cluster; Argo CD drives everything else through GitOps; the agent layer runs securely on top. The diagram below is the system as it stands.

Explore the portfolio wiki →
tf-modulesTerraform · OpenTofuk8s-gitopsArgo CD manifeststerraform applygit sync◇ Civo · K3s ClusterArgo CDapp-of-apps · GitOpssyncns: platformIstioPrometheusGrafanans: agentskagentAnthropic SDKMCPns: sandboxcrashdummy (Go)gVisorexperiments
component provisioningGitOps sync

Build phases

Each phase, decision, and repo is documented in the portfolio wiki.

05 /

Projects

A public portfolio at the Kubernetes + AI infrastructure intersection, built in the open across several phases, from Terraform foundations to a kagent-managed AI agent running on a Civo K3s cluster.

tf-modules

In progress

Phase 0 · Terraform Foundation

Reusable, opinionated Terraform module library. Provider-agnostic where possible, with clean documentation, pre-commit hooks and CI on every PR. Foundation for the Civo platform work.

TerraformOpenTofuGitHub Actionspre-commit
View on GitHub →

civo-infrastructure

In progress

Phase 1 · Platform Bootstrap

Live K3s cluster on Civo, provisioned via tf-modules. Argo CD-driven GitOps, namespace strategy, service mesh and observability stack as the platform that everything else runs on.

CivoK3sTerraformArgo CDHelmPrometheusGrafana
View on GitHub →

k8s-ai-agent

Planned

Phase 3 · K8s + AI intersection

Kubernetes-native AI agent. Starts as a Python Anthropic SDK deployment with MCP tools for PromQL and kubectl-style operations, then migrates to kagent (CNCF Sandbox) as a custom resource managed by Argo CD.

PythonAnthropic SDKMCPkagentArgo CD
Repo coming soon

Phases 2 (service mesh + observability) and 4 (kagent migration) run between the current milestones. The full roadmap lives in the portfolio wiki.

06 /

Contact

Working on something interesting in Kubernetes, platform engineering or AI infrastructure? Hiring for a Staff-track role at a globally distributed, engineering-led product company? I'm always up for a focused chat, so drop me a line.

Madrid, Spain (GMT+1) · Remote-ready · Open to full-time and contract