This list contains only the countries for which job offers have been published in the selected language (e.g., in the French version, only job offers written in French are displayed, and in the English version, only those in English).
We're building a Site Reliability Engineering center in Mexico City, and we're hiring a Senior Manager-level SRE to serve as the technical anchor for the site - defining the reliability vision, driving cross-team execution, and pioneering automation and AI-driven approaches that transform how we operate three payment networks at scale. This is a strategic technical leadership role. You won't manage people directly, but you'll shape how multiple teams work - setting architectural direction for observability, automation, and operational excellence, alert signal reduction, and reliability platform convergence. You'll be the most senior IC engineer in Mexico City, partnering with the Director (people leader) to translate organizational goals into technical roadmaps and ensuring the engineering quality bar stays high as the site scales. You'll operate across the full landscape: batch settlement systems processing every domestic and international credit/debit transaction, real-time observability platforms that must detect failures before customers do, and AI-powered automation that eliminates the toil standing between us and a proactive reliability culture.
Job Responsibility
Define and maintain a 12-18 month technical vision and roadmap for GPN SRE in Mexico City - decompose destination architecture into deliverable steps, sequence investments, and align execution across teams
Drive reliability transformation across settlement, observability, and automation domains - establish SLOs, error budgets, severity frameworks, and operational standards that teams build against
Pioneer AI and agentic automation approaches - design and build AI-driven solutions (using Claude Code, Copilot CLI, and LLM frameworks) for alert classification, runbook generation, automated remediation, and incident analysis
set patterns that other engineers extend
Own the technical strategy for domain-specific knowledge ramp-up: identify which domain expertise requires deep engineering investment vs. documentation, and architect systems that reduce reliance on tribal knowledge
Lead cross-team technical initiatives - drive observability platform convergence, standardize on COF tooling, and eliminate arbitrary uniqueness across towers
Serve as the senior escalation point for complex production incidents - diagnose cascading failures across distributed systems (storage, network, application), drive resolution, and ensure durable fixes land
Architect automation for high-risk operational processes - certificate rotation, compliance artifact generation, settlement cycle validation - ensuring security and reliability are built in from design
Mentor and elevate engineers across teams - conduct design reviews, establish engineering standards, coach on debugging and system thinking, and create an environment where Principal Associates and Managers grow into domain experts
Introduce and advocate for engineering practices that raise the bar - AI engineering, innersourcing, reuse over rebuild, open source contribution, blameless postmortems, and chaos engineering
Influence beyond the CDMX site - partner with US and UK leadership on architectural decisions, represent CDMX engineering in cross-org forums, and shape GPN-wide reliability strategy
Requirements
Professional English fluency
Bachelor's degree
At least 8+ years of experience in SRE, production operations, or reliability engineering
Experience in DevOps Engineering (internship experience does not apply)
8+ years of experience in at least one of the following: Java, Python, Go
At least 6 years of experience with Cloud Native technologies (Amazon Web Services, Microsoft Azure, Google Cloud Platform)
5+ years of experience with container orchestration services including Docker or Kubernetes
Experience with Shell or Bash scripting
At least 5 years of Unix or Linux system administration experience
Nice to have
Experience developing automation solutions using agentic AI tools (Claude Code, Copilot CLI)
Troubleshooting and debugging skills across distributed systems
Familiarity with payments, financial services, or other regulated high-availability domains
Knowledge or experience of Networking concepts (TCP/DNS/TLS)