This list contains only the countries for which job offers have been published in the selected language (e.g., in the French version, only job offers written in French are displayed, and in the English version, only those in English).
We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success. You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation. In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise.
Job Responsibility:
Implementing a World-Class Observability Platform
Defining and Upholding Reliability
Leading Incident Response
Automating for Scale and Security
Eliminating Toil and Scaling the Team
Requirements:
An active Top Secret clearance
5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus
Proven partner to DevOps/Platform and application teams
A deep understanding of incident response processes
Technical expertise in Infrastructure as Code (Terraform, Ansible)
Technical expertise in Containers and orchestration (Kubernetes)
Technical expertise in CI/CD (GitLab CI/CD, Jenkins, GitHub Actions)
Technical expertise in Scripting (Python, Go, or Bash)
Technical expertise in Cloud (AWS or AWS GovCloud)
Technical expertise in Observability (Grafana stack, ELK stack, or Datadog)
Technical expertise in Networking fundamentals
Nice to have:
Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503)
GitOps practices and toolchains
Security‑minded design for sensitive environments
Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems
Familiarity with on‑prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc)