Site Reliability Engineer Job at Chabez Tech (Pflugerville)

Site Reliability Engineer

We are currently seeking a Site Reliability Engineer to join our team in Guadala...

Location

Mexico , Guadalajara

Salary:

Not provided

NTT DATA

Expiration Date

Until further notice

Requirements

Perform L1.5 activities such as monitoring, deployment, rollback
Monitor the efficiency of the Azure cloud systems to prevent outages and initiate an Incident Management bridge in case of an outage
Troubleshoot Azure resources, escalate to Level 3 (Software Development Team)
Understand the Microsoft Azure Cloud - ideally Azure Fundamentals certified OR Computer Science/Information Systems Management degree
Familiar with PaaS and IaaS - VMs, Storage, EventHub, Service Fabric Cluster (SFC), Azure Kubernetes Service (AKS), CosmosDB, SQL Server, IoT Hub, Databricks, KeyVault, Datalake
Understand the concept of Internet of Things (IoT) - telemetry, ingestion, processing, data storage, reporting
Understand the concept tools - Octopus, Bamboo, Terraform, Azure DevOps, Jenkins, Github, Ansible
Understand the concept of container orchestration platforms (e.g. Kubernetes)
Understand the concept of scripts: Powershell, Python
Understand the difference between NoSQL and SQL databases, and how to maintain them

Job Responsibility

Perform L1.5 activities such as monitoring, deployment, rollback
Monitor the efficiency of the Azure cloud systems to prevent outages and initiate an Incident Management bridge in case of an outage
Troubleshoot Azure resources, escalate to Level 3 (Software Development Team)

Fulltime

Site Reliability Engineer

Location

South Africa , Johannesburg

Salary:

Not provided

Nintex

Expiration Date

Until further notice

Requirements

You provide guidance on infrastructure architecture and contribute to high-quality and successful product releases.
You contribute to your team and domain through successfully leading and consistently delivering on projects of ambiguous scope, high complexity, and critical business impact.
You contribute to relevant guilds, practice forums and other initiatives to improve Nintex’s DevOps and SRE discipline.
You have an in-depth understanding of distributed systems architecture, as well as monitoring and observability practices and tools.
You quickly resolve priority infrastructure issues and help other technical team members or Product Managers understand how to avoid them in the future.
You provide detailed estimates for work items you propose or assigned.
You assist in decision-making around tooling, automation practices, and testing solutions.
You stay up-to-date with technology trends and use this knowledge help your team and the broader Engineering practice.
You run Nintex infrastructure with IaC tools (as Terraform) and GitHub Actions for automation, containerize our environments (Kubernetes) and leverage cloud technologies to meet our goals
You build monitoring that alerts on symptoms rather than outages using tools like Prometheus, Grafana, Alertmanager and PagerDuty

Job Responsibility

You are highly skilled and sufficiently experienced in Nintex DevOps tools and processes to own a long-term program or technology such as Kubernetes, etc.
You write scripts, tools and utilities that support and integrate with delivery pipelines and you integrate telemetry where appropriate.
You are called into incidents and bring trusted knowledge in your platform domain.
You debug and fix infrastructure issues on production environments quickly using the relevant tools and guidelines to prevent recurrence.
You build, promote and support infrastructure patterns and practices within Nintex.
You provide coaching/mentoring to other Engineers on the team
You lead or contribute to post-mortems for incidents, including root cause analysis and identification of preventative and remedial actions.
You continuously monitor our platform performance and take immediate action to improve it
You review and advise on appropriate design patterns to solve automation and infrastructure problems without creating technical debt.
You design and build complex infrastructure components for distributed systems as Kubernetes.

What we offer

Global Gratitude and Recharge Days
Flexible, paid time off policy
Employee wellness programs and counseling resources
Meaningful peer recognition and awards
Paid parental leave
Invention/patenting assistance
Community impact, paid volunteer time, and opportunities
Intercultural learning and celebration
Multiple tools through which to learn and grow, and an incredible global community

Site Reliability Engineer

An Elite FinTech Firm is looking for a highly talented DevOps Engineer/Systems S...

Location

Hong Kong , Hong Kong

Salary:

1200000.00 HKD / Year

Hunter Bond

Expiration Date

Until further notice

Requirements

Genuine passion in Linux & Open-source
Excellent knowledge of Python
Use of CI/CD, Docker, Ansible, Chef, Puppet
Knowledge of large-scale storage systems (on-prem)

Job Responsibility

Help architect a resilient, multi-petabyte storage solutions & build new data centres
Automate anything and everything with Python & config tools
Innovate whilst bringing in new ideas

What we offer

Flexible hours/work options
Working in one of the world’s most elite teams
Invest heavily in cutting-edge and next-gen tech
Technologists only report to other technologists
Brand new skyline Manhattan office
Start-up style environment

Fulltime

Site Reliability Engineer

As a Staff Software Engineer, you will play a key role in designing, building, a...

Location

United States , San Jose

Salary:

120500.00 - 243000.00 USD / Year

Hewlett Packard Enterprise

Expiration Date

Until further notice

Requirements

Minimum of 5 years of hands-on experience in Infra Ops, Dev Ops, or Site Reliability Engineering (SRE)
Proficiency with Linux systems, especially Debian-based distributions
Strong experience with cloud platforms such as AWS and GCP
Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible
Solid programming skills in Python and/or Golang
Deep understanding of containerization (Docker, Container) and orchestration tools (AWS EKS, GCP GKE)
Experience with GitOps workflows
Proven track record in implementing and maintaining CI/CD pipelines
Strong background in security and familiarity with security programs
Experience with monitoring and logging tools (Prometheus, Grafana, ELK)

Job Responsibility

Enhance Infrastructure as Code (IAC) and enforce best practices
Optimize cloud infrastructure for scalability, security, and cost-effectiveness
Develop internal tools to support and streamline cloud platform operations
Improve CI/CD pipelines and deployment workflows using FluxCD and Jenkins
Address container image vulnerabilities and standardize remediation processes
Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks
Strengthen monitoring, alerting, and observability using Prometheus, Grafana, and logging tools
Troubleshoot complex production issues to ensure system reliability and customer satisfaction
Fine-tune distributed systems such as Apache Kafka and Cassandra
Collaborate with development, security, and operations teams to align infrastructure with application needs

What we offer

Health & Wellbeing
Personal & Professional Development
Unconditional Inclusion

Fulltime

Site Reliability Engineer

An Elite FinTech Firm is looking for a highly talented DevOps Engineer/Systems S...

Location

United Kingdom , London

Salary:

150000.00 GBP / Year

Hunter Bond

Expiration Date

Until further notice

Requirements

Genuine passion in Linux & Open-source
Excellent knowledge of Python
Use of CI/CD, Docker, Ansible, Chef, Puppet
Knowledge of large-scale storage systems (on-prem)

Job Responsibility

Help architect a resilient, multi-petabyte storage solutions & build new data centres
Automate anything and everything with Python & config tools
Innovate whilst bringing in new ideas

What we offer

Flexible hours/work options
Working in one of the world’s most elite teams
Invest heavily in cutting-edge and next-gen tech
Technologists only report to other technologists
Brand new skyline Manhattan office
Start-up style environment

Fulltime

Site Reliability Engineer

An Elite FinTech Firm is looking for a highly talented DevOps Engineer/Systems S...

Location

Canada , Montreal

Salary:

200000.00 CAD / Year

Hunter Bond

Expiration Date

Until further notice

Requirements

Genuine passion in Linux & Open-source
Excellent knowledge of Python
Use of CI/CD, Docker, Ansible, Chef, Puppet
Knowledge of large-scale storage systems (on-prem)

Job Responsibility

Help architect a resilient, multi-petabyte storage solutions & build new data centres
Automate anything and everything with Python & config tools
Innovate whilst bringing in new ideas

What we offer

Flexible hours/work options
Working in one of the world’s most elite teams
Invest heavily in cutting-edge and next-gen tech
Technologists only report to other technologists
Brand new skyline Manhattan office
Start-up style environment

Fulltime

Site Reliability Engineer

Microsoft Substrate is the foundational cloud platform that powers many of Micro...

Location

United States , Redmond

Salary:

102100.00 - 202200.00 USD / Year

Microsoft Corporation

Expiration Date

Until further notice

Requirements

Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience
4+ years technical experience in software engineering, network engineering, or systems administration
ability to meet Microsoft, customer and/or government security screening requirements
ability to obtain and maintain favorably adjudicated Tier 3 (T3) background investigation
ability to meet Criminal Justice Information Services (CJIS) eligibility requirements
must pass Microsoft Cloud background check upon hire/transfer and every two years thereafter

Job Responsibility

Own reliability and operational health for one or more Substrate components or services in highly regulated environments
Serve as an actively engaged on-call engineer (OCE), participating in an on-call rotation and independently responding to incidents for owned services
Respond to, diagnose, and resolve production incidents with minimal supervision
Design and implement automation to reduce operational toil and improve service stability
Develop and maintain monitoring, alerting, and telemetry to support SLOs and operational metrics
Lead post-incident reviews for owned incidents, focusing on root cause analysis and durable fixes
Collaborate with software engineering teams to embed reliability and operability into service design
Write and maintain production-quality code and automation that improves reliability, scalability, and operational efficiency

What we offer

Benefits and other compensation may be eligible
additional benefits and pay information available at https://careers.microsoft.com/us/en/us-corporate-pay

Fulltime

Site Reliability Engineer

Microsoft is a company where passionate innovators come to collaborate, envision...

Location

India , Bangalore

Salary:

Not provided

Microsoft Corporation

Expiration Date

Until further notice

Requirements

Master's Degree in Computer Science, Information Technology, or related field AND 3+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience
Ability to meet Microsoft, customer and/or government security screening requirements
Must pass Microsoft Cloud background check upon hire/transfer and every two years thereafter

Job Responsibility

Own the end-to-end readiness of Event Stream across Azure regions, including onboarding new regions, driving deployment automation, and ensuring consistent, secure, and compliant service rollout
Work closely with platform, infrastructure, and partner teams (e.g., Event Hubs, Kusto, Fabric platform) to deliver resilient, low-latency streaming experiences on a global scale
Play a key role in advancing our reliability posture, improving availability, monitoring, and incident response across regions
Build strong observability, telemetry, and automated recovery mechanisms to meet high availability and SLA targets
Region Build-out & Deployment: Onboard new regions, drive deployment automation, and ensure consistent service configuration
Reliability & SRE: Improve availability, resiliency, and incident response
own service health across regions
Observability & Operations: Enhance telemetry, monitoring, alerting, and troubleshooting capabilities
Cross-team Collaboration: Partner with platform and infra teams to unblock dependencies and ensure smooth rollout
Production Excellence: Drive root-cause analysis, repair items, and continuous improvement on service reliability

Fulltime

Select Country

Site Reliability Engineer

Job Description

Job Responsibility

Requirements

Nice to have

Looking for more opportunities?

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Site Reliability Engineer

Our AI answers in your language