Site Reliability Engineer Job at Cohere (Toronto)

Site Reliability Engineer

Cohere

Location:
Canada; United States; United Kingdom , Toronto ▼
Montreal
San Francisco
New York
London

Category:
IT - Software Development

Contract Type:
Not provided

Salary:

Not provided

Save Job

Apply Position

Job Description:

We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.

Job Responsibility:

Build self-service systems that automate managing, deploying and operating services
This includes our custom Kubernetes operators that support language model deployments
Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems
Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation
Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback
Develop our team through knowledge sharing and an active review process

Requirements:

5+ years of engineering experience running production infrastructure at a large scale
Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
Experience with Kubernetes dev and production coding and support
Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments
Experience in compute/storage/network resource and cost management
Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork
The grit and adaptability to solve complex technical challenges that evolve day to day
Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference
Strong understanding or working experience with distributed systems
Experience in Golang, C++ or other languages designed for high-performance scalable servers

What we offer:

An open and inclusive culture and work environment
Work closely with a team on the cutting edge of AI research
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits, including a separate budget to take care of your mental health
100% Parental Leave top-up for up to 6 months
Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement
Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend
6 weeks of vacation (30 working days!)

Additional Information:

Job Posted:
February 20, 2026

Employment Type:

Fulltime

Work Type:

Remote work

Cohere - All Job Offers

Job Link Share:

Site Reliability Engineer

Cohere

Location:
Canada; United States; United Kingdom , Toronto ▼
Montreal
San Francisco
New York
London

Category:
IT - Software Development

Contract Type:
Not provided

Salary:

Job Description:

Job Responsibility:

Requirements:

Additional Information:

Job Posted:
February 20, 2026

Looking for more opportunities? Search for other job offers that match your skills and interests.

Similar Jobs for Site Reliability Engineer

Senior Site Reliability Engineer

Senior Site Reliability Engineer

Senior Site Reliability Engineer

Site Reliability Engineering Manager

Cloud Security Site Reliability Engineer

Senior Software Engineer, Site Reliability

Principal Site Reliability Engineer

Staff Site Reliability Engineer

Site Reliability Engineer

Cohere

Location:Canada; United States; United Kingdom , Toronto ▼MontrealSan FranciscoNew YorkLondon

Category:IT - Software Development

Contract Type:Not provided

Salary:

Job Description:

Job Responsibility:

Requirements:

Additional Information:

Job Posted:February 20, 2026

Looking for more opportunities? Search for other job offers that match your skills and interests.

Similar Jobs for Site Reliability Engineer

Senior Site Reliability Engineer

Senior Site Reliability Engineer

Senior Site Reliability Engineer

Site Reliability Engineering Manager

Cloud Security Site Reliability Engineer

Senior Software Engineer, Site Reliability

Principal Site Reliability Engineer

Staff Site Reliability Engineer

Location:
Canada; United States; United Kingdom , Toronto ▼
Montreal
San Francisco
New York
London

Category:
IT - Software Development

Contract Type:
Not provided

Job Posted:
February 20, 2026