AI/HPC Systems Performance Engineer Job at Meta (Menlo Park)

Sr AI/HPC Applications and Performance Engineer

Sr AI/HPC Applications and Performance Engineer role at Hewlett Packard Enterpri...

Location

United States

Salary:

161500.00 - 370500.00 USD / Year

Hewlett Packard Enterprise

Expiration Date

Until further notice

Requirements

15+ years' experience
Deep expertise in AI and HPC applications and performance engineering including simulation, modeling and emulation capabilities
Expertise in large-scale AI and HPC systems
Experience architecting, designing, and developing innovative software system design tools and languages
Excellent analytical and problem-solving skills
Experience in leading overall architecture of software systems for products and solutions
Designing and integrating efficient and scalable software systems running on multiple platform types into overall architecture
Evaluating and selecting forms and processes for software systems testing and methodology
History of innovation with multiple patents or deployed solutions in the field of software design
Excellent written and verbal communication skills

Job Responsibility

Develops organization-wide architectures, strategies, and methodologies for software systems design and development across multiple platforms and organizations
Identifies and makes informed recommendations regarding new technologies, innovations, and outsourced development partner relationships
Reviews, evaluates, and influences designs and project activities for compliance with development guidelines and standards
Provides tangible solutions that improve product quality and mitigate failure risk
Contributes to domain expertise, business acumen, and experience to influence decisions of executive business leadership
Brings creativity and innovation to the organization
Provides guidance and mentoring to less-experienced team members
Acts as an internal authority on software systems design
Contributes to the external technical community through whitepapers, patents, or other significant innovations

What we offer

Health & Wellbeing benefits
Personal & Professional Development programs
Unconditional Inclusion environment
Comprehensive benefits suite supporting physical, financial and emotional wellbeing

Fulltime

New

Ai/hpc System Performance Engineer, Phd

Meta's AI Training and Inference Infrastructure is growing exponentially to supp...

Location

United States , Menlo Park

Salary:

122000.00 - 181000.00 USD / Year

AI/HPC System Performance Engineer

Meta's AI Training and Inference Infrastructure is growing exponentially to supp...

Location

United States , Austin

Salary:

219000.00 - 301000.00 USD / Year

Software Engineer - AI/HPC Specialist

We are looking for software engineers to help scale and improve the efficiency o...

Location

Norway , Oslo

Salary:

Not provided

Senior Software Engineer

Microsoft Azure High Performance Computing & AI Engineering (HPC & AI Eng) team ...

Location

United States , Multiple Locations

Salary:

119800.00 - 234700.00 USD / Year

Microsoft Corporation

Expiration Date

Until further notice

Requirements

Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, OR Java, JavaScript, or Python OR equivalent experience
3+ years of experience in operating AI/HPC systems, developing and running AI/HPC applications on clusters, or operating Cloud Infrastructure
2+ years of specialized experience with one of AI/HPC system management OR High-Speed Networks OR HPC Storage OR managing Cloud Infrastructure
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role
Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Job Responsibility

Collaborates with appropriate stakeholders to determine user requirements for a scenario
Drives identification of dependencies and the development of design documents for a product, application, service, or platform
Creates, implements, optimizes, debugs, refactors, and reuses code to establish and improve performance and maintainability, effectiveness, and return on investment (ROI)
Leverages subject-matter expertise of product features and partners with appropriate stakeholders (e.g., project managers) to drive a workgroup's project plans, release plans, and work items
Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate
Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of products while also driving consistency in monitoring and operations at scale

Fulltime

New

Systems Software Engineer

The Crusoe Cloud Software Development team is seeking a passionate and experienc...

Location

United States , San Francisco

Salary:

137000.00 - 161000.00 USD / Year

Crusoe

Expiration Date

Until further notice

Requirements

Linux Systems Familiarity: Experience building applications on Linux kernels, specifically pertaining to virtualization, device drivers, memory management, and process scheduling
Hardware Integration: Solid understanding of hardware devices such as GPUs, CPUs, Infiniband and Ethernet NICs, Ephemeral Disks, and PCI Express
Systems Design: Strong grasp of distributed applications and highly-scalable systems design. Specific focus around communications protocols (GRPC, REST, TCP/IP, etc.), databases (Postgres, Redis), and systems design applications (Pub/Sub, Kafka)
Software Architecture: Strong experience building software applications, both at the higher (Golang, Java, Python) and lower (C, C++, Rust) levels. Keen eye for clean, maintainable code, and a unit-test driven mindset
Excellent Communication Skills: Ability to collaborate with teams across an organization, blocking out noise, and focusing on what needs to get done to get a project across the line
Rapid and Agile Learner: Capable of adapting quickly, eager to research new technology and not get overwhelmed by unfamiliar tech stacks
Virtualization Concepts: General knowledge of hypervisors, virtual machine lifecycles, and Linux KVM tooling
CI/CD and Validation: Understanding of how to build Gitlab or Github CI/CD pipelines that deliver bug-free code across a multitude of compute platforms

Job Responsibility

Compute Application Development & Scaleout: Design highly reliable and performant Linux applications used to manage our virtualization stack across thousands of AI compute servers in multiple global datacenters
AI Hardware Platform Integration: Integrate Crusoe applications with a wide variety of hardware and software AI chip-vendor stacks. Build solutions to optimize and monitor virtualized hardware (GPUs, Infiniband/ROCe NICs, Ephemeral Storage, etc.) in cutting-edge AI/HPC environments
Kernel & Hypervisor Integration - Work side by side with our Linux Kernel and Hypervisor teams to ensure our Crusoe applications are seamlessly integrated with a variety of kernels and hypervisors
Performance Analysis & Tuning: Analyze and enhance the performance of the entire virtualization stack, from the hypervisor to the virtualized guest OS, with a specific focus on optimizing AI/ML workloads. This includes profiling, bottleneck identification, and implementing low-level optimizations
System-Level Troubleshooting: Diagnose and resolve complex system issues across our virtualization stack (drivers, kernel, hypervisor, guest OS, and crusoe applications). Work closely with kernel and hypervisor teams to debug and resolve integration challenges
Code Review and Quality Assurance: Conduct thorough code reviews to ensure the highest level of software quality, reliability, and security within compute applications and virtualization stack
Cross-Functional Collaboration: Collaborate with other engineering teams, including hardware design, OS development, and AI/ML application teams, to ensure cohesive and integrated product development
Technical Leadership: Provide technical guidance and mentorship to junior engineers, fostering a culture of technical excellence and collaborative problem-solving within the compute applications team

What we offer

Restricted Stock Units in a fast growing, well-funded technology company
Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
Employer contributions to HSA accounts
Paid Parental Leave
Paid life insurance, short-term and long-term disability
Teladoc
401(k) with a 100% match up to 4% of salary
Generous paid time off and holiday schedule
Cell phone reimbursement
Tuition reimbursement

Fulltime

Principal Software Engineer

Microsoft Azure High Performance Computing & AI Engineering (HPC & AI Eng) team ...

Location

United States , Multiple Locations

Salary:

139900.00 - 274800.00 USD / Year

Microsoft Corporation

Expiration Date

Until further notice

Requirements

Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python - OR equivalent experience
5+ years hands on experience designing and developing high volume low latency pipelines using products such as AzPubSub, Event Hubs, Azure Stream Analytics, Kafka, Grafana, Event Hubs, Prometheus or equivalent products
3+ years of experience with one of AI/HPC system management OR High-Speed Networks OR HPC Storage OR managing Cloud Infrastructure
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role
Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Job Responsibility

Architect, design and develop high volume low latency end to end event pipelines that can provide first-to-know-insights on events causing job interrupts and job reliability
Conduct analysis of existing event pipelines to evaluate fidelity, granularity and latency of critical events
Contribute to improving key metrics such as Job Mean Time to Interrupt, Nodes in Service, Mean Time to Resolve on flagship supercomputers by enabling data scientists and domain experts to use the telemetry to identify events & issues at the intersection of datacenter and hardware, develop hypothesis, conduct A/B tests and synthesize results
Partner with cross organizational teams to evaluate available telemetry and latency drive architecture, design, development and deployment of end-to-end solutions to manage core infrastructure including current & next generation datacenter, IT hardware, power & cooling technologies
Drive engineering and operational excellence based on issues and learnings from strategic customers on their usage scenarios to improve product features and capabilities
Partner with teams on continuous learning and continuous improvement programs by leading the resolution of complex incidents, driving root cause analyses and championing initiatives to minimize future customer impact

Fulltime

AI Research Lab Research Associate

We are currently seeking highly qualified interns to accelerate research towards...

Location

United States , Milpitas

Salary:

43.27 - 93.15 USD / Hour

Hewlett Packard Enterprise

Expiration Date

May 26, 2026

Requirements

Pursuing PhD degree (or other degree with significant research and innovation experience) in a relevant discipline (e.g. machine learning, computer science, electrical engineering, math, statistics, etc.)
Track record of world-class innovative contributions and ideas in machine learning
Experience with innovative solution development, such as developing proofs-of-concept, first-of-a-kind solutions, and/or technology transfer
Experience in deep learning research
Experience in developing deep learning software with high proficiency in data structures and algorithms
Strong programming skills and experience with Python, C/C++, and preferably Java
Software development experience in Deep Learning, GPU acceleration, and Model Optimization
Experience in Deep Learning and Machine Learning frameworks and models like Tensorflow, PyTorch
Experience in Transformer Neural Network architectures for Generative AI and natural language processing
Experience with Agentic AI and Generative AI workflows - desired

Job Responsibility

Conduct research and come up with solutions with a fast turnaround time
Build the software and applications for Neural Networks and Machine Learning
Work with system programming, Deep Learning frameworks and models, GPU acceleration, Model optimization, real-time streaming data, distributed computing, and deployment
Provide thought leadership and technical influence both internally and externally to HPE
Collaborate with HPE Labs research teams as well as external partners
Work in alignment with HPE's broader innovation community.

What we offer

Health & Wellbeing benefits including physical, financial and emotional wellbeing support
Personal and professional development programs
Unconditional inclusion and flexibility to manage work and personal needs.

Fulltime

AI/HPC Systems Performance Engineer

Meta

Location:
United States , Menlo Park

Category:
IT - Software Development

Contract Type:
Not provided

Salary:

Job Description:

Job Responsibility:

Requirements:

Nice to have:

Additional Information:

Job Posted:
February 21, 2026

Looking for more opportunities? Search for other job offers that match your skills and interests.

Similar Jobs for AI/HPC Systems Performance Engineer