HOME Data Science & ML Staff Machine Learning Engineer, Siri Attention and Invocation
  • Apple
  • Zurich, ZH
  • Full-Time
  • 32 days ago
Apple VERIFIED EMPLOYER

Staff Machine Learning Engineer, Siri Attention and Invocation.

Data Science & ML Full-Time

Staff Machine Learning Engineer, Siri Attention and Invocation: our view in 3 lines...

  • The Role:This role is for a Staff Machine Learning Engineer focused on generative audio and video systems for Siri Attention and Invocation.
  • The Person:The person will lead the development of audio and video generation capabilities, drive the technical vision for synthetic speech and visual representations, mentor engineers, and shape the roadmap for multimodal generative experiences.
  • Requirements:The ideal candidate has a Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, plus deep hands-on experience with generative audio and/or video architectures and experience evaluating generative model outputs.

About the role

As part of Siri Attention and Invocation, we collaborate to deliver the next revolution in human-computer interaction, to inspire and create groundbreaking technology for large-scale systems spanning speech, vision, and generative AI to overcome real-world challenges through innovation and user-centered design that improves the daily experience of millions of our customers.

Description

We are seeking an exceptional Staff Machine Learning Engineer to lead the development of audio and video generation capabilities that bring conversational agents to life. In this role, you will drive the technical vision for generating realistic, expressive synthetic speech and visual representations, mentor senior and junior engineers, and shape the roadmap for multimodal generative experiences across our products.

Minimum Qualifications

Master’s or PhD in Computer Science, Electrical Engineering, Machine Learning, or a related field, or equivalent practical experience
Deep hands-on experience with generative audio and/or video architectures (e.g., diffusion models, autoregressive models, GANs, VAEs, end-to-end neural synthesis)
Demonstrated ability to lead complex, ambiguous projects from research through production, and to make sound technical tradeoffs under real-world constraints (quality, latency, compute)
Experience evaluating generative model outputs, including both objective metrics and perceptual/subjective quality assessment

Preferred Qualifications

Proven experience building and shipping machine learning systems in production, with significant focus on generative modeling
Excellent collaboration and communication skills, with a track record of working across research, engineering, and product teams
Strong software engineering skills, with experience designing scalable ML systems and pipelines (e.g., Python, PyTorch/TensorFlow, distributed training infrastructure)
Experience with speech synthesis (TTS), voice conversion, audio acoustics/background modeling, or conversational AI systems
Experience with generative video/animation techniques (e.g., facial animation, lip-sync, avatar rendering, video diffusion)
Publications in generative modeling, speech, audio, or computer vision at top-tier venues (e.g., NeurIPS, ICML, ICASSP, CVPR, Interspeech)
Experience deploying real-time or low-latency generative models at scale
Familiarity with multimodal modeling (joint audio-visual generation, cross-modal conditioning)
Prior experience mentoring engineers or leading technical direction for a team

Published September 1, 2026
Location Zurich, Switzerland
Category Data Science & ML  
Job Type Full-Time