Skip to main content
Apple

Staff ML Infrastructure Engineer

Apple Cupertino, California, United States Full-time 4 days ago
AI & ML Engineering
Join a team at the forefront of ML infrastructure and generative AI, where data and model workflows come together to enable the next generation of intelligent experiences on Apple products and services. We build robust systems that connect scalable data pipelines with advanced ML workflows, accelerating the development of real-world AI applications. Our work spans the full ML lifecycle, from experimentation to deployment, and you’ll play a key role in shaping how AI models are built, optimized, and scaled. We develop a platform for ML data and features that powers advanced GenAI applications. This includes embeddings (generation, evaluation, ANN search, multimodal support), AI Ops, efficient inference, and a modern feature platform designed to streamline experimentation and drive innovation. We’re looking for engineers and researchers passionate about generative models, data-centric ML, and intelligent systems across diverse real-world use cases. With the autonomy to experiment, the scale to make an impact, and the support to take ideas from prototype to production, you’ll work alongside a world-class team to build intelligent, flexible systems that make ML development faster, more reliable, and more creative. \\n

The Apple AI Platform team gives Apple"s ML engineers and researchers the data systems and large-scale compute they need to build and ship models at Apple"s bar for quality and privacy. Our team owns the data layer that large-scale model training depends on: ingestion, versioning, lineage, and governance on the way in, and high-throughput data loading into the training fleet on the way out. As a Staff ML Infrastructure Engineer, you will set the technical direction for that platform and own its hardest system-level problems, the architecture other engineers and teams build on.

Own the architecture of the platform behind Apple"s largest model builds: define how ingestion, immutable versioning, lineage, and governance work across structured, unstructured, and multimodal data at petabyte scale, so every model run is reproducible from a versioned dataset.\\nSet the technical direction for high-throughput data delivery to Apple"s largest GPU and TPU fleets: define the data access and loading architecture that keeps training compute-bound, not I/O-bound.\\nMake the hard system-level and format calls that the whole platform inherits, columnar and lakehouse strategy, the dataset abstraction spanning structured and multimodal data, the shape of the SDK and core libraries, backed by design and proof, not just opinion.\\nDrive technical direction and influence across the platform and partner teams (data, embeddings, features, research), and define the interfaces and contracts between them.\\nRaise the technical bar across the team: mentor senior engineers, lead design reviews, and be the escalation point for the problems no one else can crack.\\nPartner with research and product leadership to shape the platform roadmap for next-generation workloads: foundation models, multimodal data, and retrieval-augmented systems.\\nDrive efficiency, reliability, and automation across the data plane and control plane that power Apple"s ML fleet.

10+ years of work experience in machine learning infrastructure, distributed data systems, or a related field.\\n10+ years of experience building and shipping large-scale data or ML infrastructure and platforms in production.\\nExtensive experience architecting and delivering large-scale distributed data or ML infrastructure that multiple teams or products depend on in production.\\nA track record of setting technical direction and driving it to delivery across teams, not just within a single component.\\nDeep systems engineering: strong Python plus a systems language (Rust strongly preferred; C++ or Go acceptable), and hands-on performance engineering for I/O-bound workloads (Arrow, zero-copy, memory mapping, async I/O, high-throughput object storage).\\nDeep familiarity with columnar and lakehouse formats (Parquet, Iceberg, Delta, or Lance) and the judgment to choose between them at scale.\\nStrong working knowledge of the end-to-end ML workflow and how training and inference consume data, enough to architect data systems that serve them.\\nFamiliarity with modern ML and generative techniques (transformers, diffusion, retrieval-augmented generation, fine-tuning) at the level needed to design for those consumers.\\nDemonstrated ability to design highly available, easy-to-use systems and to mentor and elevate the engineers around you.\\nStrong collaboration and communication, with the ability to align multiple teams around a technical direction.\\nB.S., M.S., or Ph.D. in Computer Science, Computer Engineering, or equivalent practical experience.

Experience defining data or ML platform architecture that was adopted across an organization.\\nDeep experience with the data-loading and dataset-access layer of a modern ML framework (PyTorch, JAX, or TensorFlow).\\nDistributed data-loading frameworks for ML: Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset.\\nExperience feeding data to GPU or TPU fleets at scale and keeping them saturated.\\nData lineage and governance systems: DataHub, OpenLineage, Unity Catalog, or equivalent.\\nContributions to or operational experience with Spark, Daft, Polars, or DuckDB internals.\\nContainerization and orchestration (Docker, Kubernetes).
Apply now
Cupertino, California, United States
On-site
Full-time
4 days ago

Share this job