Aniruddh Chandratre

Palo Alto, CA · +1 (480) 584-2618 · achandratre22@gmail.com
linkedin.com/in/aniruddhchandratre · github.com/C-Aniruddh · Google Scholar

O-1A (current) · EB-1A approved

ML infrastructure engineer building distributed data platforms, GPU scheduling, and cluster-efficiency systems for FSD, Optimus, and Digital Optimus across large heterogeneous GPU fleets.

Experience

Senior Engineer, ML Infrastructure — Tesla AI

Feb 2024 – Present · Palo Alto, CA

Distributed data platform · GPU scheduling · cluster efficiency

  • Lead Tesla AI's internal data processing platform for acquisition, curation, mining, auto-labeling, training, and evaluation; it runs 10B+ jobs and 3,000+ compute-years/month across 10 GPU clusters.
  • Re-architected the platform core, distributed worker coordination, and memory model, increasing throughput from 940M to 15B+ data items/month.
  • Built scheduling and live defragmentation for a heterogeneous GPU fleet, increasing usable compute capacity from 24% to 88% and fleet-wide effective utilization to 98%.
  • Added end-to-end lineage across artifacts, datasets, training jobs, and models for multi-stage pipeline provenance.
  • Built internet-scale acquisition and curation pipelines across Tesla AI programs.

Software Engineer, Automation Software — Tesla

Jan 2023 – Jan 2024 · Fremont, CA

Edge-to-cloud data infrastructure for the gigafactories

  • Designed and shipped an edge-to-cloud data broker processing 10B+ manufacturing data points/day across Austin, Berlin, and Fremont.
  • Developed an mmap-backed durable queue sustaining 5M writes/s on NFS-backed Kubernetes volumes and a distributed cache processing 10M entries/s, cutting storage 80% across 50 deployments.
  • Built and operated 20 high-availability services for factory insights.

Graduate Research Assistant — ASU, Cyber-Physical Systems Lab

Jan 2021 – Dec 2022 · Tempe, AZ
  • Researched formal verification under DARPA ARCOS with Lockheed Martin; co-authored PSY-TaLiRo and operated the Kubernetes cluster for high-volume simulation workloads.

Founding ML Engineer — Aftershoot

Aug 2020 – Jan 2021 · Remote
  • Built the image-ranking model at the core of AfterShoot's photo-culling product, now used by approximately 250K photographers.

Education

M.S. Robotics & Autonomous Systems (AI) — Arizona State University · 4.0/4.0, Distinction Medal 2022
B.Tech Electronics & Telecommunications — NMIMS MPSTME, Mumbai · Dean's List 2020

Research, Awards & Tooling

Publications  8 peer-reviewed papers · 174 citations — HSCC, EMSOFT, FMICS, ARCH. First author, HSCC 2023: Stealthy Attacks Formalized as STL Formulas for Falsification of CPS Security.

Awards  5× national hackathon winner in India — Smart India Hackathon 2019, Mastek Deep Blue (×2), IET Hack N Code, RAENG.

Languages  Go, Python, C, C++, JavaScript/TypeScript, Java

Systems  Slurm, Kubernetes, Helm, Docker, AWS, GCP, Prometheus, Grafana, Kafka, ClickHouse, InfluxDB, MQTT, SQL