About the Role
We are seeking highly experienced Senior AI Scientists to lead the design, implementation, and scaling of AI infrastructure. In this role, you will bridge the gap between data science and software engineering by building robust, automated pipelines and establishing best practices for model development, deployment and lifecycle management.
We are looking for someone who is passionate about building highly efficient, reusable, and developer-friendly AI systems.
Why Join Us
At the Agency for Science, Technology and Research (A*STAR), Singapore’s leading public sector R&D agency, you will work at the vibrant intersection of frontier scientific research and real-world industrial translation. Engineers at A*STAR have the unique opportunity to design and build AI infrastructure that scales across incredibly diverse, multi-disciplinary domains—from advanced manufacturing and digital healthcare to sustainability and transportation.
A*STAR heavily invests in its engineers’ growth, offering a highly collaborative research culture, competitive benefits, robust pathways for continuous learning, and the unique chance to work on nationally-significant Smart Nation initiatives that impact lives globally.
What You Will Do
- Architect the Engine: Design and build the scalable, end-to-end MLOps infrastructure required to seamlessly train, deploy, and monitor machine learning models in high-stakes production environments.
- Automate the Lifecycle: Pioneer advanced CI/CD/CT (Continuous Training) pipelines specifically engineered for ML workflows, ensuring a flawless, zero-friction transition from experimental research to live production.
- Force-Multiply Teams: Champion massive organizational efficiency by building centralized feature stores, model registries, and reusable workflow templates that drastically accelerate time-to-market across all data science pods.
- Optimize Inference: Deploy lightning-fast, high-throughput model inference services. You will relentlessly optimize models for peak performance, extreme scalability, and operational cost-efficiency.
- Ensure Reliability: Implement deep, comprehensive observability across all ML systems. You will rigorously track data drift and model degradation, establishing ironclad governance and testing protocols for all AI artifacts.
- Lead and Mentor: Act as the strategic bridge across Data Science, Engineering, and Product. You will translate complex model requirements into robust engineering solutions, mentor junior engineers, and define the gold standard for MLOps company-wide.
Requirements
- Proven Scale: 6+ years of experience in Software Engineering, Platform Engineering, or DevOps, with at least 3+ years strictly dedicated to MLOps, AI Infrastructure, or ML Engineering in a high-traffic, production environment. Fresh graduates with great interest in mastering advanced AI capabilities are welcome to apply
- Architectural Vision: Demonstrated experience designing, building, and maintaining end-to-end ML platforms from the ground up, moving beyond single-model deployments to centralized organizational infrastructure.
- Engineering Rigor: Deep proficiency in Python and strong knowledge of at least one high-performance backend language (e.g., Go, C++, Rust, or Java). You write clean, testable, and production-grade code.
- Cloud & Container Mastery: Experience with cloud-native architectures (AWS, GCP, or Azure) and deep fluency in containerization and orchestration (Docker, Kubernetes, Helm).
For The MLOps Stack
- Serving & Optimization: Hands-on expertise deploying low-latency, high-throughput inference services using frameworks like Triton Inference Server, Ray Serve, BentoML, TF Serving, or TorchServe.
- Pipelines & Orchestration: Strong experience building advanced CI/CD/CT pipelines and workflow orchestration using tools like Kubeflow, Apache Airflow, MLflow, ArgoCD, or Metaflow.
- Component Reusability: Proven track record of implementing and managing Feature Stores (e.g., Feast, Hopsworks) and Model Registries to accelerate data science workflows.
- Observability: Experience setting up comprehensive ML monitoring systems to detect data drift, concept drift, and performance degradation (using tools like Prometheus, Grafana, Arize, Evidently, or Fiddler).
- Infrastructure as Code (IaC): Solid understanding of Terraform, Ansible, or CloudFormation to ensure infrastructure is reproducible and scalable.