Latest job information from Singtel for the position of GPU-Driven DevOps Engineer - AI/HPC Infra, Flexible Work. If the GPU-Driven DevOps Engineer - AI/HPC Infra, Flexible Work vacancy in Singapore matches your qualifications, please submit your latest application or CV directly through the updated Jobkos job portal.
Please note that applying for a job may not always be easy, as new candidates must meet certain qualifications and requirements set by the company. We hope the career opportunity at Singtel for the position of GPU-Driven DevOps Engineer - AI/HPC Infra, Flexible Work below matches your qualifications.
About Singtel Digital InfraCo – RE:AI Singtel Digital InfraCo’s RE:AI division is building Asia’s most advanced and sustainable AI infrastructure ecosystem. RE:AI enables enterprises, research institutions, and digital-native businesses to accelerate innovation through responsible, high-performance AI compute and connectivity solutions.
About Singtel Digital InfraCo – RE:AI Singtel Digital InfraCo’s RE:AI division is building Asia’s most advanced and sustainable AI infrastructure ecosystem. RE:AI enables enterprises, research institutions, and digital-native businesses to accelerate innovation through responsible, high-performance AI compute and connectivity solutions. Be a Part of Something BIG! As an DevOps Engineer for SingTel’s GPU-as-a-Service (GPUaaS), you will help in implementing processes and integration of operations to advance customer’s AI and HPC capabilities. You will be exposed to both physical data center implementation and software solutions in a Singtel GPU-as-a-Service (GPUaaS). This position requires a forward-thinking individual who thrives in dynamic environments and is committed to driving continuous improvement in GPU for AI and HPC environments. T his role is suitable for professionals looking to develop their expertise in DevOps and AI/HPC cloud platforms. Responsibilities
Design, deploy and support large-scale, distribute GPU clusters for AI and ML workloads.
Manage and automate provisioning of GPU resources in both on-prem and cloud platforms.
Design, implement and manage CI/CD pipelines for AI models and GPU-accelerated applications.
Monitor cluster usage, health, performance and availability.
Improve infrastructure provisioning, management, and monitoring through automation.
Troubleshoot compute resource system level issues such as Slurm, Kubernetes, GPU drivers, CUDA, IB networking.
Optimize system parameters (e.g., OS, drivers, networking, library) for AI workload performance.
Conduct GPU cluster benchmark and keeping up with the latest advancements in GPU technology.
Set up monitoring and logging for GPU resources using Zabbix, Prometheus, NVIDIA DCGM and other tools.
Implement security best-practices for multi-tenant GPU-as-a-Service (GPUaaS) environment.
Collaborate with software and administrator to to streamline workflows and improve collaboration.
Providing technical support and guidance to users of GPU-accelerated systems.
Work with senior DevOps engineer to identify bottlenecks and improve development and operational processes for AI and HPC GPU cloud.
Learning to solve problems in high-performance distributed computation for AI and HPC GPU cloud computing.
Participate in rotational or scheduled shift work as required to support platform operations.
Requirements
Bachelor’s degree in Computer Science/Engineering, Information Technology, Systems Engineering, or a related field.
Strong Linux system administration skills in Ubuntu/CentOS/Rocky Linux, etc.
Experience with DevOps tools such as Jenkins, Kubernetes, Ansible and Terraform.
Solid understanding of DevOps practices, including CI/CD, automation, and monitoring.
Proficiency in scripting languages (e.g., Python, Bash).
Experience in implementing monitoring solutions such as Zabbix, Prometheus.
Familiarity with AI frameworks such as TensorFlow, PyTorch.
Understanding of cloud architectures (IaaS, PaaS), GPU architecture and NVIDIA GPUs.
Strong verbal, written, and presentation skills in English.
Team player with experience in cross-functional coordination.
Strong technical problem solving and analytical skills for system optimization.
Desirable Qualifications
Understanding of how collective communications (MPI, RDMA, and NCCL) works, as well as an understanding of GPU specific aceleration works on GPU cluster.
Knowledge of DevOps/ML Ops technologies in GPU cluster such as Docker/containers, Kubernetes, data center deployments
Familiarity with Slurm or other HPC workload managers to manage GPU clusters.
Understanding of AI & HPC networking technologies such as InfiniBand, RoCE, DPUs.
System-level experience specifically GPU-based systems (NVIDIA GPU and SDKs)
Understanding how AI and HPC workloads interact with both GPU HW and SW infrastructure.
Rewards that Go Beyond
Flexible work arrangements
Full suite of health and wellness benefits
Ongoing training and development programs
Internal mobility opportunities
Your Career Growth Starts Here. Apply Now! #J-18808-Ljbffr
Job Info:
Company: Singtel
Position: GPU-Driven DevOps Engineer - AI/HPC Infra, Flexible Work
Work Location: Singapore
Country: SG
How to Submit an Application:
After reading and understanding the criteria and minimum qualification requirements explained in the job information GPU-Driven DevOps Engineer - AI/HPC Infra, Flexible Work at the office Singapore above, immediately complete the job application files such as a job application letter, CV, photocopy of diploma, transcript, and other supplements as explained above. Submit via the Next Page link below.
THIS JOB POSTING HAS EXPIRED (Over 30 days ago).
Please search for the latest job opportunities on our
Homepage.
Desc: The Nomad Recruiter Collective is a place for digital Nomad Recruiters to hang out and help each other thrive.Join our LinkedIn Group on We are a group for Digital Nomad Recruiters. Share jobs, leads,...
Desc: Are you a recruiter looking to hire? HR manager with loads of vacancies?Do you want to earn a residual income? is a revolutionary platform designed to redefine the recruitment industry. We offer a com...
Desc: Responsibilities:Conduct ICT workshops and basic/intermediate ICT lessons for staff and students.Provide support for Google Workspace, Chrome/Android, MS Teams and Google Meet.Support teachers and stu...
Desc: About the jobMercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'A...
Desc: About the jobMercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'A...