Job Description
Role Overview:
We are looking for an AI Infrastructure Engineer to build, deploy, and maintain scalable infrastructure for AI/ML and Generative AI workloads.
Key Responsibilities:
Design and manage infrastructure for AI/ML model training and inference.
Deploy and scale LLMs, AI agents, and ML services in production.
Manage GPU/CPU compute, containers, Kubernetes, and cloud infrastructure.
Build CI/CD pipelines for AI/ML applications and model deployments.
Implement monitoring, logging, observability, and performance optimization.
Optimize GPU utilization, latency, scalability, and infrastructure costs.
Work with data, AI, and software engineering teams to support production AI systems.
Ensure infrastructure security, reliability, and high availability.
Required Skills:
2–5+ years of experience in Cloud, DevOps, SRE, or AI Infrastructure.
Strong knowledge of AWS/Azure/GCP.
Hands-on experience with Docker and Kubernetes.
Experience with GPU infrastructure and AI/ML workloads.
Strong Linux, Python, networking, and system administration skills.
Experience with Terraform/IaC, CI/CD, monitoring, and observability.
Understanding of ML/LLM deployment and inference.
Preferred:
Experience with NVIDIA GPUs, CUDA, Ray, Kubeflow, MLflow, model serving, LLMOps, and distributed systems.