MLOps Engineer
What is the role?
At Tensorplay, every AI project we deliver needs bulletproof infrastructure. As our MLOps Engineer, you’ll design the pipelines, tooling, and platforms that allow models to be trained, versioned, deployed, and monitored at scale. You’ll be the person who ensures our clients’ AI systems stay up — and get better — over time.
About you
You love the intersection of software engineering and machine learning operations. You’ve built CI/CD pipelines that automatically retrain and deploy models, set up Prometheus/Grafana dashboards for model metrics, and debugged Kubernetes scheduling nightmares at 2am. You know that good MLOps is invisible — things just work.
Your Responsibilities
- Build and maintain model training pipelines using tools like Kubeflow, MLflow, or Prefect
- Design and operate scalable inference infrastructure on Kubernetes with GPU support
- Implement automated model evaluation, regression testing, and deployment gates
- Set up observability stacks — metrics, logs, traces, and semantic drift detection
- Manage cloud infrastructure (AWS/GCP/Azure) using Terraform or Pulumi
- Collaborate with ML engineers to optimize model serving latency and cost
Requirements
- 3+ years of DevOps or MLOps experience, with at least 1 year focused on ML systems
- Strong knowledge of Kubernetes, Docker, and container orchestration
- Experience with cloud platforms — AWS SageMaker, Vertex AI, or Azure ML
- Familiarity with MLflow, DVC, or equivalent experiment tracking and model registry tools
- Proficiency in Python and infrastructure-as-code (Terraform preferred)
- Understanding of GPU scheduling and CUDA workload management
Nice to Have
- Experience with NVIDIA Triton Inference Server or Ray Serve
- Knowledge of Prometheus, Grafana, and OpenTelemetry
- Background in data engineering and pipeline orchestration (Airflow, Dagster)
We Offer
- Competitive salary with equity participation
- Fully remote, async-first work culture
- GPU credits and cloud environment for learning and exploration
- Access to a diverse portfolio of AI projects across industries
- Annual conference and certification budget