Machine Learning Engineer
What is the role?
As a Machine Learning Engineer at Tensorplay, you’ll be the bridge between research and production. You’ll take models — whether fine-tuned LLMs, vision models, or classical ML systems — and build the infrastructure, pipelines, and APIs that make them reliable, fast, and observable in the real world.
You’ll work directly with our clients’ AI codebases, diagnose production failures, and architect solutions that actually scale.
About you
You care deeply about how models actually behave in production — not just on eval benchmarks. You’ve seen LLM hallucinations, inference latency spikes, and embedding drift firsthand, and you know how to design systems that handle these gracefully. You’re comfortable with ambiguity and excited to work on genuinely hard engineering problems.
Your Responsibilities
- Design and optimize model serving infrastructure (Triton, vLLM, TorchServe)
- Build RAG pipelines and vector search systems at scale
- Fine-tune and evaluate LLMs for specific client use cases
- Implement monitoring and observability for model performance in production
- Work with clients to understand their AI needs and translate them into engineering solutions
Requirements
- 3+ years of experience in ML engineering or applied ML research
- Strong Python skills and experience with PyTorch or TensorFlow
- Hands-on experience with LLMs (OpenAI API, Hugging Face Transformers, or similar)
- Familiarity with vector databases (Pinecone, Weaviate, pgvector, Qdrant)
- Experience deploying models on cloud platforms (AWS, GCP, or Azure)
- Understanding of REST API design and async backend systems
Nice to Have
- Experience with Kubernetes and GPU workload scheduling
- Knowledge of RLHF, DPO, or other LLM alignment techniques
- Contributions to open-source ML projects
We Offer
- Competitive salary benchmarked against global rates
- Full remote flexibility with async-first culture
- Work on real production AI systems used by real companies
- Direct mentorship from founding engineers