Certified Remote
PUBLISHED
Aug 10, 2026
Lead the design and development of scalable, real-time AI systems at a Chennai-based technology solutions company. This role involves architecting end-to-end solutions from model development to production deployment, with a focus on low-latency inference and MLOps best practices.
Recode Solutions is a technology-driven organization based in Chennai, Tamil Nadu, India. The company specializes in building advanced software and AI-powered systems for diverse business domains. With a focus on innovation and scalability, Recode Solutions partners with clients to deliver production-ready solutions that leverage the latest in artificial intelligence, machine learning, and real-time data processing. The company fosters a collaborative culture where engineers, data scientists, and product teams work together to solve complex problems.
We are seeking an experienced AI/ML Architect to lead the design and development of scalable, real-time AI systems. This is a full-time, senior-level position (10-12 years of experience) based in Chennai. The successful candidate will work closely with product, data, and engineering teams to architect end-to-end solutions — from model development and deployment to system integration and production monitoring. The role demands deep technical expertise in building low-latency inference pipelines, selecting appropriate infrastructure, and ensuring reliability and observability of deployed systems. You will also mentor engineers and help establish best practices for MLOps across the organization.
6+ years of experience building and deploying ML systems in production. Proven expertise in real-time, low-latency system design (streaming inference, event-driven pipelines). Strong understanding of scalable architectures — microservices, message queues, distributed training/inference. Proficient in Python and ML/DL frameworks (scikit-learn, TensorFlow, PyTorch). Hands-on experience with LLM inference optimization using vLLM, TensorRT-LLM, and SGLang. Familiarity with vector databases, embedding-based retrieval, and RAG pipelines. Experience with containerized environments (Docker, Kubernetes) and managing multi-container applications. Working knowledge of cloud platforms (AWS, GCP, or Azure) and CI/CD practices for ML workflows.
Opportunity to work on cutting-edge AI systems and real-time inference pipelines. Collaborate cross-functionally with product, data, and engineering teams in a dynamic environment. Mentorship opportunities and the chance to set best practices for ML lifecycle (MLOps). Exposure to edge deployments and model compression/optimization techniques.