Full Details
<h2>About the Company</h2>
<p>Level AI is a Series C conversational intelligence company headquartered in Mountain View, CA, backed by top-tier venture capital firms and experienced Silicon Valley operators. The company helps enterprise contact centers understand every customer conversation — using advanced speech AI, NLP/NLU, and retrieval systems to transform millions of unstructured interactions into actionable business decisions. With a strong engineering team that includes talent from Amazon, Google, and Meta, Level AI is at an inflection point where scaling infrastructure and architecture is critical for the next stage of growth.</p>
<h2>Role Overview</h2>
<p>We are seeking a Principal Software Engineer — Backend & Infrastructure to lead the technical direction for backend and ML infrastructure across multiple teams. This role is not about building inside existing systems; it is about owning the transition from Series C to the next stage. You will make architectural decisions that are expensive to reverse, raise the engineering bar through design reviews and mentorship, and set standards that outlast any single project. You will report to the VP of Engineering and work closely with ML, Product, and Infrastructure leads across both our Noida and Bengaluru sites.</p>
<p>This position is ideal for an engineer who thrives on solving unsolved problems, enjoys real ownership over systems used by enterprises worldwide, and wants to work alongside engineers from top tech companies who chose to build here because the challenges are genuine and the autonomy is real.</p>
<h2>Key Responsibilities</h2>
<ul>
<li>Own the architecture for real-time data processing at scale by designing and evolving distributed messaging systems that handle high-throughput streaming workloads with strict latency requirements.</li>
<li>Build the ML platform that enables faster model shipping — define and execute the technical roadmap for training and serving infrastructure as models grow in size, complexity, and inference cost.</li>
<li>Scale GPU infrastructure by owning capacity planning, scheduling, and utilization across training and inference fleets, deciding what runs where and how to handle burst demand while controlling costs.</li>
<li>Scale inference by driving down latency and cost per request through batching, routing strategies, quantization, compilation, autoscaling, and caching — without degrading output quality.</li>
<li>Make reliability a property of the system by driving uptime, observability, and incident response for serving systems that enterprise customers depend on in production.</li>
<li>Turn ambiguous business problems into executable technical plans by partnering with Product and GTM teams to scope large cross-functional initiatives and break them into work that other teams can execute.</li>
<li>Multiply the team through leading design reviews, mentoring senior engineers, and shaping engineering practices that endure beyond any single project.</li>
<li>Bring outside perspective by evaluating emerging tools and techniques with judgment — adopting only what earns its complexity.</li>
</ul>
<h2>Tech Stack</h2>
<ul>
<li><strong>Programming Languages & Frameworks:</strong> Python (Django), Celery for task queues</li>
<li><strong>Databases & Storage:</strong> PostgreSQL, Redis</li>
<li><strong>Cloud & Infrastructure:</strong> Google Cloud Platform (GCP)</li>
<li><strong>Messaging & Streaming:</strong> High-throughput distributed messaging systems, real-time job queues (exact technologies not specified but implied Kafka/PubSub-like)</li>
<li><strong>ML & Inference:</strong> GPU infrastructure (capacity planning, scheduling), vLLM, TensorRT, Triton, Ray Serve</li>
<li><strong>Observability:</strong> Monitoring, incident response tooling</li>
<li><strong>Other:</strong> Speech AI, NLP/NLU, information retrieval systems</li>
</ul>
Requirements
10+ years building backend and infrastructure systems with a track record of owning architecture and design at scale. Deep hands-on experience with large-scale databases, high-throughput messaging systems, and real-time job queues. Proven ability to navigate complex codebases and reason about architectural tradeoffs. Experience mentoring senior engineers and driving technical decisions through influence. Strong written communication for cross-timezone and executive audiences. BTech/MTech/PhD in Computer Science or equivalent. Bonus: experience with Django, Celery, Redis, PostgreSQL, Google Cloud; scaling GPU infrastructure and model inference; specific inference tooling (vLLM, TensorRT, Triton, Ray Serve); scaling a platform through Series C to D; background in speech, NLP, or information retrieval.