About Navana.ai

At Navana, we are at the forefront of developing Voice AI solutions tailored for Indic languages, driving growth and innovation for large enterprises.

We build real-time voice AI systems that power critical customer interactions for leading enterprises. Our stack is built for low-latency, high-throughput workloads and is deployed both in our cloud and on customer infrastructure. We train and serve our own Speech and Language models (STT, TTS, and SLMs) and as we scale, we are investing in the platform those models are trained on, evaluated against, released through, and run on in production.


Core Responsibilities

  1. Training and fine-tuning pipelines. Build reproducible, multi-node pipelines for STT, TTS, and SLM training: checkpoint/resume, spot-interruption tolerance, and per-run cost visibility utilising tools such as Kubeflow Pipelines, Argo Workflows, Flyte, Metaflow, Airflow, etc.
  2. Experiment tracking, model registry, and lineage. Stand up the system of record for every run and every artifact (MLflow, Weights & Biases, ClearML) so any model serving live traffic traces back to its exact code, config, hyperparameters, data snapshot, and base checkpoint.
  3. Audio data and model versioning at scale. Version multi-terabyte speech corpora and their derived manifests (DVC, lakeFS, Delta Lake), and turn dataset curation into pipelines with lineage instead of one-off notebooks.
  4. Model evaluation as CI. Make quality a release gate, not a spreadsheet: automated WER/CER by language, dialect, and acoustic condition; TTS intelligibility, latency, RTF, and memory budgets enforced per candidate model; golden test sets that block a bad promotion before a customer finds it.
  5. Low-latency inference and optimization. Productionize and tune STT, TTS, and SLM serving under hard latency budgets (NVIDIA Triton, TensorRT, vLLM, ONNX Runtime, etc), optimizing for time-to-first-token and first-audio latency.
  6. Release, rollback, and on-prem delivery. Ship versioned, repeatable deployments into our cloud and into constrained or air-gapped customer environments.
  7. GPU infrastructure end to end. Drivers, CUDA / cuDNN / NCCL, NVIDIA GPU Operator (driver, container toolkit, device plugin, DCGM), MIG partitioning and time-slicing, mixed-GPU scheduling across different GPU families, and DCGM-based utilization, saturation, and thermal metrics you actually act on.
  8. The practices themselves. Define how ML and engineering train, evaluate, and ship. You are building the paved road the rest of the company drives on.

Must Have Qualifications