Senior AI/ML Data Scientist or Engineer - LLM Fine-Tuning, Evaluation & Inference
Software Engineering, Data Science · Full-time
India
Location: India only | Remote
Type: Full-time | Permanent
Role overview
Own the complete ML lifecycle for an enterprise LLM system that selects and executes browser-based tools. You will turn raw execution logs into high-quality training data, run fine-tuning experiments, build rigorous evaluations, and deploy and optimise models on AWS. This is an applied ML engineering role - not prompt engineering, analytics or model-ops-only work.
Core responsibilities
ML data and evaluation
- Build pipelines to ingest, parse, clean, normalise, enrich and deduplicate execution logs and documents.
- Convert raw logs into structured supervised examples with context, candidate tools, expected actions and reference labels.
- Create versioned training, validation and holdout datasets with privacy, schema, label and quality checks.
- Design leakage-aware splits using task- or session-level boundaries.
- Build reproducible evaluations for tool-selection accuracy, reference accuracy, hallucinations, disagreements and failure slices.
- Trace misleading metrics back to inputs, labels, dataset versions and model versions.
Model development
- Benchmark open-weight models such as Qwen, Llama and Mistral on quality, training effort, long-context behaviour and inference cost.
- Run supervised fine-tuning, knowledge distillation and parameter-efficient fine-tuning experiments.
- Apply sound experiment design: clear baselines, controlled comparisons, reproducible configurations and error analysis.
- Identify overfitting, leakage, sampling bias, label noise and uneven task coverage.
- Maintain versioned datasets, prompts, training configurations, checkpoints and evaluation code.
- Recommend which experiments justify additional training or inference cost.
ML platform and MLOps
- Take models through deployment, testing, monitoring and controlled release.
- Benchmark serving configurations for throughput, latency, GPU memory, concurrency and cost at an agreed quality level.
- Evaluate vLLM or comparable runtimes, INT8 quantisation, checkpoint compatibility and KV-cache requirements.
- Measure cost per successful task or action, not just cost per token.
- Implement model versioning, CI/CD, shadow inference, disagreement logging, feature-flag releases and rollback.
- Monitor data drift, quality regressions, serving errors, latency and cost.
LLM workflow engineering
- Build pipelines involving document parsing, chunking, embeddings, vector indexing and retrieval.
- Implement function or tool calling, multi-step workflows, retries, fallbacks and state handling.
- Design observability across prompts, retrieved context, model outputs, tool calls and downstream failures.
- Write production-grade, tested Python for large-scale document and execution-data processing.
Essential qualifications
- 7+ years of Python experience.
- 5+ years of AWS experience.
- 7+ years across ML, data, backend or applied AI engineering, with production delivery experience.
- Strong supervised ML fundamentals, including validation design, overfitting, leakage, sampling bias and error analysis.
- Hands-on transformer fine-tuning using PyTorch, Hugging Face or equivalent tooling.
- Direct ownership of a completed model training, fine-tuning, distillation or benchmarking project.
- Experience preparing ML/LLM datasets and building model evaluation code.
- Experience with experiment tracking, model versioning, CI/CD, deployment, monitoring and rollback.
- Working knowledge of GPU inference, quantisation, batching, concurrency, tail latency and cost optimisation.
- Experience with large-scale data processing using Spark, Kafka, Flink or comparable technologies.
- Strong SQL and practical experience with relational and NoSQL systems.
- Production experience with RAG, embeddings, vector databases, function calling or agentic workflows.
- Ability to work autonomously and own outcomes without detailed specifications.
Preferred background
- Startup or similarly high-ownership environment.
- AWS S3, Glue, IAM, EC2 GPU instances, CloudWatch and model registries.
- vLLM, NVIDIA GPU optimisation or AWS G5-family instances.
- Qwen, Llama or Mistral model families.
- LoRA, QLoRA, knowledge distillation or long-context evaluation.
- Browser automation, test execution logs or tool-selection workflows.
- Enterprise or regulated environments where privacy, auditability and governance matter.
What to include with your application
- A model you trained, fine-tuned, distilled or benchmarked, including the dataset, validation approach, metrics and your contribution.
- An ML data, evaluation, deployment or inference system you owned, including scale, failure handling and measurable results.
- Your current Indian city, earliest start date and interview availability.