Senior AI/ML Data Scientist or Engineer - LLM Fine-Tuning, Evaluation & Inference

Vamstar
Vamstar

Software Engineering, Data Science · Full-time

India

Posted on Sep 19, 2026

Location: India only | Remote

Type: Full-time | Permanent

Role overview

Own the complete ML lifecycle for an enterprise LLM system that selects and executes browser-based tools. You will turn raw execution logs into high-quality training data, run fine-tuning experiments, build rigorous evaluations, and deploy and optimise models on AWS. This is an applied ML engineering role - not prompt engineering, analytics or model-ops-only work.

Core responsibilities

ML data and evaluation

  • Build pipelines to ingest, parse, clean, normalise, enrich and deduplicate execution logs and documents.
  • Convert raw logs into structured supervised examples with context, candidate tools, expected actions and reference labels.
  • Create versioned training, validation and holdout datasets with privacy, schema, label and quality checks.
  • Design leakage-aware splits using task- or session-level boundaries.
  • Build reproducible evaluations for tool-selection accuracy, reference accuracy, hallucinations, disagreements and failure slices.
  • Trace misleading metrics back to inputs, labels, dataset versions and model versions.

Model development

  • Benchmark open-weight models such as Qwen, Llama and Mistral on quality, training effort, long-context behaviour and inference cost.
  • Run supervised fine-tuning, knowledge distillation and parameter-efficient fine-tuning experiments.
  • Apply sound experiment design: clear baselines, controlled comparisons, reproducible configurations and error analysis.
  • Identify overfitting, leakage, sampling bias, label noise and uneven task coverage.
  • Maintain versioned datasets, prompts, training configurations, checkpoints and evaluation code.
  • Recommend which experiments justify additional training or inference cost.

ML platform and MLOps

  • Take models through deployment, testing, monitoring and controlled release.
  • Benchmark serving configurations for throughput, latency, GPU memory, concurrency and cost at an agreed quality level.
  • Evaluate vLLM or comparable runtimes, INT8 quantisation, checkpoint compatibility and KV-cache requirements.
  • Measure cost per successful task or action, not just cost per token.
  • Implement model versioning, CI/CD, shadow inference, disagreement logging, feature-flag releases and rollback.
  • Monitor data drift, quality regressions, serving errors, latency and cost.

LLM workflow engineering

  • Build pipelines involving document parsing, chunking, embeddings, vector indexing and retrieval.
  • Implement function or tool calling, multi-step workflows, retries, fallbacks and state handling.
  • Design observability across prompts, retrieved context, model outputs, tool calls and downstream failures.
  • Write production-grade, tested Python for large-scale document and execution-data processing.

Essential qualifications

  • 7+ years of Python experience.
  • 5+ years of AWS experience.
  • 7+ years across ML, data, backend or applied AI engineering, with production delivery experience.
  • Strong supervised ML fundamentals, including validation design, overfitting, leakage, sampling bias and error analysis.
  • Hands-on transformer fine-tuning using PyTorch, Hugging Face or equivalent tooling.
  • Direct ownership of a completed model training, fine-tuning, distillation or benchmarking project.
  • Experience preparing ML/LLM datasets and building model evaluation code.
  • Experience with experiment tracking, model versioning, CI/CD, deployment, monitoring and rollback.
  • Working knowledge of GPU inference, quantisation, batching, concurrency, tail latency and cost optimisation.
  • Experience with large-scale data processing using Spark, Kafka, Flink or comparable technologies.
  • Strong SQL and practical experience with relational and NoSQL systems.
  • Production experience with RAG, embeddings, vector databases, function calling or agentic workflows.
  • Ability to work autonomously and own outcomes without detailed specifications.

Preferred background

  • Startup or similarly high-ownership environment.
  • AWS S3, Glue, IAM, EC2 GPU instances, CloudWatch and model registries.
  • vLLM, NVIDIA GPU optimisation or AWS G5-family instances.
  • Qwen, Llama or Mistral model families.
  • LoRA, QLoRA, knowledge distillation or long-context evaluation.
  • Browser automation, test execution logs or tool-selection workflows.
  • Enterprise or regulated environments where privacy, auditability and governance matter.

What to include with your application

  • A model you trained, fine-tuned, distilled or benchmarked, including the dataset, validation approach, metrics and your contribution.
  • An ML data, evaluation, deployment or inference system you owned, including scale, failure handling and measurable results.
  • Your current Indian city, earliest start date and interview availability.