sainyam@skynet:~$ building-at-the-frontier

Sainyam Kapoor

Principal Engineer (Acting) · Frontier AI Systems

I build the evaluation, data, and inference systems that turn frontier foundation models into reliable products — from agentic benchmarks and RL environments to production delivery at Bedrock scale.

Portrait of Sainyam Kapoor
// selected impact

Turning model capability into dependable products

At Amazon Nova, I lead systems that measure, improve, and ship foundation models. My work spans verifier-backed RL data, agentic evaluation, pre-production inference, and model delivery — turning science goals into repeatable engineering systems used across teams.

Nova 1 & 2 Launch and benchmarking leadership
50+ Agentic benchmarks enabled
3 wks → 3 days Post-training-to-release cycle
1000s of evals Across 100s of benchmarks daily
// experience

A decade of applied AI, from research to production

  1. Principal Engineer (Acting)

    Jan 2026 — Present

    Amazon Nova · AGI Foundations

    • Leading RL Gym data and validation systems for simulated app environments, producing verifier-backed interaction traces for model training.
    • Building one extensible evaluation harness for agentic and non-agentic workloads, enabling 50+ agentic benchmarks with context engineering and benchmark-specific adaptation.
  2. Senior Manager, AI

    Jun 2023 — Jan 2026

    Amazon Nova · AGI Foundations

    • Owned end-to-end evaluation for Amazon Nova and Titan, from pre-training through runtime, across agentic, multimodal, reasoning, and generation workloads.
    • Helped launch Nova 1 at re:Invent 2024 and led Nova 2 benchmarking for re:Invent 2025, incorporating real workloads from 20+ customers.
    • Co-created OneClickEval, eliminating thousands of hours of weekly manual evaluation and reducing the model-feedback loop from two weeks to 1.5 days. It grew to 500+ metrics and 300+ datasets used by 300+ scientists.
    • Led pre-production inference and Bedrock delivery across 30+ engineers, supporting multimodality, tools, extended thinking, and 1M-token context at 100M+ daily API-call scale.
    • Reduced the post-training-to-release cycle from three weeks to three days by aligning teams on a common model specification and standardized evaluation.
  3. Engineering Manager (SDE → Sr. SDE → EM)

    Nov 2019 — Jun 2023

    Amazon · Alexa NU

    • Led delivery of Alexa's context-aware NLU stack — +5% intent prediction in English, +14% across FR, HI, ES, JP and more.
    • Deployed large-transformer models at sub-100ms latency across 100M+ active devices and billions of daily utterances.
    • Cut ML-to-prod launch time from 3–4 weeks to 3 days via an ECS service with automated safeguards and A/B rollouts. 2 patents filed, 1 granted.
  4. AI Research Resident

    Oct 2018 — Sep 2019

    Facebook AI Research (FAIR)

    • Built graph embedding models using BERT representations to study entity-content relationships and social influence at Facebook scale — +6% content integrity detection.
  5. Software Development Engineer

    Aug 2017 — Sep 2018

    Amazon · Alexa Info Federation

    • Improved Alexa's knowledge-answering accuracy by +4% via utterance paraphrasing across 50M daily utterances; containerized the ML-to-prod A/B rollout pipeline.
  6. Global Alpha Researcher

    Jan 2017 — Aug 2017

    Trexquant Investment LP

    • Built trading signals for medium-frequency trading.
// expertise

What I bring to frontier-model teams

01

Measure what matters

Evaluation strategies for agentic, multimodal, reasoning, and long-context models—designed to explain behavior, not just produce a score.

02

Shorten the learning loop

Evaluation platforms, RL data pipelines, and verifier-backed quality systems that help researchers iterate with speed and confidence.

03

Ship the model

Pre-production inference and model delivery that bridge research architectures with reliable, observable production systems.

Working toolkit
  • Python
  • PyTorch
  • vLLM
  • SGLang
  • TensorRT-LLM
  • AWS Neuron
  • Bedrock
  • ECS
  • Kubernetes
// education

Education

M.Tech & B.Tech, Computer Science & Engineering

IIT (BHU), Varanasi

8.5 CGPA · May 2017

// contact

Working on a hard model-systems problem?

I enjoy exchanging notes on evaluation design, agentic systems, RL environments, and the engineering it takes to move models from research to production.

Start a conversation