# Sainyam Kapoor

Principal Engineer (Acting) · Frontier AI Systems

I build the evaluation, data, and inference systems that turn frontier foundation models into reliable products — from agentic benchmarks and RL environments to production delivery at Bedrock scale.

- Email: <hello@sainyam.me>
- GitHub: https://github.com/saikpr
- LinkedIn: https://linkedin.com/in/sainyamkapoor
- Blog: https://blog.sainyam.me/
- Résumé: https://link.sainyam.me/resume

## Selected impact

At Amazon Nova, I lead systems that measure, improve, and ship foundation
models. My work spans verifier-backed RL data, agentic evaluation,
pre-production inference, and model delivery — turning science goals into
repeatable engineering systems used across teams.

| Metric | What it covers |
| --- | --- |
| Nova 1 & 2 | Launch and benchmarking leadership |
| 50+ | Agentic benchmarks enabled |
| 3 wks → 3 days | Post-training-to-release cycle |
| 1000s of evals | Across 100s of benchmarks daily |

## Experience

### Principal Engineer (Acting) — Amazon Nova · AGI Foundations

_Jan 2026 — Present_

- Leading **RL Gym** data and validation systems for simulated app environments,
  producing verifier-backed interaction traces for model training.
- Building one extensible evaluation harness for agentic and non-agentic
  workloads, enabling 50+ agentic benchmarks with context engineering and
  benchmark-specific adaptation.

### Senior Manager, AI — Amazon Nova · AGI Foundations

_Jun 2023 — Jan 2026_

- Owned end-to-end evaluation for Amazon Nova and Titan, from pre-training
  through runtime, across agentic, multimodal, reasoning, and generation
  workloads.
- Helped launch **Nova 1** at re:Invent 2024 and led **Nova 2** benchmarking for
  re:Invent 2025, incorporating real workloads from 20+ customers.
- Co-created **OneClickEval**, eliminating thousands of hours of weekly manual
  evaluation and reducing the model-feedback loop from two weeks to 1.5 days. It
  grew to 500+ metrics and 300+ datasets used by 300+ scientists.
- Led pre-production inference and Bedrock delivery across 30+ engineers,
  supporting multimodality, tools, extended thinking, and 1M-token context at
  100M+ daily API-call scale.
- Reduced the post-training-to-release cycle from three weeks to **three days**
  by aligning teams on a common model specification and standardized evaluation.

### Engineering Manager (SDE → Sr. SDE → EM) — Amazon · Alexa NU

_Nov 2019 — Jun 2023_

- Led delivery of Alexa's context-aware NLU stack — +5% intent prediction in
  English, **+14%** across FR, HI, ES, JP and more.
- Deployed large-transformer models at **sub-100ms latency** across 100M+ active
  devices and billions of daily utterances.
- Cut ML-to-prod launch time from 3–4 weeks to 3 days via an ECS service with
  automated safeguards and A/B rollouts. 2 patents filed, 1 granted.

### AI Research Resident — Facebook AI Research (FAIR)

_Oct 2018 — Sep 2019_

- Built graph embedding models using BERT representations to study
  entity-content relationships and social influence at Facebook scale — **+6%**
  content integrity detection.

### Software Development Engineer — Amazon · Alexa Info Federation

_Aug 2017 — Sep 2018_

- Improved Alexa's knowledge-answering accuracy by **+4%** via utterance
  paraphrasing across 50M daily utterances; containerized the ML-to-prod A/B
  rollout pipeline.

### Global Alpha Researcher — Trexquant Investment LP

_Jan 2017 — Aug 2017_

- Built trading signals for medium-frequency trading.

## Expertise

1. **Measure what matters** — evaluation strategies for agentic, multimodal,
   reasoning, and long-context models, designed to explain behavior, not just
   produce a score.
2. **Shorten the learning loop** — evaluation platforms, RL data pipelines, and
   verifier-backed quality systems that help researchers iterate with speed and
   confidence.
3. **Ship the model** — pre-production inference and model delivery that bridge
   research architectures with reliable, observable production systems.

Working toolkit: Python, PyTorch, vLLM, SGLang, TensorRT-LLM, AWS Neuron,
Bedrock, ECS, Kubernetes.

## Education

**M.Tech & B.Tech, Computer Science & Engineering** — IIT (BHU), Varanasi.
8.5 CGPA · May 2017.

## Contact

I enjoy exchanging notes on evaluation design, agentic systems, RL environments,
and the engineering it takes to move models from research to production.

Email: <hello@sainyam.me>

---

Machine-readable resources: [/llms.txt](https://sainyam.me/llms.txt) ·
[/profile.jsonld](https://sainyam.me/profile.jsonld) ·
[/.well-known/api-catalog](https://sainyam.me/.well-known/api-catalog) ·
[/robots.txt](https://sainyam.me/robots.txt)

Content policy: `search=yes, ai-input=yes, ai-train=yes` (see /robots.txt).
