Measure what matters
Evaluation strategies for agentic, multimodal, reasoning, and long-context models—designed to explain behavior, not just produce a score.
sainyam@skynet:~$ building-at-the-frontier
Principal Engineer (Acting) · Frontier AI Systems
I build the evaluation, data, and inference systems that turn frontier foundation models into reliable products — from agentic benchmarks and RL environments to production delivery at Bedrock scale.
At Amazon Nova, I lead systems that measure, improve, and ship foundation models. My work spans verifier-backed RL data, agentic evaluation, pre-production inference, and model delivery — turning science goals into repeatable engineering systems used across teams.
Amazon Nova · AGI Foundations
Amazon Nova · AGI Foundations
Amazon · Alexa NU
Facebook AI Research (FAIR)
Amazon · Alexa Info Federation
Trexquant Investment LP
Evaluation strategies for agentic, multimodal, reasoning, and long-context models—designed to explain behavior, not just produce a score.
Evaluation platforms, RL data pipelines, and verifier-backed quality systems that help researchers iterate with speed and confidence.
Pre-production inference and model delivery that bridge research architectures with reliable, observable production systems.
IIT (BHU), Varanasi
I enjoy exchanging notes on evaluation design, agentic systems, RL environments, and the engineering it takes to move models from research to production.
Start a conversation