Yash Maurya

AI Safety & Evaluation

I'm a Research Engineer at Scale AI, building post-training, evaluation, and red-teaming systems for frontier language models, with a focus on alignment, safety, privacy, and dependable behavior in enterprise deployments.

I care about making advanced models trustworthy in production, especially under adversarial and edge-case behavior. My work sits at the intersection of model capability, policy constraints, and operational risk.

I completed my master's at Carnegie Mellon University in Privacy Engineering, where I worked on privacy, LLM safety, and evaluation research across academia and industry collaborations.

Previously, I worked on AI governance and evaluation at BNY, and earlier on large-scale ML and privacy-preserving systems at Samsung and Dynamo AI.

Here's my resume in case you need it.

Want to chat? Send me an email or text on LinkedIn!

Privacy is not something that I'm merely entitled to, it's an absolute prerequisite.

Recent News

Research

* indicates equal contribution

2026

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue
arXiv preprint arXiv:2608.06301 · 2026
TL;DR: Benchmarks whether frontier LLMs can improve an agent's prompts, tools, memory, and orchestration under a fixed evaluation budget. Results show that optimizer models matter more than the coding harness they use, with gains varying widely by task and seed.
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
Akshay Manglik, Apaar Shanker, Kaustubh Deshpande, Jason Qin, Yash Maurya, Veronica Chatrath, Vijay S. Kalmath, Levi Lentz, Yuan Xue
arXiv preprint arXiv:2605.21347 · 2026
TL;DR: A multi-agent system that turns large corpora of LLM-agent traces into evidence-backed diagnostic insights. Experts using its reports improved scaffold performance by 30.4 percentage points over the unmodified baseline.
Yu Ying Chiu*, Michael S. Lee*, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Raphaël Millière, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q. Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L. Gordon, Sydney Levine
International Conference on Learning Representations (ICLR 2026)
TL;DR: Evaluates model reasoning on morally ambiguous scenarios, where multiple conclusions can be defensible. Expert rubrics test whether models identify trade-offs, justify recommendations, and reason across normative ethics frameworks.
LHAW: Controllable Underspecification for Long-Horizon Tasks
George Pu, Michael S. Lee, Udari Madhushani Sehwag, David J. Lee, Bryan Zhu, Yash Maurya, Mohit Raghavendra, Yuan Xue, Samuel Marc Denton
Lifelong Agents: Learning, Aligning, Evolving Workshop at ICLR 2026
TL;DR: Builds controllably underspecified versions of long-horizon agent tasks and validates their ambiguity through actual agent execution. The 285-task release measures whether agents seek clarification when missing information is genuinely outcome-critical.
Michael S. Lee*, Yash Maurya*, Drew Rein, Bert Herring, Jonathan Nguyen, Kyungho Song, Udari Madhushani Sehwag, Jiyeon Cho, Kaustubh Deshpande, Yeongkyun Jang, Jiyeon Joo, Minn Seok Choi, Evi Fuelle, Christina Q. Knight, Joseph Brandifino, Max Fenkell
TAIGR Workshop at ICML 2026 · Best Paper Award
TL;DR: A bilingual safety benchmark that separates language effects from geopolitical grounding through an English-Korean transcreation matrix. It exposes safety and over-refusal behaviors that translation-only evaluations miss.

2025

Yash Maurya, Ibrahim Mohamed Anis Chhaya, Hana Habib
IEEE Symposium on Privacy Expectations (ISoPE) 2025 & SUPA 2025 Workshop on Societal & User-Centered Privacy in AI
TL;DR: Practitioner-oriented framework that organizes concrete ML privacy mitigations, tools, and design patterns across the ML lifecycle to help teams operationalize privacy-preserving AI in real-world deployments.
When Privacy Guarantees Meet Pre-trained LLMs: A Case Study in Synthetic Data
Yash Maurya*, Aman Priyanshu*
2025 USENIX Conference on Privacy Engineering Practice and Respect (PEPR'25)
TL;DR: Shows how document formatting and contextual patterns can create privacy leakage in differentially private synthetic data pipelines built on opaque pre-trained LLMs, even at conservative privacy budgets.

2024

Beyond the Accept Button: How Information and Control Shape Data Sharing and AI Engagement
Ibrahim Chhaya*, Yash Maurya*, Zuofei Hong*, Limin Ge*
Sponsored by Meta - MSIT-PE Capstone Report 2024 (Advisor: Professor Lorrie Cranor)
TL;DR: Studied how consent flow design impacts AI engagement and data sharing, examining length and control options in social media data sharing.
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Pratiksha Thaker, Shengyuan Hu, Neil Kale, Yash Maurya, Zhiwei Steven Wu, Virginia Smith
IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2025
TL;DR: Shows that modest, benign changes to common unlearning benchmarks can expose recoverable target information or substantially worse retention loss, making reported progress look more reliable than it is.
Designing a Benefit Assessment Protocol for AI Systems
Rachel Kim*, Yash Maurya*, Goutam Mukku*
Course Project for Responsible AI Course(10-735) at CMU (Advisor: Professor Hoda Heidari)
TL;DR: A structured protocol for systematically assessing AI benefits to enable more comprehensive AI evaluation.
Unified Locational Differential Privacy Framework
Aman Priyanshu*, Yash Maurya*, Suriya Ganesh*, Vy Tran*
arXiv preprint arXiv:2405.03903 (Advisor: Professor Hana Habib)
TL;DR: A privacy framework for aggregating sensitive location-based data while protecting individual privacy through differential privacy mechanisms.
AI Governance and Accountability: An Analysis of Anthropic's Claude
Aman Priyanshu*, Yash Maurya*, Zuofei Hong*
arXiv preprint arXiv:2407.01557 (Advisor: Professor Norman Sadeh)
TL;DR: Case study on Anthropic's Claude using AI governance and accountability frameworks, examining compliance with NIST and EU AI Act standards.
Guardrail baselines for unlearning in LLMs
Pratiksha Thaker, Yash Maurya, Shengyuan Hu, Zhiwei Steven Wu, Virginia Smith
ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
TL;DR: Simple prompting and filtering baselines can match fine-tuning on common unlearning evaluations, motivating stronger metrics that distinguish inference-time safeguards from genuine parameter-level forgetting.
Xinran Alexandra Li*, Yu-Ju Yang*, Yash Maurya*, Tian Wang*, Hana Habib, Norman Sadeh, Lorrie Faith Cranor
Twentieth Symposium on Usable Privacy and Security (SOUPS 2024 Posters) & SOUPS 2024 Societal & User-Centered Privacy in AI Workshop (SUPA 2024)
TL;DR: UsersFirst taxonomy outperforms LINDDUN PRO in detecting privacy notice and choice threats in user study.
Tian Wang*, Xinran Alexandra Li*, Miguel Rivera-Lanas*, Yash Maurya*, Hana Habib, Lorrie Faith Cranor, Norman Sadeh
Twentieth Symposium on Usable Privacy and Security (SOUPS 2024 Posters) & SOUPS 2024 Workshop on Privacy Threat Modeling (WPTM 2024)
TL;DR: UsersFirst: A user-centric framework for identifying and mitigating privacy notice and choice threats, extending beyond LINDDUN
Through the Lens of LLMs: Unveiling Differential Privacy Challenges
Aman Priyanshu*, Yash Maurya*, Vy Tran*
2024 USENIX Conference on Privacy Engineering Practice and Respect(PEPR'24)
TL;DR: LLMs demonstrate stronger privacy attacks on Google's Topics API, bypassing differential privacy safeguards.
Is it Worth Storing Historical Gradients?
Joong Ho Choi*, Yingxin Liu*, Yash Maurya*
Course Project for Federated and Collaborative Learning Course(10-719) at CMU (Advisor: Professor Virginia Smith)
TL;DR: Current weights beat historical gradients for detecting FL attacks, saving storage and enhancing privacy.

2022

Federated Learning for Colorectal Cancer Prediction
Yash Maurya*, Prahaladh Chandrahasan*, G Poornalatha
2022 IEEE 3rd Global Conference for Advancement in Technology (GCAT), 1-5
TL;DR: Federated learning enables privacy-preserving colorectal cancer prediction across hospitals with centralized-level accuracy

2021

Rakshit Naidu, Haofan Wang, Soumya Snigdha Kundu, Ankita Ghosh, Yash Maurya, Shamanth R Nayak K, Joy Michael
Responsible Computer Vision (RCV) Workshop at CVPR 2021
TL;DR: Slightly modified version of IS-CAM (described below)

2020

IS-CAM: Integrated Score-CAM for axiomatic-based explanations
Rakshit Naidu, Ankita Ghosh, Yash Maurya, Shamanth R Nayak K, Soumya Snigdha Kundu
arXiv preprint arXiv:2010.03023
TL;DR: Enhanced CNN interpretability through IS-CAM, integrating Score-CAM to produce sharper attribution maps