Abhay Gupta

Abhay Gupta ML Engineer & AI Researcher

ML Engineer @ Kita (YC W26) · Research Intern @ Stanford AI Lab

Contact: abhaygupta1266@gmail.com

I build production AI systems and conduct research on large language models, alignment, fairness, and long-context reasoning. I am currently an ML Engineer at Kita (YC W26), where I build AI infrastructure for commercial lending, and a research intern at the Stanford Artificial Intelligence Laboratory (SAIL) working under Prof. Yejin Choi and Liwei Jiang.

My work spans production document intelligence, synthetic alignment data, counterfactual medical AI evaluation, dialect robustness, and multi-hop reasoning. It has appeared at EMNLP, NeurIPS, and AACL venues.

Large Language Models AI Fairness & Alignment Natural Language Processing

News

Selected Publications (* indicates equal contribution)

NovelHopQA Benchmark

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts

In collaboration with Meta and UC Berkeley

A Gupta*, K Zhu, V Sharma, S O'Brien, M Lu

Proceedings of the Association for Computational Linguistics: EMNLP 2025

Current LLMs struggle to answer questions spanning tens of thousands of tokens. We introduce NovelHopQA, an innovative benchmark to evaluate 1-4 hop QA over 64k-128k-token excerpts from 83 full-length public domain novels. Our research evaluates six SOTA models, revealing that scale alone does not guarantee robust multi-hop reasoning.

EnDive Benchmark

EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models

A Gupta*, J Cheung, P Meng, S Sayyed, K Zhu, A Liao, S O'Brien

Findings of the Association for Computational Linguistics: EMNLP 2025

EnDive (English Diversity) addresses the lack of intra-language evaluation in standard benchmarks. Translating SAE datasets into five underrepresented dialects via few-shot prompting, we created a challenging diagnostic that uncovers persistent model biases against non-standard language speakers across reasoning, logic, and math tasks.

AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark

A Gupta*, E Yurtseven, P Meng, K Zhu

Proceedings of the Third Workshop on NLP for Positive Impact, 2024

To develop more inclusive NLP systems, we introduced AAVENUE, a benchmark to evaluate LLMs on NLU tasks in African American Vernacular English (AAVE). The benchmark uses human-verified LLM translation to reliably map GLUE and SuperGLUE metrics out to AAVE.

MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning

R Ghosh*, A Gupta*, H McBride, A J Vaidya, F Mahmood

Proceedings of SciProdLLM, IJCNLP-AACL 2025

Evaluates how demographic cues influence clinical reasoning in frontier LLMs by holding critical symptoms constant while perturbing patient pronouns in 69,000 parallel test items, exposing localized divergences in downstream medical rationales.

CivicParse: A Benchmark and Pipeline for Structured Online Deliberation

A Gupta, M Klein

Accepted @ NeurIPS 2025 LLM Evaluation Workshop

Introducing a two-stage NLP extraction and classification pipeline to structure raw, free-form online discussions into a coherent deliberation map of core issues, barriers, and solutions. Outperforms prompt-only baselines and establishes a standardized task for modeling collective civic intelligence.

Experience

Industry & Startups

Cluely

ML Intern

Aug 2025 - Oct 2025 · Remote

  • Built production LLM systems for enterprise customers, designing custom AI assistants and prompt pipelines tailored to customer workflows and brand voice.
  • Developed a CLI-based LLM evaluation framework for prompt benchmarking and regression testing, improving pipeline accuracy and iteration speed by approximately 40%.
  • Collaborated directly with enterprise customers to design, deploy, and iterate AI solutions across real-world business workflows.

Mathos AI

ML Intern (YC W24)

May 2025 - Jul 2025 · Remote

  • Built AI-powered quiz and flashcard generation tools used by more than 1 million students, developing LLM pipelines for automated educational content creation and retrieval.
  • Improved content quality through prompt engineering, evaluation, and iterative testing while optimizing latency and scalability for production deployment.

Research

Stanford Artificial Intelligence Laboratory (SAIL)

Research Intern

Jul 2025 - Present · Remote

  • Developing PluralisticDataSmith under Prof. Yejin Choi: an open-source agentic framework for generating synthetic alignment data across diverse value systems.
  • Built the constitutional data-generation pipeline, filtering 243 value specifications and synthesizing a 1.2-million-example training corpus through model ensembling and iterative critique and refinement.
  • Designed a retrieval-augmented benchmark validated by 324 human annotators across more than 5,800 judgments; LoRA fine-tuning improved strict alignment pass rates by up to 14 percentage points. The work is supported by Schmidt Sciences' AI2050 program with 100,000 H100 GPU hours.

Harvard-MIT Health Sciences and Technology (HST)

Research Intern

Nov 2024 - Oct 2025 · Remote

  • Led MedEqualQA, a 69,000-example counterfactual benchmark evaluating reasoning stability and demographic robustness in medical LLMs.
  • Built large-scale data-generation and evaluation pipelines using minimally perturbed clinical narratives to isolate demographic variation and uncover hidden reasoning instabilities in frontier models.

Massachusetts Institute of Technology (MIT CCI Lab)

Research Intern

Nov 2024 - Dec 2025 · Remote

  • Collaborating with Prof. Mark Klein at the MIT Center for Collective Intelligence.
  • First author on CivicParse, an NLP extraction and classification pipeline bridging argumentation theory and AI to map large-scale deliberative discussions.

Algoverse

LLM Researcher

Dec 2023 - Aug 2025 · Remote

  • Developed AAVENUE, EnDive, and NovelHopQA, benchmarks for evaluating fairness, dialect robustness, and long-context reasoning in large language models.
  • Built evaluation pipelines spanning five English dialects and 64k-128k-token multi-hop reasoning tasks, uncovering systematic failures in frontier LLMs.
  • Research was accepted to EMNLP 2025 Main, EMNLP 2025 Findings, and EMNLP NLP4PI 2024 and has collectively received 43 citations from researchers at leading universities and industry labs.

Education

John Jay Senior High School

Hopewell Junction, New York · Graduated June 2026

SAT: 1550 (790 Math; 760 Reading & Writing)

Grants & Fellowships

AI2050 Compute Grant (100K hours of NVIDIA H100 compute)

Schmidt Sciences (Jan 2026)

Davidson Fellow Scholarship ($25,000)

Davidson Institute (Jul 2025)

Personal

My journey in AI began with a simple curiosity about how machines process human language, but it was Kevin Zhu who truly opened the door to the world of research for me. He was the very first person to guide my early steps, teaching me through everything and becoming my most impactful mentor. That initial curiosity quickly evolved into a passion for ensuring these systems are fair, inclusive, and safe for everyone. As I continue to grow, I am also incredibly grateful to Prof. Yejin Choi and Liwei Jiang for continuously inspiring me to tackle the difficult sociotechnical problems in AI alignment.

"Growth happens the moment you decide to step into the unknown. The challenges we face are not roadblocks, but the very stepping stones that build our resilience and character."

Outside of running evaluations and writing papers, I enjoy exploring the intersection of technology and linguistics, keeping up with the rapid pace of open-source AI, and finding new ways to make complex machine learning concepts accessible to my peers.