April Yang

April Yang

Hi, I'm April! I'm a research engineer at NVIDIA, working on multimodal understanding and agentic search, currently with a domain focus in autonomous vehicles but in general interested in Physical AI.

Outside work, I enjoy photography, traveling, driving with music, and spending time with my three cats.

Current Work

NOW

Multimodal search and data flywheels for physical AI

We have taken VLM-based video search from prototype to production: models caption driving clips, and the captions become embeddings for semantic retrieval over a petabyte-scale video datalake. We build the retrieval foundation to support data flywheel that powers end to end training loops.

Training and evaluating VLMs for AV scene understanding

We post-train video VLMs for driving-scene understanding. My focus is the evaluation side: metrics and benchmarks covering caption quality, temporal reasoning, retrieval behavior, and how well caption embeddings separate driving scenarios, so we can compare fine-tuned checkpoints with confidence.

Agent loops that curate long-tail driving scenarios

We're building an agentic search-refinement loop where a rewriter reformulates the query, a search agent retrieves candidates, a judge verifies them, and a refiner sharpens the query until precision converges. It replaces a curation process that used to take hours of manual search and review per scenario, with an evaluation harness that keeps the agents measurable, reproducible, and self-improving.

Publications

SELECTED
CVPRW 2026

Scalable Parallel Prompting for Complex AV Video Captioning

April Yang*, Roberto Amoroso*, Nikita Durasov, Devansh Bisla, Sandipan Kundu, Elmar Haussmann, Ruchi Bhargava, Maying Shen, Nadine Chang, Jose M. Alvarez

Small VLMs caption the scene, road entities, and key driving actions in parallel, then consolidate everything into one structured caption, curbing the compounding hallucinations and cost of large-model captioning and improving top-10 retrieval precision by up to 11 points.

Whitepaper 2026

SIL-Wheel: A Multi-Modal Search and Curation Platform for Physical AI Systems

NVIDIA (core contributor)

An open-source system for large-scale video data curation that fuses complementary retrieval signals over tens of millions of clips, closing the loop from model failure discovery to training and evaluation.

Experience

EXPERIENCE
2025 - Present
NVIDIA Corporation
Research Engineer

Vision-language model post-training and evaluation, agentic multimodal search, and large-scale data curation for autonomous vehicles

2024
NVIDIA Corporation
Software Engineer Intern, LLMs

Multi-agent RAG system for AV engineering Q&A, test generation, and code generation, deployed on Kubernetes and used by thousands of engineers

2023
AsyncHealth
Software Engineering Intern
2022
Moody's Investors Service
Data Engineer Intern

APAC Corporate Finance Group

2021 - 2022
China International Capital Corporation (CICC)
Risk Management Intern

CICC Wealth Management

2021
Ant Group
Data Analyst Intern

Consumer Credit Business Group

2021
ByteDance
Data Scientist Intern

Real Estate Product and Strategy Analysis Division

EDUCATION
2026 - Present
Graduate Coursework, Robotics & Autonomous Systems
Stanford University
2023 - 2024
Master of Science, Software Engineering
Carnegie Mellon University
2022 - 2023
Master of Engineering, Industrial Engineering & Operations Research
(FinTech)
University of California, Berkeley
2018 - 2022
Bachelor of Science, Management Science
(Business Analytics)
Renmin (People's) University of China
2021
Visiting Student, Statistics & Data Science
Columbia University,
Columbia College
2019
Summer Student, Data Mining
Harvard University
Download CV