Undergraduate · NYU Shanghai

Zhuoyan Yu.

I work on language models for code
and how we evaluate them.

At NYU Shanghai, I work on code generation, LLM-generated tests, and behavior-preserving repair of research software. I also work on computer vision, with experience in medical imaging and multimodal systems.

Selected research

Publication
Explore the CodeFORGE framework
CodeFORGE framework: task specification, candidate generation, sandbox verification, benchmark construction, and model evaluation.
Framework overview from the paper, Figure 2. Candidate tasks pass through isolated execution and feedback-driven verification before entering the benchmark.

Research & experience

2025 — present
Sep 2026 — presentOngoing research

RepoRevive: Behavior-Preserving Repository Repair

Advanced Capstone Project · NYU Shanghai

Building tooling for behavior-preserving repair of research code. Reconstructed historical and modern environments to compare model outputs, gradients, and one-step training updates during a PyTorch 1.4-to-2.8 migration.

An API-only migration ran all 14 configurations but preserved behavior in 10; a reference patch passed all 14 behavior checks.

Jun 2026 — presentOngoing research

LLM Test Reliability & Cross-Model Fault Detection

Serendipity Research Group · NYU Shanghai

Studying how same-model and cross-model tests expose faults in fixed code candidates, with and without access to candidate code. Built isolated execution and a proxy oracle requiring agreement among three qualified reference programs.

A 36-problem expansion yielded 52 qualified faulty programs across 21 problems for strict test-strategy comparisons; fault detection and structural coverage are analyzed separately.

May — Aug 2026Research internship

Ant Technologies U.S.

Algorithm Engineer Intern (Research) · Medical and Healthcare AI Lab

Developed nnU-Net v2 pipelines for cerebral microbleed and lacune detection in multimodal MRI. Worked on spatial alignment, data validation, controlled ablations, and lesion-level evaluation.

Sep 2025 — Apr 2026Research assistant

CodeFORGE

Serendipity Research Fellowship · NYU Shanghai

Built benchmark generation and verification tooling for rare and legacy programming languages, including Docker-based pass@k evaluation, feedback-driven resampling, and API cost controls.

Earlier research includes long-video understanding and chest X-ray interpretation. See the full CV

Selected projects

Vision & multimodal systems
Qualitative examples from the MYGO report: baseline and MYGO generated images, conditioning depth, and predicted depth.
Qualitative comparison · project report, Fig. 3

Course project Sep — Dec 2025

MYGO: Geometry-Aware Depth ControlNet

Exploring how depth representation affects image generation. Co-designed an 18-channel conditioning signal for Stable Diffusion 1.5 ControlNet, combining raw depth, 16 ordinal layers, and occlusion boundaries.

Compared with a single-channel baseline on 50k COCO images under a matched 20k-step training budget.

Method illustration: OCR, segmentation, and visual language understanding convert paper figures into editable slide elements.
Method illustration

Research project May — Aug 2025

Agentic Figure Transcription for Scientific Papers

A prototype that turns academic figures into editable PowerPoint slides. Combines OCR, SAM segmentation, region-level vision-language classification, and recursive segmentation of small objects.

Research Assistant · CSDSE Department, NYU Shanghai

Get in touch

Let’s talk research.

For research conversations and collaboration, email is the best way to reach me.