Biostatistics · Harvard Chan School

Haoyu Yang

Postdoctoral Research Fellow in Biostatistics. I build methods that extract more discovery from the data we already have — and prove they stay valid even when the AI, the borrowed population, or the fitted model is wrong.

Currently, I'm on the job market. Please don't hesitate to contact me if you are interested in my research or would like to share any ideas 🥳!

Portrait of Haoyu Yang
01

About

I am a Postdoctoral Research Fellow in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health, working with Professor Xihong Lin.

My work starts from one observation: every dataset contains more information than a standard analysis extracts.

Some of it sits in missing measurements — the 90% of a biobank whose proteome was never assayed, though its genotypes and clinical records are complete. Some sits in heterogenous populations - the target population is too small, and the available source populations are too heterogeneous to combine. Some is discarded by the analysis itself, when half a million hypothesis tests are evaluated one at a time, and everything the ensemble knows is thrown away. And some information is ignored when participants are enrolled, by a design that leaves balance to chance.

I build methods that convert this potential information into scientific discovery. The difficulty is that each of these sources is untrustworthy in its own specific way — a generative model can be misspecified, another population can carry different biology, a composite null distribution is hard to characterize — which is why the interesting part of the problem is statistical rather than computational.

Before Harvard, I earned my Ph.D. in Statistics at Renmin University of China.

Position
Postdoctoral Research Fellow
Affiliation
Harvard T.H. Chan School
Department of Biostatistics
Location
Boston, USA
02

Education & Appointment

Jul 2023 — Present
Postdoctoral Research Fellow
Harvard University · Department of Biostatistics — Boston, USA
Advisor: Xihong Lin
Sep 2018 — Jun 2023
Ph.D. in Statistics
Renmin University of China — Beijing, China
Advisor: Yang Li
Sep 2014 — Jun 2018
B.S. in Mathematics & Applied Mathematics
Tianjin University — Tianjin, China
Minor in Financial Management
03

Research

Biomedical research routinely discards four critical sources of information: unmeasured samples, excluded populations, the hidden structure of large-scale testing, and the study’s own design. My research builds methodological frameworks for each, united by a single guarantee: borrowing strength must never cost you the error rate.

01 · Multiplicity

Multiplicity is a blessing.

When a study runs half a million tests, the ensemble of statistics is enormously informative about which hypotheses are true. I turn the multiplicity into the thing that makes the test nearly optimal.

M-DACT & Tail likelihood ratio test
02 · Synthetic data

AI-generated data is Useful, not necessarily true.

90% of a biobank has no proteomic assay but complete genotypes and records. Rather than substituting predictions for data, I model them jointly inside a mixed model — AI model misspecification costs power, never validity.

Syn-PALM
03 · Populations

Safely borrowing strength across heterogeneous populations.

A minority cohort is underpowered while a cohort a hundred times larger sits beside it. Pooling naively is worse than not borrowing at all; I decompose the target null instead of assuming transportability away.

HEART-GWAS
04 · Design

Decide how data can be collected.

Before anyone is enrolled, the design fixes how much a study will ever reveal. Covariate-adaptive randomization and design strategies for networked experiments help to get better treatment assignment.

Adaptive randomization & Network interference

“The auxiliary information is allowed to be arbitrarily wrong.”

This is the rule I hold myself to in my project. Validity — type I error, false discovery rate, unbiasedness for the target estimand — is established by derivation, never inherited from an assumption that the model, the borrowed population, or the imputation happens to be correct. Accuracy buys power, and nothing else. The result is a guarantee worth stating plainly: provably never worse than the classical analysis that throws the extra information away, and strictly better whenever that information is real.

Methodology

Inference with AI-generated Data Large-scale Composite-null Testing Cross-population Transfer Learning Covariate-adaptive Randomization Causal Inference under Interference

Applications

Genome-wide Association Studies Proteomics & pQTL Discovery Epigenome-wide Studies Population Biobanks Clinical Trials Networked & Economic Experiments
Where I'm headed

Over the next five years I want to build a general theory, and a usable software stack, for turning imperfect auxiliary information into valid scientific discovery — together with adversarial benchmarks, because this literature still has no shared standard for demonstrating that a method fails safely.

Mediation with AI-generated mediators

We can measure exposures and outcomes on everyone, but the proteomic and epigenomic layers in between on almost no one — so most candidate pathways stay permanently unexplored. Calibrated synthetic mediators, propagated through an augmented-IPW construction and tested by a synthetic tail likelihood ratio test, so pathway discovery stops being capped by assay budgets.

A unified theory of auxiliary evidence

A prediction model, a heterogeneous population, a historical control arm, an external registry — all the same formal object. Decorrelate it from the target, decompose the null into the configurations it can misrepresent, and let an ensemble-estimated mixture decide how far to trust it, with how much to borrow answered decision-theoretically rather than by convention.

Foundation models as measurement instruments

Single-cell, protein-language and clinical models are becoming primary scientific measurement devices — black-box, silently updated, and differently calibrated across the very subpopulations we study. Conformal layers under distribution shift, drift diagnostics, and selective-inference corrections: an end-to-end reliability audit for AI-assisted discovery.

04

Publications

* corresponding author; † contributed equally to the first author.

Book Chapters

  • Variable Selection for Missing Data. In: Modern Survey Analysis.
  • Statistical Principles of Data Science. In: Introduction to Data Science.
05

Talks & Awards & Teaching & Service

Selected Presentations
Nov 2025 · Mar 2026 · Aug 2026
Robust synthetic data–assisted genome-wide association studies boost proteome-wide genetic discovery with partially observed data in biobanks
Statistical Genetics Meeting; ILCCO Biostatistics Meeting; Joint Statistical Meetings — Boston, USA
Aug 2024
Tail likelihood ratio method for large-scale causal mediation hypothesis testing
Joint Statistical Meetings — Portland, USA
Jun 2024
Causal mediation analysis for integrating exposure, genomic and phenotype data
Statistical Genetics Meeting — Boston, USA
May 2024
Design strategies for networked experiments via interference balancing
37th New England Statistics Symposium — Connecticut, USA
Jun 2021
Adaptive randomization via Mahalanobis distance
Ph.D. Forum on Biostatistical Methods & Applications — Beijing, China
Dec 2020
Balancing covariates in multi-arm trials via adaptive randomization
6th Academic Seminar, Beijing Biomedical Statistics & Data Management Research Association
Honors & Awards
2026Poster Award, 2026 ICSA Applied Statistics Symposium
2022Top Ten Papers, 8th National Postgraduate Statistics Forum
2021Top Ten Papers, 2nd Data Science and Modern Economic Statistics Forum
2020First Prize, "BeiGene" Excellent Paper for Youth
Teaching
2021Teaching Assistant — Complex Statistical Analysis
2021Teaching Assistant — Probability Theory
2020Teaching Assistant — Introduction to Data Science
Service
Journal reviewer — Statistica Sinica, Biostatistics, Journal of Data Science
06

Beyond the Data

Always in motion — on the water, on the court, on the snow. Away from the data, I'm happiest when my heart rate is up. I run marathons, play tennis, swim, and carve down ski slopes; I love the rhythm of rowing and the pull of a good race. And because the best moments are shared, I'm all in on team sports too — volleyball, dragon boat, and ultimate frisbee. Whatever the season, I'll keep moving.