Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

Hello — A Brief Intro

3 minute read

Published:

A snapshot of who I was in my last year of high school, written when I first started this blog.

My First SAT — Test-taking Experience in Macau

less than 1 minute read

Published:

A short travelogue and logistics log from my first SAT, taken at St. Joseph University in Macau. (The original practice-correction screenshots live in my offline study archive and are omitted here.)

Notes on Calculus III — Vector Basics

1 minute read

Published:

My study notes on the vector foundations of Calculus III — length, dot product, cross product, and the inequalities and geometric interpretations that fall out of them.

portfolio

iGEM — Synthetic Biology (BNDS-China)

Two consecutive gold medals and top-10 high-school team at the International Genetically Engineered Machine competition, as team leader and instructor.
synthetic biology · genetic circuits · dry lab

Mathematical Modeling — IMMC & MCM

Outstanding (top <1%) in the 2025 IMMC Greater-China round and Meritorious internationally; MILP scheduling models built and typeset end-to-end.
MILP · PuLP · LaTeX

publications

Semi-Supervised High-Order Relation Learning for Drug–Combination–Disease Prediction

Manuscript in preparation, 2026

Predicting and analyzing higher-order drug-combination-disease relationships from a comprehensive dataset, using semi-supervised learning to exploit the large space of unlabeled combinations.

Recommended citation: Yonghan Yang, et al. Semi-Supervised High-Order Relation Learning for Drug–Combination–Disease Prediction. Manuscript in preparation, 2026.

A Survey on Discrete Diffusion Models

Survey · manuscript in preparation, 2026

A comprehensive review of the formulation, training, and applications of discrete diffusion models — unifying notation across masked, uniform, and absorbing-state processes and surveying their use in language, biology, and combinatorial design.

Recommended citation: Yonghan Yang, et al. A Survey on Discrete Diffusion Models. Manuscript in preparation, 2026.

Surrogate-Guided Memory Retrieval for Autonomous Agents

Manuscript in preparation, 2026

Optimizing what an autonomous agent recalls: we train memory retrieval offline with a learned surrogate that scores which past experiences most improve downstream task success, rather than relying on raw semantic similarity.

Recommended citation: Yonghan Yang*, Ye Yuan*, et al. Surrogate-Guided Memory Retrieval for Autonomous Agents. Manuscript in preparation, 2026. (* equal contribution)

Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization

Published in International Conference on Machine Learning (ICML), 2026

SPADE casts forward surrogate modeling as a calibrated conditional diffusion problem and adds a kNN support-proximity prior, so an offline optimizer stays expressive without exploiting unsupported, out-of-distribution regions. We prove the regularizer is equivalent to Bayesian inference under a valid design prior, and it tops Design-Bench and LLM-optimization tasks.

Recommended citation: Yonghan Yang*, Ye Yuan*, Zipeng Sun, Linfeng Du, Bowei He, Haolun Wu, Can Chen, and Xue Liu. (2026). "Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization." Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), PMLR 306, Seoul, South Korea. (* equal contribution)
Download Paper | Download Slides | Download Bibtex

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks

arXiv preprint, 2026

A biomedical ML coding benchmark testing whether agents can write task-specific model-building code for heterogeneous, often multi-modal biomedical data. 76 end-to-end tasks across 9 domains, each with hidden labels, held-out graders, and biology-aware metrics on a common 0–1 scale — agents must write runnable code, train models, and submit predictions on private test samples.

Recommended citation: Loka Li, Duzhen Zhang, Xingbo Du, et al. (including Yonghan Yang), Bin Zhang, and Le Song. (2026). "BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks." arXiv:2605.15766.
Download Paper

Herculean: An Agentic Benchmark for Financial Intelligence

arXiv preprint · under review at NeurIPS, 2026

Herculean evaluates AI agents on practical financial workflows across four domains — Trading, Hedging, Market Insights, and Auditing — each a standardized environment with its own tools and success criteria. Frontier agents do well on Trading and Market Insights but struggle badly with Hedging and Auditing, where long-horizon coordination, state consistency, and structured verification matter.

Recommended citation: Xueqing Peng, Zhuohan Xie, Yupeng Cao, et al. (including Yonghan Yang). (2026). "Herculean: An Agentic Benchmark for Financial Intelligence." arXiv:2605.14355. Under review at NeurIPS 2026.
Download Paper

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

arXiv preprint, 2026

Social deduction games are a popular testbed for reasoning and deception in LLM agents, but scoring only win rates cannot tell whether an agent’s words are grounded in what it actually saw and did. QUACK audits agents at three levels — outcomes, trajectories, and utterance-level consistency — and finds even the strongest VLM hallucinates 15.1% of its verifiable spatial claims.

Recommended citation: Ye Yuan, Rui Song, Weien Li, et al. (including Yonghan Yang), and Xue Liu. (2026). "QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents." arXiv:2605.27068.
Download Paper

talks

teaching