BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks

Published in arXiv preprint, 2026

Status: preprint on arXiv · My role: key contributions to the formulation of the paper — shaping how the benchmark was framed, which task families it should cover, and how agent performance is argued for and presented.

Paper (arXiv)

Summary

BioXArena asks a harder question than most agent benchmarks: not whether an agent can answer a biomedical question, but whether it can build a working model for one. Agents must write runnable code, train a model on a real dataset, and submit predictions for held-out private test samples.

  • 76 end-to-end tasks across 9 domains: sequence, single-cell, structure, network biology, chemical biology, perturbation dynamics, phenotype–disease, imaging, and text-integrated tasks.
  • Each task is curated from primary sources into a unified public capsule with hidden labels and held-out graders.
  • Biology-aware metrics are normalized onto a common 0–1 scale, so scores are comparable across very different data modalities.

Work with the GenBio AI group under Prof. Le Song.

Recommended citation: Loka Li, Duzhen Zhang, Xingbo Du, et al. (including Yonghan Yang), Bin Zhang, and Le Song. (2026). "BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks." arXiv:2605.15766.
Download Paper