SciGroveBeta
Genetics

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Edward De Brouwer, Carl Edwards, Alexander Wu, Jenna Collier, Graham Heimberg, Xiner Li, Meena Subramaniam, Ehsan Hajiramezanali, David Richmond, Jan-Christian Hütter, Sara Mostafavi, Gabriele Scalia

Featured May 19, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Scientists created AssayBench, a new challenge for AI models to predict which genes will cause specific cell changes in experiments, using a special scoring system called AnDCG to fairly compare how well different AIs perform.

In depth
The paper introduces AssayBench, the first large-scale benchmark for in silico phenotypic screen prediction, addressing a critical gap in drug discovery. It formalizes the task as a gene ranking problem for 1,920 CRISPR screens and proposes adjusted nDCG (AnDCG), a novel metric that enables fair performance comparison across diverse biological assays.

Key Takeaways

  • 1
    The authors introduce AssayBench, the first large-scale benchmark for phenotypic screen prediction, comprising 1,920 publicly available CRISPR screens.
  • 2
    The task is formalized as a gene ranking problem for each screen, conditioned on a free-text description, mirroring real-world screening workflows.
  • 3
    A novel evaluation metric, adjusted nDCG (AnDCG), is proposed to enable continuous and comparable performance assessment across heterogeneous assays by correcting for screen-specific random baselines.

Conceptual Flow

HIGH LEVEL
1
Methodology: How was it done?

The scientists gathered many cell experiment results and turned them into a game where AI models guess the most important genes, then used a special score to see how good the guesses were.

Many Cell Experiments
Process & Organize
AI Challenge: Rank Genes
Fair Scoring System
2
Results: What did they find?

They found that even the best AI models are still far from perfect at predicting cell changes, but general-purpose AIs did better than specialized ones.

AI Models Try Ranking
Compare Performance
AIs Not Perfect Yet
General AIs Lead