SciGroveBeta
Cheminformatics

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Edward De Brouwer, Carl Edwards, Alexander Wu, Jenna Collier, Graham Heimberg, Xiner Li, Meena Subramaniam, Ehsan Hajiramezanali, David Richmond, Jan-Christian Hütter, Sara Mostafavi, Gabriele Scalia

Featured May 19, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

AssayBench tests how well AI models can predict which genes control specific cell behaviors by turning complex biological experiments into a simple ranking task for language models.

In depth
The paper introduces AssayBench, a large-scale benchmark designed to evaluate the ability of LLMs to predict phenotypic outcomes of genetic perturbations. By framing the task as gene ranking conditioned on natural language descriptions of experimental screens, the study provides a standardized testbed for assessing how well models generalize across diverse biological contexts using the AnDCG metric to account for assay-specific difficulty.

Key Takeaways

  • 1
    AssayBench provides a standardized, large-scale dataset of 1,920 CRISPR screens for evaluating phenotypic prediction.
  • 2
    The AnDCG metric enables fair performance comparison across heterogeneous biological assays by correcting for random baselines.
  • 3
    Frontier LLMs demonstrate strong zero-shot performance, though they remain significantly below the empirically estimated performance ceiling.

Conceptual Flow

HIGH LEVEL
1
Methodology

The system takes a description of a biological experiment and asks the AI to rank genes based on their expected impact.

Experiment Description
Rank Genes
Gene Priority List
2
Results

The AI's performance is measured against a theoretical limit to see how much room for improvement remains.

AI Predictions

Compare to Reality

Performance Gap