SciGroveBeta
Genetics

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

Xin Luo, Yicheng Tao, Haoxuan Zeng, Suyuan Wang, Chenzi Ouyang, Meiqi Zhu, Kai Liu, Shuibing Chen, Jie Liu

Featured August 22, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

A new AI model called VOICE learns to guess which genes are active in single cells just by looking at cheap tissue pictures, combining direct guesses with finding similar cells to make super accurate predictions.

In depth
The paper introduces VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images. It achieves this by first aligning cell morphology representations from H&E images with transcriptomic embeddings in a shared latent space using contrastive learning. Subsequently, VOICE employs two complementary prediction branches: one directly regresses gene expression from morphology, and another retrieves expression from similar reference cells. A per-gene gated fusion mechanism then optimally combines these two predictions, adapting to how well each gene's expression is associated with visible morphology.

Key Takeaways

  • 1
    VOICE establishes a panel-independent vision-omics representation by contrastively aligning cell morphology from H&E images with single-cell expression embeddings, enabling generalization across diverse gene panels.
  • 2
    The model utilizes two complementary prediction branches—a direct regression branch for morphologically predictable genes and a retrieval branch for genes lacking strong morphological signals—to enhance predictive accuracy.
  • 3
    A per-gene gated fusion mechanism dynamically combines the outputs of these two branches, optimizing predictions based on the varying morphological predictability of individual genes.

Conceptual Flow

HIGH LEVEL
1
Aligning Images to Gene Profiles

The model first learns to match what a cell looks like in a picture with its hidden gene activity, then uses this match to guess gene levels in two smart ways, finally blending the guesses.

Tissue Image
Gene Activity Data
Learn Connections
Cell Appearance Info
Gene Profile Info
2
Better Gene Prediction

By combining image analysis and finding similar cells, the new method predicts gene activity much better than older ways, even for genes it hasn't seen before.

Old Prediction Methods
VOICE Model
Compare Accuracy
Lower Accuracy
Higher Accuracy

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.