In this comprehensive review, the authors critically evaluate the transformative impact of foundation models and vision-language systems on computational pathology, framing their analysis through a multiscale lens—from cellular morphology and tissue microenvironments to whole-slide image (WSI) analysis and patient-level multimodal integration. Moving beyond traditional task-specific deep learning solutions, the review examines how these emerging data-driven systems are redefining pathological analysis across diverse spatial scales and bridging the gap between pixel-level features and patient-centric clinical decision-making.
Key highlights
A paradigm shift in computational pathology — The review documents the evolution from fragmented, supervised convolutional neural networks to unified, general-purpose foundation models and vision-language systems that leverage massive unlabeled datasets for robust, transferable feature representations. Unlike their predecessors, these next-generation systems transcend specific organs or stains, enabling cross-scale reasoning from nuclear atypia to WSI-level prognostic profiling.
Comprehensive multiscale framework — The authors systematically survey applications across four biological hierarchies: cell-level (segmentation and phenotyping), tissue/region-level (microenvironment modeling and spatial architecture), WSI-level (weakly supervised prediction and retrieval), and patient-level (multimodal integration of histology, genomics, radiology, and clinical records). Representative architectures, including attention-based multiple instance learning (MIL), hierarchical vision transformers, and graph neural networks are critically assessed for their roles at each scale.
Foundation models as unifying enablers — The review provides an in-depth analysis of pathology foundation models (PFMs) such as Virchow, UNI, CHIEF, GigaPath, and TITAN, which construct unified representation spaces where multiscale features—from nuclear morphology to tumor microenvironment topology—are mapped onto a common manifold. These models support diverse downstream tasks, including tumor subtyping, mutation prediction, and prognosis assessment without task-specific architectural modifications.
Vision-language systems bridging semantic gaps — The emergence of pathology-specific vision-language models (VLMs), including PLIP, CONCH, and BiomedCLIP, is highlighted as a pivotal advancement. By aligning visual morphology with textual medical knowledge, these systems enable zero-shot classification, visual question answering, and automated report generation—transforming computational pathology from a "one model per task" paradigm toward general‑purpose, multimodal reasoning systems.
Critical translational perspective — The authors candidly identify key barriers to clinical adoption, including model hallucinations in generative VLMs, adversarial vulnerabilities, fragmented evaluation benchmarks, interpretability and uncertainty calibration challenges, data bias and domain shifts, and regulatory and workflow integration hurdles. They emphasize that clinical readiness requires moving beyond leaderboard accuracy to assess robustness, safety, explainability, factuality, and workflow-level accountability.
A vision for patient-centric "pan-optic" perception — The review articulates a forward-looking framework in which future computational pathology systems integrate histology with radiology, genomics, and clinical records to achieve holistic patient representation. Such systems must not only aggregate multimodal data but also preserve traceable evidence linking local morphology, tissue architecture, WSI-level patterns, and patient-level conclusions.
Clinical and Research Implications
This review offers a timely and authoritative synthesis for researchers, clinicians, and regulatory stakeholders navigating the rapidly evolving landscape of AI-driven pathology. By organizing recent progress around how pathological evidence is represented, transferred, integrated, and potentially lost across scales, the authors provide a structured roadmap for developing context-aware, interpretable, and clinically trustworthy systems. The critical emphasis on validation design, uncertainty handling, human-AI interaction, and post-deployment monitoring makes this review particularly valuable for those seeking to translate computational innovation into real-world patient care.
Full article available on ScienceDirect:
https://doi.org/10.1016/j.intonc.2026.100074
Contact Information for Intelligent Oncology:
LinkedIn: @IntelligentOncology
X: @IntelligentOnco
Facebook: @intelligentoncology
Email Address: editorialoffice@intelligent-oncology.net
Official Website: https://www.sciencedirect.com/journal/intelligent-oncology
Submission Link: https://www2.cloud.editorialmanager.com/intonc/default2.aspx