Introduction
Artificial intelligence (AI)—encompassing machine learning (ML), deep learning (DL), natural language processing (NLP), and foundation models—has emerged as a transformative technology for oncology biomarker discovery, enabling systematic extraction of actionable signals from increasingly complex, multimodal datasets 17. Traditional biomarker development relied on hypothesis-driven, single-modality assays and predominantly retrospective analyses. AI now enables integration of genomic, imaging, pathological, and clinical data at scale, accelerating discovery across the cancer care continuum from early detection to minimal residual disease (MRD) monitoring 115. As of mid-2026, the field has matured beyond exploratory research into early clinical implementation, supported by regulatory frameworks from the U.S. Food and Drug Administration (FDA) and clinical guidance from the European Society for Medical Oncology (ESMO), while critical validation gaps persist 2829.
A foundational distinction for clinical audiences is the developmental stage of any AI biomarker: exploratory discovery (hypothesis-generating, unvalidated), analytical validation (technical performance demonstrated), clinical validation (association with outcomes confirmed prospectively), and clinical implementation (integrated into care with demonstrated utility). Most published AI biomarker studies remain in exploratory or analytical validation phases; rigorous prospective evidence supporting clinical utility is limited 1529.
AI Methods and Approaches
Machine learning methods—including random forests (RF), support vector machines (SVM), and gradient boosting—identify patterns in structured tabular data and remain workhorses for risk stratification and treatment selection, offering computational efficiency and interpretability 3. Deep learning models, particularly convolutional neural networks (CNNs) and transformers, learn hierarchical representations directly from high-dimensional inputs such as medical images, whole-slide images (WSIs), and genomic sequences without manual feature engineering, enabling discovery that conventional statistics cannot achieve 213. Radiomics systematically extracts quantitative imaging features (shape, texture, intensity) from computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET), which can be linked to genomic or clinical phenotypes non-invasively 5.
Digital pathology applies DL to WSIs to quantify histomorphological biomarkers and predict molecular subtypes from routine tissue sections—an approach exemplified by EAGLE, a fine-tuned pathology foundation model that predicts epidermal growth factor receptor (EGFR), KRAS, and anaplastic lymphoma kinase (ALK) mutation status from pathology images, reducing dependence on molecular assays 21. NLP mines unstructured electronic health record (EHR) clinical notes to extract biomarker-relevant data (e.g., programmed death-ligand 1 [PD-L1] expression status, performance status, adverse events) with high precision and recall, enabling population-level biomarker studies without manual abstraction 4. Foundation models—large, pre-trained architectures adaptable across tasks—represent the current frontier, with multimodal versions integrating pathology, genomics, and radiology signals simultaneously emerging at major conferences including the American Association for Cancer Research (AACR) Annual Meeting 2026 32. Spatial AI, applying graph neural networks (GNNs) to spatial transcriptomics and proteomics data, decodes tumor microenvironment (TME) architecture with unprecedented resolution, identifying topological biomarkers of immune resistance that surpass conventional PD-L1 quantification 27.
Data Modalities for Biomarker Discovery
The richness of AI-driven biomarker discovery depends critically on the input data modalities available. Each carries distinct strengths and limitations summarized in Table 2 below. Multi-omics integration—combining genomics, transcriptomics, proteomics, epigenomics, and metabolomics—offers the most comprehensive tumor characterization but introduces substantial data harmonization complexity 13. The Molecular Twin platform in pancreatic cancer demonstrated that a parsimonious model using 589 of 6,363 multi-omic features maintained survival prediction accuracy comparable to the full model, illustrating the principle that feature reduction preserves clinical utility while reducing analytical cost 9. Single-cell RNA sequencing (scRNA-seq) datasets integrating 223 immunotherapy-treated patients across nine cancer types provide cell-state resolution of immune resistance mechanisms, with direct implications for biomarker discovery 4. Wearable devices and remote monitoring generate longitudinal digital biomarkers (activity, heart rate, sleep) that are feasible in older oncology patients (adherence 74–100% across reviewed studies) but lack standardized outcome definitions and prospective validation 4.
Clinical Applications Across the Cancer Care Continuum
A systematic review and meta-analysis of 34 studies on AI for lung cancer biomarker prediction reported pooled sensitivity of 0.77 (95% confidence interval [CI]: 0.72–0.82) and specificity of 0.79 (95% CI: 0.78–0.84) 2. The most frequently targeted biomarkers were EGFR mutations, PD-L1 expression, and ALK rearrangements, predominantly using CNN-based models applied to CT imaging, WSIs, and genomic data 25.
Early detection has been meaningfully advanced by the PRAIM study—a nationwide real-world implementation of CE-certified AI for population-based mammography screening across 12 German sites (461,818 women). AI support increased breast cancer detection rate by 17.6% relative (6.70 vs. 5.70 per 1,000; 95% CI: +5.7%, +30.8%), improved positive predictive value of recall from 14.9% to 17.9%, and raised biopsy positive predictive value from 59.2% to 64.5%, with noninferior recall rates—a concrete model of AI-human collaborative clinical implementation 20.
Immunotherapy response prediction has incorporated spatial AI, with a 2025 hepatocellular carcinoma (HCC) study achieving area under the receiver operating characteristic curve (AUROC) >0.90 using multimodal spatial transcriptomic–proteomic inputs and GNNs, substantially outperforming unimodal approaches 27. In pan-cancer immunotherapy cohorts, the Clinical Transformer—an interpretable multimodal AI framework—achieved a hazard ratio of 0.29 comparing high- versus low-risk groups across 12 datasets and >150,000 patients, with a permutation-based feature importance module enabling clinician-interpretable outputs 23.
MRD monitoring via circulating tumor DNA (ctDNA) tumor fraction (TF), quantified through comprehensive genomic profiling, demonstrated independent prognostic value in two large precision medicine cohorts (BIP: n=965; STING: n=947), with a multivariate nomogram achieving 3-month AUROC of 0.80 in development and 0.76 in validation cohorts 24. The FDA has approved multiple ctDNA-based companion diagnostics for targeted therapies but has not yet approved ctDNA as an immunotherapy response biomarker, emphasizing the evidence gap 25.
Multi-omics AI models for prognosis, exemplified by an 18-gene signature integrating programmed cell death and metabolic pathways in HCC (identified via consensus clustering and ML), demonstrated strong overall survival prediction and immunotherapy sensitivity correlation, while also enabling drug repurposing discovery 6. The China Protocol for lung cancer—integrating AI-based CT analysis (AUROC improved to 0.824 with self-supervised multitask learning) and methylation-based liquid biopsy (LunaCAM-S AUROC 0.90 for screening)—observed increases in early-stage detection from 46.3% (2019) to 65.6% (2023) and 5-year survival to 59% overall across 11,844 patients, though the authors explicitly flag the need for multi-ethnic, multicenter prospective validation 22.
Selected clinical applications with validation considerations are detailed in Table 3.
Methodological and Clinical Validation Challenges
The most critical gap in the field is the predominance of retrospective, post hoc analyses. Among 90 studies on AI for immuno-oncology biomarkers, no study incorporated AI-based methodologies prospectively from inception 15. The ESMO Basic Requirements for AI-based Biomarkers in Oncology (EBAI) framework—the first comprehensive clinical guidance in this area—addresses this by classifying biomarkers into Class A (quantification of established biomarkers), Class B (prediction of novel phenotypes requiring prospective validation), and Class C (exploratory discovery) 2930. This tiered framework is essential for clinicians evaluating commercial platforms.
Data quality and bias remain foundational challenges. Small sample sizes, batch effects, overfitting, and data leakage are widespread; a 2025 perspective on ML for clinical proteomics explicitly cautioned against uncritical application of complex DL architectures in typical clinical proteomics datasets with limited interpretability gains 8. The Image Biomarker Standardization Initiative (IBSI) and FAIR (Findability, Accessibility, Interoperability, Reusability) data principles provide partial frameworks for radiomics and omics standardization, respectively 1. Marked staining variability across institutions—e.g., MSH6 interpretation in prostate cancers—illustrates how pre-analytical variation can bias AI models 32.
Interpretability and explainability are regulatory and clinical imperatives. The FDA's draft guidance on AI-enabled device software mandates algorithm transparency and post-market performance monitoring 28. Explainable AI (XAI) methods—including SHapley Additive exPlanations (SHAP), integrated gradients, layer-wise relevance propagation, and saliency maps—enable clinicians to understand which features drive predictions 127. Commercially deployed AI digital pathology platforms (ArteraAI for prostate, Vesta-Bladder, VisioCyt) vary substantially in their explainability support and strength of supporting evidence 31.
Federated learning—training models across decentralized institutional datasets without centralizing patient data—offers a promising solution for privacy-preserving multi-site validation but remains limited in oncology biomarker settings, with regulatory frameworks still nascent 428. Regulatory approvals as of 2026 include 204 FDA-approved companion diagnostic (CDx) devices, including notable recent approvals such as MI Cancer Seek (Caris Life Sciences, approved November 2024) for melanoma and non-small cell lung cancer (NSCLC) tissue analysis 19. A prospective randomized trial (n=20,707) testing AI-triggered notifications for genomically matched clinical trials found no increase in trial enrollment, indicating that effective AI-driven trial matching requires broader recruitment strategy integration 4.
Future Opportunities
The AACR Annual Meeting 2026 highlighted several strategic directions: multimodal foundation models integrating spatial transcriptomics, proteomics, radiology, and clinical data; agentic AI systems capable of autonomous hypothesis generation and experimental prioritization; and virtual cell models predicting cellular responses to perturbations without extensive experimental validation 32. Spatial AI combining sequencing-based platforms (10x Genomics Visium HD) with imaging-based systems (Xenium, CosMx) and AI-driven GNN analysis is yielding spatial phenotypes—immune exclusion zones, dysfunctional inflamed niches, tertiary lymphoid structure maturation states—that outperform conventional biomarkers in predicting immunotherapy response 27. Federated learning, synthetic data generation, and liquid biopsy ctDNA methylation integration represent additional high-priority investments 2526. Industry partnerships, such as the Lunit-Labcorp collaboration launched November 2025, are accelerating commercial translation of spatial profiling AI biomarkers in NSCLC 31.
Table 1: AI Methods and Typical Oncology Biomarker Applications
| AI Method | Data Input | Typical Biomarker Application | Key Strengths | Key Limitations |
|---|---|---|---|---|
| ML (Random Forest, SVM, XGBoost) | Structured tabular: genomics, clinical variables | Risk stratification, treatment selection, toxicity prediction | Interpretable; computationally efficient; validated | Requires feature engineering; limited for unstructured data |
| Deep Learning (CNN, RNN, Transformer) | Images, sequences, raw omics, text | Image-based diagnosis, genomic signature discovery, EHR mining | Learns representations automatically; handles high-dimensional data | "Black-box"; requires large datasets; computationally intensive |
| Radiomics + ML | CT, MRI, PET imaging features | Tumor prognosis, genotype prediction, treatment response | Non-invasive; integrable with clinical imaging | Platform-dependent; standardization challenges |
| Digital Pathology (CNN, Foundation Models) | Whole-slide histopathology images | Molecular subtype prediction (EGFR, KRAS, ALK), spatial immune profiling | No additional tissue; leverages routine slides | Slide quality variability; limited explainability in commercial tools |
| NLP | EHR clinical notes | PD-L1 status extraction, adverse event detection, trial eligibility | Automates manual abstraction; enables population-level studies | Domain-specific vocabulary; privacy concerns |
| Foundation Models | Multimodal (text, images, omics) | Multi-omics integration, biomarker panel design, clinical trial matching | Pre-trained representations; few-shot learning | Limited prospective validation; regulatory uncertainty |
| Spatial AI (GNN) | Spatial transcriptomics, spatial proteomics | TME architecture, immune exclusion phenotyping, immunotherapy response | Captures topological biology; surpasses bulk biomarkers | Expensive platforms; computational burden; limited standardization |
| Federated Learning | Decentralized multi-site data | Privacy-preserving multi-site validation | Preserves patient privacy; enables large-scale validation | Infrastructure complexity; regulatory frameworks nascent |
| Bayesian Networks | Multi-omics, clinical data | Joint outcome and toxicity prediction, causal inference | Interpretable; models causal relationships | Computationally complex; requires domain expertise |
Table 2: Data Modalities, Strengths, and Limitations
| Data Modality | Biomarker Type | Strengths | Limitations | Typical AI Application |
|---|---|---|---|---|
| Genomics (WES, targeted panels, ctDNA) | Mutations, CNAs, fusions, TMB | Well-established; FDA-approved CDx exist; actionable for targeted therapy | Static snapshot; tumor heterogeneity; clonal evolution | Treatment selection; MRD monitoring |
| Transcriptomics (RNA-seq, scRNA-seq) | Gene-expression signatures, cell-state programs | Functional relevance; cell-type resolution | Batch effects; tissue accessibility; cost | Immune profiling; subtype classification |
| Proteomics (mass spectrometry) | Protein abundance, phosphorylation, PTMs | Closer to functional phenotype; therapeutic target identification | Cost; limited coverage; dynamic range challenges | Prognosis; drug-target prediction |
| Epigenomics / ctDNA methylation | DNA methylation patterns; cancer-specific markers | Early detection potential; non-invasive (liquid biopsy) | Assay standardization lacking; population variability | Early detection; MRD monitoring |
| Metabolomics | Circulating metabolites, pathway activity | Reflects metabolic phenotype; drug response monitoring | Standardization limited; platform cost | Prognosis; treatment response |
| Imaging (CT, MRI, PET, Mammography) | Radiomics features, morphology | Non-invasive; repeatable; large-scale screening | Limited molecular information; operator-dependent | Detection; prognosis; response assessment |
| Digital Pathology (WSI) | Histomorphology, spatial immune features | High resolution; quantitative; integrable with genomics | Annotation burden; stain variability; computational cost | Diagnosis; prognosis; genotype prediction |
| EHR / Clinical Notes | Performance status, adverse events, treatment response | Longitudinal; real-world; comprehensive phenotypes | Unstructured; variable documentation; privacy concerns | NLP extraction; trial enrichment |
| Wearable / Longitudinal | Activity, heart rate, sleep, vital signs | Continuous; real-time; patient-centric | Adherence variability; technical artifacts; no standardized definitions | Toxicity detection; recovery monitoring |
| Spatial Multi-omics | Cell-cell relationships, TME topology | Captures spatial architecture; surpasses bulk-level correlations | Expensive; limited clinical availability; computational burden | Immunotherapy response; TME profiling |
| Multi-Omics Integration | Integrated biomarker panels | Holistic tumor biology; improved predictive accuracy | Data harmonization complexity; sample burden; computational cost | Comprehensive biomarker signatures; drug repurposing |
Table 3: Selected Clinical Applications with Validation Considerations
| Clinical Application | Cancer Type | AI Method | Biomarker | Key Performance | Validation Status | Regulatory Status | Actionability |
|---|---|---|---|---|---|---|---|
| Mammography screening | Breast | CE-marked CNN (Vara MG) | Imaging AI signal | BCDR +17.6% relative; PPV recall +3.0 pp (14.9% to 17.9%) | Prospective real-world (n=461,818; PRAIM) | CE-marked | Clinically implemented (select EU sites) |
| EGFR mutation detection | NSCLC | CNN, ML | EGFR mutations | Sensitivity 61–90%; Specificity 41–97% | Retrospective + external validation; some real-world deployment (EAGLE) | FDA-approved CDx (cobas EGFR v2; plasma) | Companion diagnostic for TKI selection |
| KRAS G12C detection | NSCLC | NGS-based AI | KRAS G12C | Analytically validated | FDA-approved CDx (Agilent Resolution ctDx FIRST, 12/2022) | FDA-approved CDx | Guides sotorasib/adagrasib selection |
| PD-L1 expression assessment | NSCLC, Melanoma | CNN | PD-L1 expression | Sensitivity 68–95%; Specificity 96–97% | Retrospective; internal validation predominant | FDA-approved (immunohistochemistry assays; not AI-specific) | Exploratory AI prediction; not yet standard |
| Immunotherapy response prediction | Pan-cancer | Clinical Transformer (multimodal) | Multi-feature signature | HR 0.29 (high vs. low risk; 12 datasets, n>150,000) | Clinical validation; retrospective multicohort | Exploratory; CDx pathway emerging | Research/pilot use |
| MRD monitoring (ctDNA TF) | Pan-cancer | Genomic profiling + ML | ctDNA tumor fraction | 3-month AUROC 0.80 (dev.); 0.76 (val.) | Analytical + clinical validation (n=1,912 across 2 cohorts) | Companion diagnostic pathway (FDA guidance) | Trial enrichment; not yet standard of care |
| Early lung cancer detection | Lung | Self-supervised AI + methylation | LunaCAM-S (methylation) | AUROC 0.90 (screening); 5-yr OS 59% overall | Real-world prospective (n=11,844; China Protocol) | Multi-ethnic multicenter validation needed | Promising; requires prospective validation |
| HCC prognosis | HCC | Consensus ML clustering + multi-omics | 18-gene PCD-metabolic signature | Strong OS prediction; immunotherapy sensitivity correlation | Retrospective; limited external validation | Exploratory | Research tool |
| Spatial immunotherapy response | HCC | GNN (spatial transcriptomics + proteomics) | TME topology (immune exclusion) | AUROC >0.90 | Analytically validated; prospective validation pending | Exploratory | Research tool; emerging clinical interest |
| Digital pathology genotype prediction | NSCLC | Pathology foundation model (EAGLE) | EGFR, KRAS, ALK from WSI | Early real-world deployment | Early real-world | Emerging; likely CDx pathway | Accelerates molecular turnaround time |
Conclusion
AI-driven biomarker discovery has advanced from exploratory research to early clinical implementation in oncology, supported by FDA-cleared companion diagnostics, CE-marked imaging AI systems, ESMO's EBAI classification framework, and prospective real-world studies such as PRAIM. Pooled analytical performance for lung cancer biomarker AI is high (sensitivity ~0.77, specificity ~0.79), and spatial AI, multimodal foundation models, and ctDNA-based MRD tools represent the next generation of clinically actionable technologies 22724. However, the fundamental gap between retrospective analytical validation and prospective clinical utility remains the most pressing challenge. Medical professionals evaluating AI biomarkers should demand explicit statement of developmental stage, external validation evidence, regulatory status, and interpretability support before incorporating them into clinical decision-making. Multidisciplinary collaboration among oncologists, pathologists, data scientists, biostatisticians, and regulatory experts—alongside prospective clinical trial designs with a priori planned AI methodologies—is essential to ensure that AI-enabled precision oncology benefits patients equitably and sustainably 151729.