What PRS does well
Stratifies inherited risk at population scale, often improving discrimination when combined with age, family history, biomarkers, and clinical factors. [7]
A practical library on the scientific foundation of polygenic scores: from publications, model specifications, and reference data to validation, calibration, technical quality, and responsible implementation.
This page shows how Starling reads evidence for polygenic-score applications. Not every publication, model, or analytical capability is automatically ready for routine use; value depends on intended use, data quality, external validation, calibration, and local context.
The central theme is caution with purpose. Polygenic scores are probabilistic instruments, and conclusions must never exceed what the data, model performance, and current evidence support.
Responsible PRS implementation connects discovery evidence, model specification, genotype and imputation quality, external validation, calibration, ancestry context, and comparison with established clinical predictors.
Starling therefore makes clear which sources support each score, what validation is needed, which conclusion can be reported, and where population, pathway, or territory-specific limitations apply.
Across most clinical areas, PRS is best supported as a risk modifier that complements standard clinical assessment. Evidence is strongest where PRS is integrated with existing pathways (screening, monitoring, classification), and weakest where PRS is proposed as a stand-alone decision tool. [2]
Stratifies inherited risk at population scale, often improving discrimination when combined with age, family history, biomarkers, and clinical factors. [7]
It does not diagnose disease, and a low score does not rule out disease. Clinical context remains essential. [2]
Ancestry portability and calibration: performance can drop when a score is transferred across populations without local validation. [13]
Meaningful evidence is rarely a single headline result. It usually combines discovery studies, population biobanks, ancestry-reference resources, and external validation cohorts that show whether the score remains interpretable beyond its original training environment.
For real implementation questions, the strongest studies are the ones that compare polygenic models with existing predictors, test calibration carefully, and examine how reporting could work inside routine pathways rather than in isolation.
We use a simple three-part framework for each domain: clinical validity, incremental value, and implementation readiness. [3]
Can the PRS distinguish relatively higher-risk versus lower-risk individuals in relevant populations?
Does PRS materially improve prediction beyond established predictors already used in care?
Is the score reproducible, calibrated, ancestry-appropriate, and interpretable for safe reporting?
Different PRS for the same endpoint can perform similarly at population level but classify some individuals differently, which matters for hard percentile cutoffs. [5]
The currently available indications sit within a broader PGS literature that is rarely built from a single dataset. Most usable models combine discovery GWAS, ancestry reference resources, and external validation cohorts for development, calibration, and transfer testing.
Large resources such as UK Biobank and All of Us are commonly used to train, benchmark, or validate broad disease and quantitative trait scores at scale.
Datasets such as 1000 Genomes are commonly used to support ancestry assignment, allele-frequency checks, and more careful interpretation across diverse populations.
Prospective or deeply phenotyped cohorts such as MESA and ARIC are especially valuable because they test whether performance holds up outside the original discovery environment.
Some indications also rely on specialist consortia or registries, for example CARDIoGRAMplusC4D, Biobank Japan, CIMBA, and WHI, to strengthen domain-specific risk modeling.
The wider literature across oncology, ophthalmology, immune disease, bone health, neurology, psychiatry, renal medicine, and quantitative traits remains highly relevant. It shows where polygenic methods are gaining traction and where translational potential is beginning to appear.
At the same time, those domains differ substantially in readiness. Some have strong population-level signal but weak pathway integration; others have promising screening or surveillance relevance but still depend on better calibration, ancestry transfer, or prospective outcome data.
Quick comparison of where evidence is strongest today and what still limits routine use. Ratings are synthesis-level judgments based on the cited studies and consensus statements on this page.
| Domain | Evidence Strength | Near-Term Clinical Readiness | Main Caveat |
|---|---|---|---|
| Oncology | High | High | Outcome impact and equity across ancestry groups still need stronger prospective data. [9] [10] |
| Cardiometabolic & Vascular | High | Moderate | Incremental value over strong clinical models varies by endpoint and care setting. [6] [7] |
| Renal | Moderate | Moderate | Integration with APOL1/clinical data helps, but prospective decision-impact data is limited. [8] |
| Bone Health | Moderate | Moderate | Thresholds and pathway standardization are still developing across populations. [14] |
| Gastroenterology & Immune | Moderate | Moderate | Actionable clinical pathways are uneven despite strong genetic signal. [15] |
| Ophthalmology | Moderate | Moderate | Calibration and validated treatment or surveillance thresholds remain key constraints. [11] |
| Neurology & Psychiatry | Moderate | Low-Moderate | Effect sizes, portability, and ethical complexity limit routine susceptibility use. [16] |
| Quantitative Traits | High | Moderate | Strong statistical signal does not automatically support direct treatment decisions. [12] |
Trait PRS can be statistically strong because traits are directly measured at scale. Clinical value depends on whether trait information connects to validated intervention pathways. [12]
Key sources underlying this page, including standards, major evaluations, and domain-focused studies.
We discuss evidence, model selection, cohort context, and interpretation questions with clinics, healthcare providers, laboratories, and research teams implementing polygenic scores.