Google study finds genomic transfer learning can hurt as local datasets grow
Google Research found that European genomic data improved risk prediction when Japanese samples were scarce but could reduce accuracy as local cohorts grew.
Google Research found that transferring genomic information from a European-ancestry cohort improved polygenic-risk prediction when Japanese training samples were scarce, but could reduce accuracy as the target-population dataset grew. In the company’s experiments, pooling more out-of-population data was not uniformly better.
The Google Research evaluation drew European-ancestry samples from UK Biobank and Japanese samples from BioBank Japan, covering eight traits: body mass index, systolic and diastolic blood pressure, red and white blood cell counts, HDL and LDL cholesterol, and blood glucose. Every model was tested on the same held-out BioBank Japan samples. The work is separate from Google’s recently released methane-plume mapping system.
Polygenic risk scores combine information from many genomic variants to estimate a person’s relative likelihood of a condition or trait. The National Human Genome Research Institute says most genomic studies have focused on people of European ancestry, which can make scores less valid for other populations and contribute to health disparities.
Google compared three approaches: one identified variants in UK Biobank before fitting an elastic-net model; another combined variant discoveries across populations before elastic-net training; and PRS-CSx jointly modeled data from multiple populations. Google said models trained only on the Japanese cohort performed better once target-population training sets reached 15,000 samples, while co-training with European data constrained further gains as Japanese datasets expanded.
Where that crossover fell depended on each trait’s genetic architecture. For HDL cholesterol, adding more than 5,000 UK Biobank samples hurt performance once the BioBank Japan training set reached 15,000. Traits whose genetic effects were more similar across populations kept benefiting from European data until the Japanese cohort reached roughly 25,000 to more than 40,000 samples. HDL, LDL and blood glucose crossed over sooner.
Method choice shifted with sample size as well. Google reported that cross-population meta-analysis substantially improved HDL and LDL prediction and improved blood-glucose prediction to a lesser degree. PRS-CSx trailed the strongest elastic-net model below 25,000 target samples for every trait except body mass index, then matched or exceeded the best model near 100,000 target samples for every trait except blood glucose.
The findings are author-reported and were not independently reproduced in the supplied evidence. Google’s post linked no paper, preprint, code repository or statistical supplement, and it did not give uncertainty intervals or exact crossover points for every trait and method.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
