PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY
Orgo-Life the new way to the future Advertising by Adpathway-
Loading metrics
Open Access
Community Page
The Community Page is a forum for organizations and societies to highlight their efforts to enhance the dissemination and value of scientific knowledge.
- Thomas Mock,
- Gust Bilcke,
- Eliott Flaum,
- Shunan Fu,
- Lilian Hoch,
- Kevin Moog,
- Nadine Rijsdijk,
- Elizabeth C. Ruck,
- Ian W. Bishop,
- Carole Duchene
x
- Published: September 1, 2026
- https://doi.org/10.1371/journal.pbio.3003947
Figures
One hundred diatom species have been selected for genome and transcriptome sequencing. The 100 Diatom Genomes Project aims to provide a scalable framework for understanding diatom biodiversity, ecology and evolution, and for investigating their use in biotechnology.
Citation: Mock T, Bilcke G, Flaum E, Fu S, Hoch L, Moog K, et al. (2026) The 100 Diatom Genomes Project. PLoS Biol 24(9): e3003947. https://doi.org/10.1371/journal.pbio.3003947
Published: September 1, 2026
This is an open access article, free of all copyright, and may be freely reproduced, distributed, transmitted, modified, built upon, or otherwise used by anyone for any lawful purpose. The work is made available under the Creative Commons CC0 public domain dedication.
Funding: The work (proposal: https://doi.org/10.46936/10.25585/60001393) conducted by the U.S. Department of Energy Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science User Facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231.TM acknowledges support from the Natural Environment Research Council UK (NE/Z504130/1, NE/W005654/1, NE/R015449/2, and NE/T000848/1). GB, ST, MVB, NR, KV, EPi, EPo, DL, and WV acknowledge funding from Research Foundation Flanders (1228423N, 11L2325N, 1221323N, I000323N, G001521N, and 01G01323), UGent (BOF-GOA 01G01323, BOF/STA/202409/019), and the European Research Council (DIADAPT, 101160805), as well as the Belgian Science Policy (Belspo) for supporting the BCCM/DCG diatoms culture collection, and EMBRC Belgium-FWO project GOH3817N for infrastructure funding. NC, ZC, and SL were supported by the National Natural Science Foundation of China (NSFC, No. 42561144237) and e-ASIA (No. DM1414). TM, SF, and YZ acknowledge support by the National Natural Science Foundation of China (42276134). SF acknowledges support by China Scholarship Council (202106330017). GR, VDD, and FDC acknowledges support by Ministero degli Affari Esteri e della Cooperazione Internazionale Italia (PGR05972). EF is supported by the EIPOD-LinC postdoctoral fellowship programme. GD acknowledges the European Commission (KaryodynEVO, 101078291) and the support of the EMBO Young Investigator Programme. GD and EF acknowledge EMBL for core funding. NP was supported by the Deutsche Forschungsgemeinschaft (PO 2256/1-1). CJ was supported by an MSCA Postdoctoral Fellowship (GA101153556). KEH acknowledges support by the Biotechnology and Biological Sciences Research Council UK (BB/W006286/1), European Research Council (101170086) and Natural Environment Research Council UK (NE/R015449/2). GW acknowledges support from Natural Environment Research Council UK (NE/T000848/1). LT acknowledges support from Région Pays de la Loire Connect Talent EpiAlg and the Epicycle ANR project (ANR-19-CE20- 00). MIF was supported by the Italian Ministry of University and Research (No. 202295S3WC). JML acknowledges support from the National Research Foundation of Korea (RS-2023-00209930) and from the Korea Environment Industry & Technology Institute, funded by Ministry of Climate, Energy and Environment (RS-2021-KE001788). FV, JR, JBK, and TB acknowledge support by the Hellenic Foundation for Research and Innovation (RADIO/483), the European Union (GHaNA/734708) and the European Regional Development Fund co-financed by Greece and the European Union (NSRF 2014-2020/CMBR/MIS 5002670). TB was funded by the German Research Foundation (DFG) through the Transregional Collaborative Research Centre ‘Roseobacter’ (TRR 51). AJA, ECR, and WRR were supported by the U.S. National Science Foundation (Grant DEB-2331644). TAR acknowledges support from the US National Science Foundation (award 2227425). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Diatoms, the most species-rich algal division, are found in all aquatic and some terrestrial environments, and often dominate primary production. They contribute ~20% of annual global carbon fixation and drive the biological carbon pump sequestering organic carbon to the deep sea [1], yet current genomic sampling is heavily biased toward a limited number of model species [2]. Recognizing their fundamental role in sustaining Earth’s habitability, the 100 Diatom Genomes Project (100DGP) was established as an international, interdisciplinary project, bringing together 108 researchers from 11 countries to leverage diatom genomics across a broad range of scientific disciplines. Expanding the genome representation across diverse diatom clades, ecological niches and life histories will facilitate high-resolution comparative genomics to trace evolutionary innovations.
Approaches to studying diatoms in the 100DGP
Species selection was based on practical considerations, including strain availability in culture collections and diatom laboratories from around the world, intrinsic growth rates, and the feasibility of minimizing prokaryotic contaminants. Further checkpoints were the quality and quantity of extractable gDNA and total RNA, both of which can differ substantially between species. So far, representative strains of 104 diatom species have passed our suitability assessments and, of those, 20 annotated assemblies are available (Fig 1a; https://doi.org/10.5281/zenodo.21360381). Although these 104 diatoms cover all major clades, it is not possible to comprehensively sample the entire diversity of diatoms owing to strain availability and cultivability. To mitigate these constraints, we aim to continue the project beyond 100 diatom genomes.
Fig 1. Species selection and gene space diversity of the 100 Diatoms Genome Project.
a. Simplified drawing of the 10 clades of the diatom phylogeny following [2]. Within each clade, the number of species in the pipeline of the 100 Diatom Genomes Project (100DGP) is indicated, as well as the number of species currently described in AlgaeBase (searched on 3 May 2026). The names of species for which both assembly and annotation of genomes have been completed are shown on the right. b. Example of transcriptomic data generated as part of the 100DGP, illustrated for Craspedostauros australis. The heatmap shows the log2 fold changes of 193 diel-responsive genes identified for cells grown under a 14 h light:10 h dark photoperiod (adjusted P < 0.001; log2FC > 2 or log2FC < –2). Cells were harvested for RNA isolation at the midpoint of the light phase and the midpoint of the dark phase on day 4, when cultures were growing exponentially (‘Exp.’), and at the midpoint of the light phase on day 12, when cultures had reached stationary-phase and formed a biofilm (‘Stat.’). Gene expression patterns reveal transcriptional responses associated with the diel cycle as well as changes accompanying the transition from exponential growth to stationary-phase biofilm development. c. Cladogram of selected diatom and outgroup species used to assess gene space evolution, based on the phylogenetic relationships from [2]. In total, 22 diatom genomes covering 17 genera were selected for a preliminary comparative genomic analysis, combining the 9 first annotated genomes from the 100DGP and 13 publicly available assemblies. Drawings (not to scale) of selected species illustrate the morphological and ultrastructural diversity of diatoms. d. Bar plots showing the proportion of gene families with a distinct distribution over diatom genomes: ‘core’, appearing in all diatom genomes; ‘soft-core’, appearing in all but one or two diatom genomes; ‘dispensable’, appearing in 2-19 diatom genomes; and ‘unique’, appearing in a single diatom genome. For each distribution class, pie charts show the proportion that is annotated by at least one gene ontology (GO) term, whereas bubble plots visualize the enrichment significance of the top 3 most enriched GO terms, deduplicated with rrvgo [3]. AdjP: false discovery rate-adjusted p-value. e. Pan-genome rarefaction curve [4] showing the number of gene families identified when the number of sequenced diatom genera increases. Both the total number of gene families detected (pink, pan) and the number of families detected in all genera (blue, core) are visualized. f. Heap’s law (black line) predicts the expected total number of gene families discovered when an increasing number of diatom genera is sampled. The light pink shade depicts the prediction interval, darker pink highlights the confidence interval.
The 100DGP portal provides users with access to all genome assemblies, transcriptomes and annotations. A combination of long-read (PacBio HiFi CCS) and short-read (Illumina NovaSeq X & S4) sequencing has been applied to generate haplotype-resolved genome assemblies and organellar genomes, and a diverse set of transcriptomes (IsoSeq, RNAseq) have been generated for all species for genome annotations. Moreover, for a subset of species, transcriptomes have been generated from cultures exposed to light/dark cycles or raised temperature treatments for future comparative functional studies (Illumina NovaSeq) (Fig 1b; https://doi.org/10.5281/zenodo.21360381). All genomes sequenced at the Joint Genome Institute (JGI) were assembled with hifiasm [5], polished with Racon [6], and annotated using the JGI annotation pipeline [7]. To provide user-friendly interfaces for comparative genomics, multiple bioinformatics analyses were performed and integrated through the PhycoCosm [8] (JGI, Berkeley, USA) and PLAZA platforms (VIB, Ghent, Belgium) (Fig 1c-1f).
To bridge the gap between genetic variants and key underpinning traits, we are complementing our omics analyses with light and expansion microscopy (ExM) and high-throughput phenotyping (phenomics) (Fig 2). As a technique that enables researchers to physically enlarge diatoms, ExM will facilitate the development of molecular markers to resolve the subcellular localization of proteins with nanometer-scale precision, including those encoded by previously uncharacterized or lineage-specific genes. Our high-throughput phenotyping platform should reveal if and how selected proteins have an impact on measurable traits (e.g., carbon acquisition) and their phenotypic plasticity (Fig 2b). The latter will be quantified using the environmentally standardized plasticity index (*ESPI), a recently developed metric that integrates trait responses across controlled environmental gradients to capture the breadth and magnitude of species-specific plastic responses [9]. By applying the *ESPI, we will compare trait plasticity between species and link differences in their plastic response to ecological niches, global distribution patterns, and therefore adaptive potential. Thus, this integrated framework will lay the foundation for mechanistic trait ontology, which is essential to understand both the evolutionary history of diatoms and the adaptability of important traits.
Fig 2. Phenomics, imaging and cellular physiology.
a. High-throughput phenotyping pipeline to quantify key traits (maximum PSII quantum yield (Fv/Fm), growth rate, cell size, chlorophyl-a content, and non-photochemical quenching (NPQ)), under different environmental conditions (3x temperatures, 3x light intensities, 16x C:N:P ratios). These data will be used to quantify phenotypic plasticity using an Environmentally Standardised Plasticity Index (*ESPI) representing the breadth of the phenotypic responses of each species to changing environmental conditions. This metric is designed to facilitate cross-species comparisons of phenotype plasticity, allowing for greater understanding of phenome response to climate variation. Greatest variance in response per trait is normalized using the Euclidean distance between the environments in which these responses are found, to quantify plasticity per unit of environmental change. In the radial *ESPI plot, each red dot represents the *ESPI value for a given trait for each of the 100 diatom species; larger overall *ESPI values indicate greater phenotypic plasticity. The radial plot shows low *ESPI in the center and high towards the exterior. Information from (a) and (b) will be merged to develop a trait ontology for diatoms. b. Cell morphology of Coscinodiscus granii (left) and Craspedostauros australis (right) using light and expansion microscopy (ExM) to demonstrate the nanoscale distribution of molecular complexes: tubulin (magenta), photosystem II (green), nucleus and silica cell wall (blue), NHS ester (white). Black scale bars are 5 μm. White scale bars are 20 μm (4.65 μm biological scale) and 5 μm (1.37 μm biological scale), respectively.
Resources and preliminary insights for advancing integrative diatom research
In line with Findable, Accessible, Interoperable and Reusable (FAIR) principles of scientific data management, the 100DGP will publish laboratory methods including standard operating procedures, code, and a diverse set of resources (e.g., genomes, transcriptomes, images, and trait data), together forming an integrated dataset that can be leveraged for future functional studies of diatoms. We encourage researchers to make use of these resources, including the deposited diatom strains at the BCCM/DCG diatoms collection (Ghent, Belgium). Some strains are also available from NCMA (Bigelow, East Boothbay, Maine, USA), RCC (Roscoff, France) and CCAP (Oban, UK). Data integration from genomes, transcriptomes, nanoscale imaging and phenomics is unprecedented for any protist group and will allow us to place comparative evolutionary genomics into the context of cell biology, metabolism and ecology for questions that go beyond the role of individual genes. The first genome assemblies are already publicly available (100DGP portal), revealing a range of genome sizes (75–1,033 Mbp) and >3-fold variation in gene counts. Future comparative analyses will be released through PhycoCosm [8], DiatOmicBase [10], and PLAZA Diatoms: versatile gene-centered platforms for integrative diatom omics research.
External researchers can use this resource to link the outcome of comparative evolutionary genomics (e.g., genomic loci under selection) with nanometer-scale detection of biomolecules and how they shape key traits such as carbon acquisition, nutrient utilization, morphotype diversification, and life-cycle regulation. Beyond fundamental research, our integrative datasets can be exploited for linking genetic variants with the in-vivo function of their encoded proteins and other bioactive compounds for advancing technological applications, such as drug discovery or bioprocess optimization [11].
Preliminary insights based on a comparative analysis of 22 new and existing diatom genomes revealed that only <3% of gene families and 32% of genes are conserved amongst all diatoms (core, soft-core) (Fig 1d). Gene families were identified by performing an all-versus-all homology search between diatom and outgroup proteomes using DIAMOND with subsequent clustering into gene families using the tribeMCL algorithm. Since the rise of diatoms more than 250 mya, a high level of genetic diversity has evolved through a combination of divergent evolution in existing gene families and the appearance of novel genes, either through de-novo gene birth or horizontal gene transfer [12]. Although most gene ontology terms enriched in the diatom shared pan-genome represent genes involved in primary metabolism, the enrichment of cyclic nucleotide-mediated signaling in the diatom core genome supports previous evidence that these secondary messengers drive a multitude of essential abiotic and biotic processes in diatoms. Dispensable and unique genes, on the other hand, were enriched in receptor signaling, flagella organization during reproduction, cell division, DNA recombination, and nuclear organization, which suggest that these processes underpin speciation. Although still preliminary, these results suggest that adaptation may have led to the emergence of novel genes and the exceptional diversity of diatoms.
Based on the current number of genomes, it appears the diatom pan-genome is open (Fig 1e), and Heap’s law [13] predicts that the number of gene families, including their novel genes and variants, will double to ~40,000 by the project end (Fig 1f). To address their potential role in diatom cell biology, particularly if identified as part of known metabolic pathways such as photosynthesis, we will develop molecular markers to reveal the subcellular localization of their encoded proteins using ExM [14]. Our high-throughput phenotyping platform will show if and how those genetic variants and their encoded proteins affect measurable traits, such as light harvesting, carbon acquisition, elemental stoichiometry, or cell division. Early findings indicate that diatoms exhibit a wide range of phenotypic responses, likely driven by their genetic diversity and cellular structural complexity. Altogether, we are constructing a common ontology of traits, critical for developing an integrated approach based on evolutionary genomics, suitable for building predictive models from cell biology to ecosystem science.
Conclusions
The 100DGP aims to generate diverse biological resources for a key group of primary producers. Our project also aims to provide a test case for protist research at scale, which, for many other protist groups, is still in its infancy. To grow beyond 100 diatom genomes, we are integrating externally sequenced diatom genomes and transcriptomes from other projects (e.g., from the Chaetoceros genus sequencing project at JGI). Our preliminary results, based on all resources, provide evidence that integrated research at scale is feasible for diatoms and promise to be a potential treasure trove for discovering novel biology, including unique adaptations and interactions (e.g., endosymbionts). We hope that this work will improve our understanding of global species diversity, evolution, and ecosystem functioning, as well as providing novel genetic variants for advancing biotechnology via genetic engineering and synthetic biology.
Acknowledgments
TM acknowledges the School of Environmental Sciences, University of East Anglia, Norwich (UK), and the College of Environmental Science and Engineering, Ocean University of China for providing time to coordinate this large-scale project. NP acknowledges the CMCB light microscopy facility, a Core Facility of the CMCB Technology Platform at TU Dresden (Germany) for light microscopy support.
References
- 1. Field C, Behrenfeld M, Randerson J, Falkowski P. Primary production of the biosphere: integrating terrestrial and oceanic components. Science. 1998;281(5374):237–40. pmid:9657713
- 2. Alverson AJ, Roberts WR, Ruck EC, Nakov T, Ashworth MP, Bryłka K, et al. Phylogenomics reveals the slow-burning fuse of diatom evolution. Proc Natl Acad Sci U S A. 2025;122(22):e2500153122. pmid:40440071
- 3. Sayols S. rrvgo: a bioconductor package for interpreting lists of gene ontology terms. MicroPubl Biol. 2023;2023:10.17912/micropub.biology.000811. pmid:37151216
- 4. Zhao Y, Jia X, Yang J, Ling Y, Zhang Z, Yu J, et al. PanGP: a tool for quickly analyzing bacterial pan-genome profile. Bioinformatics. 2014;30(9):1297–9. pmid:24420766
- 5. Cheng H, Concepcion GT, Feng X, Zhang H, Li H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18(2):170–5. pmid:33526886
- 6. Vaser R, Sović I, Nagarajan N, Šikić M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 2017;27(5):737–46. pmid:28100585
- 7. Grigoriev IV, Nikitin R, Haridas S, Kuo A, Ohm R, Otillar R, et al. MycoCosm portal: gearing up for 1000 fungal genomes. Nucleic Acids Res. 2014;42(Database issue):D699-704. pmid:24297253
- 8. Grigoriev IV, Hayes RD, Calhoun S, Kamel B, Wang A, Ahrendt S, et al. PhycoCosm, a comparative algal genomics resource. Nucleic Acids Res. 2021;49(D1):D1004–11. pmid:33104790
- 9. Hoch L, Herdean A, Woodcock S, Songsomboon K, Osborne B, Ralph PJ. New method for quantification of phenotypic plasticity reveals how plasticity changes over time in the Diatom Thalassiosira weissflogii. Ecol Evol. 2026;16(2):e73072. pmid:41716591
- 10. Villar E, Zweig N, Vincens P, Cruz de Carvalho H, Duchene C, Liu S, et al. DiatOmicBase: a versatile gene-centered platform for mining functional omics data in diatom research. Plant J. 2025;121(6):e70061. pmid:40089834
- 11. Dong Z, Xiang W, Jiang W, Guo T. Expansion omics: from expansion microscopy to spatial omics. Mol Syst Biol. 2026;22(2):165–78. pmid:41326779
- 12. Vancaester E, Depuydt T, Osuna-Cruz CM, Vandepoele K. Comprehensive and functional analysis of horizontal gene transfer events in Diatoms. Mol Biol Evol. 2020;37(11):3243–57. pmid:32918458
- 13. Tettelin H, Riley D, Cattuto C, Medini D. Comparative genomics: the bacterial pan-genome. Curr Opin Microbiol. 2008;11(5):472–7. pmid:19086349
- 14. Flori S, Mikus F, Flaum E, Moog K, Guessoum S, Beavis T, et al. Diatom ultrastructural diversity across controlled and natural environments. Curr Biol. 2025;35(23):5709-5720.e4. pmid:41175869


















English (US) ·
French (CA) ·