Data philosophy #
- Three evidence tiers are kept distinct across the site: (1) curated literature — manually collected and verified records (see
longevity_literature_verified.json); (2) measured atlas data — processed single-cell / bulk / metabolomics datasets; (3) model output — predictions such as aging-clock scores, always labelled as such.
- Every site update is archived as a dated code snapshot on the Versions page, so published results can be traced back to the exact software that produced them.
- This page summarizes methods at the workflow level and intentionally omits numeric parameter values; exact processing parameters travel with each dataset (
config.json) and with the site code snapshot of the release.
Atlas construction #
- Preprocessing: cells with extreme gene/count ranges or high mitochondrial content are filtered; counts are normalized to a common depth and log-transformed; highly variable genes are selected per dataset.
- Integration & embedding: datasets are integrated with a batch-aware graph integration workflow (BBKNN/Harmony-style); a 2-D UMAP embedding is computed on the integrated PCA space and stored as
umap.bin.gz (float32 x, float32 y, uint16 cell-type index, 10 bytes/cell).
- Cell-type annotation: clusters are annotated with canonical marker panels and reference-based mapping, organized in a two-level hierarchy (category → cell type) recorded in
celltype_legend.json with names, colors and per-type cell counts.
- Gene vectors: per-gene expression vectors are exported into
genes/<GENE>.bin.gz for interactive visualization; dataset-level statistics live in config.json.
- Atlas packages for all 137 species can be downloaded from the Download Center.
- Full preprocessing protocol: the complete pipeline — two ingestion tracks (raw FASTQ vs. author matrices), QC thresholds, normalisation, the level-dependent batch-correction scheme (within-dataset / within-species / cross-species, with batch keys) and BioID cell-type annotation with validation — is documented step by step on the Data Processing page, with a per-dataset breakdown in the Dataset Processing Table.
Cross-species comparison #
- Orthology mapping: genes are mapped to human orthologs using OrthoFinder preferentially, supplemented by BioMart and DIAMOND reciprocal best hits; for one-to-many groups a single representative ortholog (best reciprocal hit) is retained for cross-species analyses, and genes without a human ortholog are excluded from cross-species comparisons but remain available in species-level data.
- Expression summarization: single-cell expression is aggregated per cell type (median and detection rate) to make species comparable despite different atlas depths.
- Comparison metrics: module pages report fold-change rank agreement and correlation of cell-type expression profiles; the comparison views in Gene Comparison and Cell Comparison display these side by side.
- Version note: the comparison modules have evolved through several implementations (the
content.v*.php series). The description here matches the current live pages; archived copies reached by old links may behave differently.
Cross-species expression similarity is descriptive, not phylogenetic — it does not by itself establish conserved function.
Module species compatibility #
- Why compatibility differs: several integrated tools (CellPhoneDB, NicheNet, SCENIC, DisGeNET, OMIM disease associations, drug–target annotations) were built primarily from human or mammalian annotations, so their applicability varies by species.
- Three tiers: (i) data-driven modules (co-expression, trajectory, DEG) are statistically supported for all species; (ii) annotation-dependent modules (cell communication, GRN, GO/KEGG & GSEA, motif, metabolism) run on one-to-one human orthologs for non-human species, with reliability decreasing with phylogenetic distance; (iii) human-annotation-bound modules (ToppGene enrichment, disease association, drug target) operate through human orthologs only for non-human species and are hypothesis-generating there.
- Per-species coverage: the ortholog mapping behind every cross-species analysis is quantified per species and is downloadable from the Download Center, flagged by mapping type (one-to-one, one-to-many, missing).
- The full module-by-module matrix cited in the LACA manuscript is published on the Module Species Compatibility page.
Outputs of annotation-dependent modules in non-human species are hypotheses, not statistically supported conclusions; disease phenotypes are not equivalent across species.
Correlation analysis #
- Gene–trait correlation (Gene Correlation, Longevity Correlation): per-gene or aggregated signatures are correlated with longevity-related traits using Spearman (default) and Pearson coefficients.
- Cell–trait correlation (Cell Correlation): cell-type abundance or per-cell expression is correlated with the trait; cell-level results are also exported as tables (e.g.
cell_level_celltype_correlation.csv).
- Significance: p-values come from asymptotic tests and are adjusted for multiple testing (Benjamini–Hochberg FDR); color scales on module pages encode both effect size and significance.
Correlations are associations, not causal claims; effect sizes shrink as the number of tested genes grows, so interpret FDR-adjusted values.
Aging clock #
- Feature selection: age-associated genes are selected by correlation with chronological age across donors; the SP121 module restricts to genes measurable in that atlas.
- Model: a regularized linear model (elastic-net-style) predicts age from expression; scores reported on Aging Clock are model outputs and are labelled as predictions.
- Validation: models are evaluated on held-out donors; reported metrics use the holdout split.
- Two implementations: the main clock (this page) and the SP121 module (aging-clock-sp121-module.php) are fitted independently on their own inputs; treat their scores as separate predictors.
Aging hallmarks #
- Gene sets: hallmark modules use the established aging-hallmark gene sets (genomic instability, telomere attrition, epigenetic alterations, loss of proteostasis, disabled autophagy, deregulated nutrient sensing, mitochondrial dysfunction, cellular senescence, stem-cell exhaustion, altered intercellular communication, chronic inflammation, dysbiosis).
- Scoring: per-cell/module scores are computed from mean z-scored expression of set genes (ssGSEA-style enrichment on aggregated profiles); the explorer on Aging Hallmarks maps these scores onto the atlas.
- Summarization: results are meta-analyzed across the included aging datasets, with each dataset's contribution shown where available.
Knowledge graph, RAG & research agent #
- Curated records: GenePedia / LiteraturePedia / ScholarPedia entries are built from manually collected papers; the verified set currently contains 131 manually checked records and is versioned in
longevity_literature_verified.json.
- RAG evidence search (Evidence Search): questions are matched to curated passages by embedding similarity; answers render the retrieved records deterministically with their citations — no free-form generation without sources. The retrieval corpus is larger than the manually verified subset, so generated answers may include content from unverified sources and should be independently confirmed.
- Research agent (Research Agent): an LLM planner constrained to the site's own APIs (atlas, knowledge base, RAG) as tools; every claim in its reports must trace back to a tool result.
Lifespan analysis #
- Lifespan gene annotation: genes associated with lifespan extension or shortening are collected from literature curation; Lifespan tables distinguish evidence types (genetic intervention, association, expression signature).
- Genotype context: population-genotype associations shown in the genotype modules are labelled with their source study and cohort.
Human immune cohort #
- Donor QC: donor-level metadata (age, sex, cohort) in
china_donor_manifest.tsv is checked for consistency; samples failing QC are excluded from aggregated views.
- Multi-omics: the cohort integrates scRNA-seq (CancerSCEM, 173 count matrices), bulk RNA-seq and metabolomics; cross-omics comparisons use donor-matched subsets only.
- Batch handling: visualization embeddings are computed within dataset after standard batch correction; bulk and single-cell results are never pooled directly.
- Ethics & privacy: all human donor data integrated in LACA are de-identified and derive from published cohorts that were collected under their own institutional ethics approvals with informed consent.