Data Processing & Cell Annotation
How every atlas view in LACA is generated — from raw sequencing reads and published expression matrices to harmonised cross-species cell-type labels. This page documents each processing step, the level at which batch correction is applied, and how annotations were validated. A per-dataset breakdown is available in the Dataset Processing Table.
01Pipeline overview
LACA integrates newly generated sequencing libraries with published single-cell and single-nucleus datasets. All datasets pass through a common core pipeline; the only difference is the entry point (raw reads vs. author-processed matrices). Every step below is recorded per dataset in the Dataset Processing Table.
Track A · Raw reads
In-house FASTQ → alignment & quantification
(seeksoultools v1.2.2, species-specific references)
Track B · Public matrices
Author-processed count matrices curated from published atlases (per-study list in Table S4.1 of the manuscript)
Shared quality control
Doublet removal (DoubletFinder) · 200–6,000 detected genes · mitochondrial fraction ≤ 15% · library-level confidence audit
Normalisation & ortholog projection
Scanpy normalise to 10,000 counts + log-transform · projection onto a unified human ortholog feature space (BioMart / OrthoFinder)
Integration & batch correction
Harmony / scVI / BBKNN benchmarked with scIB · BBKNN (species as batch key) for cross-species UMAP · level-dependent application (details)
BioID annotation & expert curation
CellTypist label transfer (tissue-matched reference) → LLM-assisted candidate interpretation (BioReasoner) → expert-curated final labels in a 148-population canonical taxonomy
02Data inputs: two ingestion tracks
Each dataset enters the pipeline through exactly one of two tracks; the track is recorded in the Input type column of the Dataset Processing Table.
- Track A — raw sequencing data (FASTQ). Newly profiled
libraries are aligned and quantified uniformly with
seeksoultools v1.2.2against species-specific reference genomes (manuscript Table S1.3). Where a species-specific assembly is unavailable, the closest available reference is used and this is recorded (e.g. hamadryas baboon reads were aligned to the olive baboon Papio anubis Panu_3.0 assembly). - Track B — previously processed expression matrices. Public datasets are curated from the studies listed in manuscript Table S4.1 (human and non-human primate atlases, rodent resources, additional mammals, and bird, reptile, amphibian, fish and other chordate resources). Author counts are re-filtered and re-normalised with the same criteria as Track A before integration — author-normalised values are never used directly.
Platform composition (10x Genomics chemistries, Smart-seq family, and others) is recorded per dataset. Platform is treated as a batch covariate during integration (see section 05), and cross-platform concordance is examined during annotation review.
03Read processing & quality control
Cell-level filtering (all datasets, both tracks)
- Doublets are identified in Seurat after normalisation, selection of the top 2,000
highly variable genes and PCA;
DoubletFinder v2.0.4is run on the first 10 principal components with parameters optimised by maximising the BCmetric and an assumed doublet rate of 5% per sample. - Cells/nuclei with fewer than 200 or more than 6,000 detected genes, or with mitochondrial gene content above 15%, are removed. Singlet matrices are carried into Scanpy.
Library-level confidence audit
Each newly generated library is additionally scored on recovered cell number, median detected genes, median UMI count, mitochondrial fraction, doublet score, annotation confidence and tissue-level plausibility. Of the 422 in-house libraries, 288 were classified as high confidence, 72 as requiring caution, and 62 as not recommended for coverage-sensitive analyses (manuscript Table S1.5). These classes are used to audit — but not automatically erase — species–tissue coverage. Per-atlas QC metrics are also browsable on the QC Statistics page.
04Normalisation & ortholog feature space
- Filtered count matrices are normalised to 10,000 counts per transcriptome and log-transformed with default Scanpy settings.
- Species-specific expression matrices are projected onto a unified
human ortholog feature space. Ortholog mappings are taken from
Ensembl BioMart where available; for absent species, orthology is inferred with
OrthoFinder v2.5.5from NCBI RefSeq annotations or primary-literature proteomes (manuscript Table S1.4). - One-to-one and many-to-one mappings are retained; counts of species-specific paralogs are summed to represent the corresponding human ortholog. Ortholog coverage and one-to-one vs. many-to-one proportions are summarised per species, and key cross-species comparisons are designed to be repeatable in a one-to-one-ortholog-only feature space.
- Datasets from multiple species are concatenated with an outer join over the union of orthologous genes; PCA is performed on the full gene space.
05Batch correction: at what level, with what
Batch correction in LACA is level-dependent by design — there is no single global correction. The table below states exactly where each method is applied and which covariate is used as the batch key.
| Level | What is corrected | Method | Batch key | Where it shows |
|---|---|---|---|---|
| Within dataset | Sequencing libraries / donors / samples of one accession | Library-level QC audit; covariates retained for downstream checks | Library / donor ID | Per-dataset QC metrics (Statistics page) |
| Within species, multi-dataset | Multiple studies/platforms of the same species | Harmony or BBKNN (selected per species by scIB evaluation) | Dataset / study accession (e.g. FolderID) |
Species atlas pages |
| Cross-species (atlas level) | Species-specific batch effects in the shared ortholog space | BBKNN (Harmony and scVI were benchmarked against it with scIB metrics plus biological interpretability of tissue/cell-type structure) | Species | Main cross-species UMAPs; sub-clustering of each lineage |
| Analyses | — | Pseudo-bulk conservation analyses are checked in both integrated and non-integrated expression spaces where feasible | — | Comparative/conservation results |
06Cell-type annotation: the BioID framework
Standardised labels in LACA are produced by BioID, a hybrid framework combining machine-learning label transfer, LLM-assisted interpretation and mandatory expert curation. Annotation is performed independently for each sequencing library, then harmonised into a unified taxonomy of 148 canonical cell populations.
- Feature selection. Top 5,000 highly variable genes per query
library (Seurat method,
scanpy.pp.highly_variable_genes). - Hierarchical, tissue-aware reference selection. For multi-organ datasets the rhesus macaque atlas is the primary reference (broad tissue coverage, close evolutionary relationship to human). For brain-only datasets the human brain atlas is used directly. When a tissue is absent from the rhesus reference, the corresponding crab-eating macaque (Macaca fascicularis) tissue is used where available, otherwise the matched human tissue atlas.
- Label transfer. Reference data are restricted to query-derived
HVGs and a library-specific
CellTypistmodel is trained to generate preliminary labels. - LLM-assisted interpretation (BioReasoner). Query libraries are over-clustered; the top 50 marker genes per cluster (ranked by log2FC) together with the distribution of CellTypist labels are converted into structured prompts. The LLM only proposes candidate biological interpretations — it never assigns final labels automatically.
- Expert curation. Candidate labels are manually reviewed against canonical marker genes (a curated panel of 532 markers), tissue context, CellTypist probabilities and cross-library consistency before final labels are assigned. Anatomical names are harmonised through a standardised tissue mapping table.
07Validation & concordance with original annotations
- Benchmarking across tissues and taxa. BioID was benchmarked on the gray mouse lemur (Microcebus murinus) brain dataset, with additional validation on zebrafish (Danio rerio) and axolotl (Ambystoma mexicanum) brain datasets — comparing standalone CellTypist label transfer, LLM-assisted interpretation and the final expert-curated labels against published expert annotations.
- Concordance metrics. Annotation concordance and marker support are evaluated with confusion matrices, Sankey diagrams and canonical marker expression. Failure modes (ambiguous clusters, low-confidence labels, marker-poor cell states) are documented explicitly.
- Functional coherence. Cell-type differentially expressed genes are checked by GO and KEGG enrichment (manuscript Tables S1.9–S1.11).
- External reference similarity. Matched pseudobulk profiles are correlated (Pearson) against Tabula Sapiens 2.0 over the shared ortholog space, summarised at cell-type, tissue and organ-system levels.
- Per-dataset concordance. Where a public dataset ships with original author annotations, the agreement between author labels and LACA standardised labels is reported per dataset in the Concordance column of the Dataset Processing Table.
Senescence-associated signature scores
The SenMayo panel and the curated 8-gene senescence-associated signature are gene-set scores computed from transcriptomes alone (scanpy score_genes with expression-matched control genes). They represent senescence-associated transcriptional programs, can overlap with inflammatory and immune-activation programs in immune populations, and do not by themselves establish cellular senescence or quantify senescent cells.
08Provisional datasets policy
A subset of atlas pages currently displays data that have not yet passed through the full BioID standardisation. These pages are explicitly marked, and their annotation status is visible both on the page and in the dataset table:
| Badge | Meaning | What you see on the page |
|---|---|---|
| BioID-standardised | Full pipeline: shared QC → ortholog projection → level-appropriate integration → BioID annotation with expert curation | Standard UMAP/gene-search functionality; canonical 148-population labels |
| Original author labels | Public dataset with author cell-type annotations, re-filtered and re-embedded by LACA but not yet re-annotated | “Provisional atlas release” disclosure banner; gene search disabled |
| Provisional clusters | Re-computed embedding (normalise → HVG → PCA → UMAP → Leiden 0.8) with unannotated clusters as placeholders | “Provisional atlas release” disclosure banner; gene search disabled; replaced when BioID standardisation completes |
09Reproducibility, versions & downloads
Per-dataset processing record
The Dataset Processing Table lists, for every dataset: ingestion track, platform, QC thresholds and pre/post-QC cell counts, normalisation, integration method and batch key, annotation source and reference, concordance with original annotations, and pipeline version/date. The full table is downloadable as CSV/JSON.
Annotation audit trail
For BioID-standardised datasets, prompt templates, model identifiers, model outputs, CellTypist probability summaries, marker tables and final curator decisions are released as a structured BioID audit table together with the analysis code.
Key software versions
| Step | Tool | Version |
|---|---|---|
| Alignment & quantification (Track A) | seeksoultools | v1.2.2 |
| Doublet detection | DoubletFinder (Seurat workflow) | v2.0.4 |
| Filtering / normalisation / embedding | Scanpy | see audit table |
| Orthology inference | OrthoFinder + Ensembl BioMart | v2.5.5 |
| Integration (benchmarked) | Harmony · scVI · BBKNN; evaluated with scIB | see audit table |
| Label transfer | CellTypist (library-specific models) | see audit table |
| LLM-assisted interpretation | BioReasoner (model IDs in audit table) | see audit table |
Questions about a specific dataset’s processing can be directed to the consortium via the contact page.
→ Open the Dataset Processing Table