Data Processing & Cell Annotation

How every atlas view in LACA is generated — from raw sequencing reads and published expression matrices to harmonised cross-species cell-type labels. This page documents each processing step, the level at which batch correction is applied, and how annotations were validated. A per-dataset breakdown is available in the Dataset Processing Table.

01Pipeline overview

LACA integrates newly generated sequencing libraries with published single-cell and single-nucleus datasets. All datasets pass through a common core pipeline; the only difference is the entry point (raw reads vs. author-processed matrices). Every step below is recorded per dataset in the Dataset Processing Table.

02Data inputs: two ingestion tracks

Each dataset enters the pipeline through exactly one of two tracks; the track is recorded in the Input type column of the Dataset Processing Table.

Platform composition (10x Genomics chemistries, Smart-seq family, and others) is recorded per dataset. Platform is treated as a batch covariate during integration (see section 05), and cross-platform concordance is examined during annotation review.

03Read processing & quality control

Cell-level filtering (all datasets, both tracks)

Library-level confidence audit

Each newly generated library is additionally scored on recovered cell number, median detected genes, median UMI count, mitochondrial fraction, doublet score, annotation confidence and tissue-level plausibility. Of the 422 in-house libraries, 288 were classified as high confidence, 72 as requiring caution, and 62 as not recommended for coverage-sensitive analyses (manuscript Table S1.5). These classes are used to audit — but not automatically erase — species–tissue coverage. Per-atlas QC metrics are also browsable on the QC Statistics page.

04Normalisation & ortholog feature space

05Batch correction: at what level, with what

Batch correction in LACA is level-dependent by design — there is no single global correction. The table below states exactly where each method is applied and which covariate is used as the batch key.

LevelWhat is correctedMethodBatch keyWhere it shows
Within dataset Sequencing libraries / donors / samples of one accession Library-level QC audit; covariates retained for downstream checks Library / donor ID Per-dataset QC metrics (Statistics page)
Within species, multi-dataset Multiple studies/platforms of the same species Harmony or BBKNN (selected per species by scIB evaluation) Dataset / study accession (e.g. FolderID) Species atlas pages
Cross-species (atlas level) Species-specific batch effects in the shared ortholog space BBKNN (Harmony and scVI were benchmarked against it with scIB metrics plus biological interpretability of tissue/cell-type structure) Species Main cross-species UMAPs; sub-clustering of each lineage
Analyses — Pseudo-bulk conservation analyses are checked in both integrated and non-integrated expression spaces where feasible — Comparative/conservation results
Study, tissue, platform and species differences are therefore handled explicitly: tissue and platform enter as covariates at the within-species level, while species is the batch key at the atlas level. The exact integration method and batch key used for each displayed dataset are recorded in the Dataset Processing Table.

06Cell-type annotation: the BioID framework

Standardised labels in LACA are produced by BioID, a hybrid framework combining machine-learning label transfer, LLM-assisted interpretation and mandatory expert curation. Annotation is performed independently for each sequencing library, then harmonised into a unified taxonomy of 148 canonical cell populations.

  1. Feature selection. Top 5,000 highly variable genes per query library (Seurat method, scanpy.pp.highly_variable_genes).
  2. Hierarchical, tissue-aware reference selection. For multi-organ datasets the rhesus macaque atlas is the primary reference (broad tissue coverage, close evolutionary relationship to human). For brain-only datasets the human brain atlas is used directly. When a tissue is absent from the rhesus reference, the corresponding crab-eating macaque (Macaca fascicularis) tissue is used where available, otherwise the matched human tissue atlas.
  3. Label transfer. Reference data are restricted to query-derived HVGs and a library-specific CellTypist model is trained to generate preliminary labels.
  4. LLM-assisted interpretation (BioReasoner). Query libraries are over-clustered; the top 50 marker genes per cluster (ranked by log2FC) together with the distribution of CellTypist labels are converted into structured prompts. The LLM only proposes candidate biological interpretations — it never assigns final labels automatically.
  5. Expert curation. Candidate labels are manually reviewed against canonical marker genes (a curated panel of 532 markers), tissue context, CellTypist probabilities and cross-library consistency before final labels are assigned. Anatomical names are harmonised through a standardised tissue mapping table.
Where a label-transfer reference also contributes cells to the integrated atlas, agreement with transferred labels is not treated as independent validation; final labels additionally require marker, tissue-context and cross-library support. Labels unsupported by canonical markers or tissue context are treated as ambiguous and either merged into broader categories or excluded from fine-grained interpretation.

07Validation & concordance with original annotations

Senescence-associated signature scores

The SenMayo panel and the curated 8-gene senescence-associated signature are gene-set scores computed from transcriptomes alone (scanpy score_genes with expression-matched control genes). They represent senescence-associated transcriptional programs, can overlap with inflammatory and immune-activation programs in immune populations, and do not by themselves establish cellular senescence or quantify senescent cells.

08Provisional datasets policy

A subset of atlas pages currently displays data that have not yet passed through the full BioID standardisation. These pages are explicitly marked, and their annotation status is visible both on the page and in the dataset table:

BadgeMeaningWhat you see on the page
BioID-standardised Full pipeline: shared QC → ortholog projection → level-appropriate integration → BioID annotation with expert curation Standard UMAP/gene-search functionality; canonical 148-population labels
Original author labels Public dataset with author cell-type annotations, re-filtered and re-embedded by LACA but not yet re-annotated “Provisional atlas release” disclosure banner; gene search disabled
Provisional clusters Re-computed embedding (normalise → HVG → PCA → UMAP → Leiden 0.8) with unannotated clusters as placeholders “Provisional atlas release” disclosure banner; gene search disabled; replaced when BioID standardisation completes

09Reproducibility, versions & downloads

Per-dataset processing record

The Dataset Processing Table lists, for every dataset: ingestion track, platform, QC thresholds and pre/post-QC cell counts, normalisation, integration method and batch key, annotation source and reference, concordance with original annotations, and pipeline version/date. The full table is downloadable as CSV/JSON.

Annotation audit trail

For BioID-standardised datasets, prompt templates, model identifiers, model outputs, CellTypist probability summaries, marker tables and final curator decisions are released as a structured BioID audit table together with the analysis code.

Key software versions

StepToolVersion
Alignment & quantification (Track A)seeksoultoolsv1.2.2
Doublet detectionDoubletFinder (Seurat workflow)v2.0.4
Filtering / normalisation / embeddingScanpysee audit table
Orthology inferenceOrthoFinder + Ensembl BioMartv2.5.5
Integration (benchmarked)Harmony · scVI · BBKNN; evaluated with scIBsee audit table
Label transferCellTypist (library-specific models)see audit table
LLM-assisted interpretationBioReasoner (model IDs in audit table)see audit table

Questions about a specific dataset’s processing can be directed to the consortium via the contact page.

→ Open the Dataset Processing Table