[[project skin]] This document describes the general workflow and analysis steps performed by the `10x_persomed_fixed.py` script for processing and analyzing 10x Genomics single-cell RNA-seq data from skin samples. ## Overview The pipeline is designed to: - Load and preprocess single-cell gene expression data from multiple 10x Genomics samples - Generate metadata and assign biological phenotypes - Normalize and scale gene expression data - Calculate immune signature scores - Perform differential expression analysis - Visualize results (UMAP, heatmap, signature plots) - Output summary statistics and processed data tables ## Data Structure The analysis expects the following directory structure for raw 10x Genomics data: ``` data/data-all/ ├── Skin_N__GEX/filtered_feature_bc_matrix/ ├── Skin_02__GEX/filtered_feature_bc_matrix/ └── Skin_03__GEX/filtered_feature_bc_matrix/ ``` Each `filtered_feature_bc_matrix` directory should contain the standard 10x files: `matrix.mtx.gz`, `features.tsv.gz`, and `barcodes.tsv.gz`. ## Main Analysis Steps 1. **Data Loading** - Loads preprocessed AnnData (`adata_gex.h5ad`) if available, otherwise reads raw 10x data from the specified directories. - Concatenates data from all samples into a single AnnData object. 2. **Metadata Generation** - Constructs a metadata table for all cells, including sample ID, condition (e.g., Normal/Lesion), and phenotype (e.g., Healthy/AD). - Adds dummy demographic data (age, gender, location) for demonstration. 3. **Gene Panel Loading** - Loads lists of housekeeping genes and immune signature genes from CSV files (or uses defaults if not found). 4. **Normalization and Scaling** - Normalizes gene expression counts per cell and applies log-transformation. - Scales gene expression values to unit variance. 5. **Signature Scoring** - Calculates immune signature scores (e.g., Th1, Th2, Th17, Neutro, Macro, Eosino, IFN) for each cell/sample based on scaled expression of marker genes. - Applies thresholds and logistic transformation to generate interpretable scores. 6. **Differential Expression Analysis** - Performs t-tests for each signature gene between different phenotypes (e.g., Healthy vs. AD) using sentinel samples. - Outputs log fold-changes and p-values for each gene. 7. **Visualization** - **UMAP**: Projects high-dimensional gene expression data into 2D for visualization of sample/phenotype clustering. - **Heatmap**: Shows expression of signature genes across samples, colored by phenotype. - **Signature Plots**: Generates density and boxplots of signature scores by phenotype. 8. **Summary Statistics and Output** - Prints summary statistics (number of samples, cells, genes, phenotype distribution, etc.). - Saves metadata, signature scores, and DEG results as CSV files in the `output/` directory. - Saves all plots (UMAP, heatmap, signature plots) as PDF files in the `output/` directory. ## Output Files - `output/metadata_10x.csv`: Cell/sample metadata - `output/signature_scores_10x.csv`: Immune signature scores - `output/deg_results_{phenotype}_10x.csv`: Differential expression results for each phenotype - `output/fig1UMAP_10x.pdf`: UMAP plot - `output/fig1Heatmap_10x.pdf`: Heatmap of signature gene expression - `output/fig2Density_10x.pdf`: Density plots of signature scores - `output/fig2Boxplot_10x.pdf`: Boxplots of signature scores ## Notes - The pipeline is modular and can be adapted to other 10x Genomics datasets by updating the sample paths. - Dummy demographic and treatment data are generated for demonstration and should be replaced with real metadata for actual studies. - The script is designed for exploratory analysis and can be extended for more advanced statistical testing or visualization as needed.