[[project skin]]
This document describes the general workflow and analysis steps performed by the `10x_persomed_fixed.py` script for processing and analyzing 10x Genomics single-cell RNA-seq data from skin samples.
## Overview
The pipeline is designed to:
- Load and preprocess single-cell gene expression data from multiple 10x Genomics samples
- Generate metadata and assign biological phenotypes
- Normalize and scale gene expression data
- Calculate immune signature scores
- Perform differential expression analysis
- Visualize results (UMAP, heatmap, signature plots)
- Output summary statistics and processed data tables
## Data Structure
The analysis expects the following directory structure for raw 10x Genomics data:
```
data/data-all/
├── Skin_N__GEX/filtered_feature_bc_matrix/
├── Skin_02__GEX/filtered_feature_bc_matrix/
└── Skin_03__GEX/filtered_feature_bc_matrix/
```
Each `filtered_feature_bc_matrix` directory should contain the standard 10x files: `matrix.mtx.gz`, `features.tsv.gz`, and `barcodes.tsv.gz`.
## Main Analysis Steps
1. **Data Loading**
- Loads preprocessed AnnData (`adata_gex.h5ad`) if available, otherwise reads raw 10x data from the specified directories.
- Concatenates data from all samples into a single AnnData object.
2. **Metadata Generation**
- Constructs a metadata table for all cells, including sample ID, condition (e.g., Normal/Lesion), and phenotype (e.g., Healthy/AD).
- Adds dummy demographic data (age, gender, location) for demonstration.
3. **Gene Panel Loading**
- Loads lists of housekeeping genes and immune signature genes from CSV files (or uses defaults if not found).
4. **Normalization and Scaling**
- Normalizes gene expression counts per cell and applies log-transformation.
- Scales gene expression values to unit variance.
5. **Signature Scoring**
- Calculates immune signature scores (e.g., Th1, Th2, Th17, Neutro, Macro, Eosino, IFN) for each cell/sample based on scaled expression of marker genes.
- Applies thresholds and logistic transformation to generate interpretable scores.
6. **Differential Expression Analysis**
- Performs t-tests for each signature gene between different phenotypes (e.g., Healthy vs. AD) using sentinel samples.
- Outputs log fold-changes and p-values for each gene.
7. **Visualization**
- **UMAP**: Projects high-dimensional gene expression data into 2D for visualization of sample/phenotype clustering.
- **Heatmap**: Shows expression of signature genes across samples, colored by phenotype.
- **Signature Plots**: Generates density and boxplots of signature scores by phenotype.
8. **Summary Statistics and Output**
- Prints summary statistics (number of samples, cells, genes, phenotype distribution, etc.).
- Saves metadata, signature scores, and DEG results as CSV files in the `output/` directory.
- Saves all plots (UMAP, heatmap, signature plots) as PDF files in the `output/` directory.
## Output Files
- `output/metadata_10x.csv`: Cell/sample metadata
- `output/signature_scores_10x.csv`: Immune signature scores
- `output/deg_results_{phenotype}_10x.csv`: Differential expression results for each phenotype
- `output/fig1UMAP_10x.pdf`: UMAP plot
- `output/fig1Heatmap_10x.pdf`: Heatmap of signature gene expression
- `output/fig2Density_10x.pdf`: Density plots of signature scores
- `output/fig2Boxplot_10x.pdf`: Boxplots of signature scores
## Notes
- The pipeline is modular and can be adapted to other 10x Genomics datasets by updating the sample paths.
- Dummy demographic and treatment data are generated for demonstration and should be replaced with real metadata for actual studies.
- The script is designed for exploratory analysis and can be extended for more advanced statistical testing or visualization as needed.