P-2026-02 FIG. 014 EPCOT Foundation Model: Data & Training Pipeline for TSS-to-Expression Prediction
An end-to-end pipeline turning raw HPAP snMultiome data into per-donor, per-cell-type pseudobulk — ATAC single-cell 1 bp signal into 600 kb-windowed HDF5 arrays, and RNA counts into CP10K-normalized 19,264-gene vectors across 7 cell types. Training is sped up 2–5× per epoch at numerical parity via I/O prefetch, pinned H2D transfers, and vectorized masking; reset-head staged fine-tuning improves generalization on limited data; and a ~50 kb TSS-bin offset bug was found and fixed for cleaner interpretability.