[
    {
        "id": "thesis:18849",
        "collection": "thesis",
        "collection_id": "18849",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06252026-142447562",
        "type": "thesis",
        "title": "Multi-Scale Systems Analysis of the Squid-Vibrio Symbiosis",
        "author": [
            {
                "family_name": "Beilinson",
                "given_name": "Vera Michelle",
                "orcid": "0000-0002-1259-733X",
                "clpid": "Beilinson-Vera-Michelle"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "McFall-Ngai",
                "given_name": "Margaret J.",
                "orcid": "0000-0002-6046-6238",
                "clpid": "McFall-Ngai-Margaret-J"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Karthikeyan",
                "given_name": "Smruthi",
                "orcid": "0000-0001-6226-4536",
                "clpid": "Karthikeyan-Smruthi"
            },
            {
                "family_name": "Mazmanian",
                "given_name": "Sarkis K.",
                "orcid": "0000-0003-2713-1513",
                "clpid": "Mazmanian-S-K"
            },
            {
                "family_name": "Ruby",
                "given_name": "Edward  G.",
                "orcid": "0000-0002-4112-4830",
                "clpid": "Ruby-Edward"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "McFall-Ngai",
                "given_name": "Margaret J.",
                "orcid": "0000-0002-6046-6238",
                "clpid": "McFall-Ngai-Margaret-J"
            }
        ],
        "local_group": [
            {
                "literal": "div_biol"
            }
        ],
        "abstract": "<p>Animal\u2013microbe symbioses are fundamental to the biology of nearly all multicellular organisms, yet many questions remain regarding how hosts recognize beneficial microbes and establish stable associations. The symbiosis between the Hawaiian bobtail squid, Euprymna scolopes, and the bioluminescent bacterium Vibrio fischeri provides a powerful model for investigating these processes.</p>\r\n\r\n<p>In this dissertation, I examined host\u2013microbe interactions across multiple biological scales, from bacterial genetic variation and tissue-level responses to cellular mechanisms of symbiosis. I demonstrate that bacterial strain identity influences host transcriptional responses. I further identify the skin as an early site of host\u2013microbe interaction, revealing that hosts respond to symbiont exposure prior to stable colonization of the light organ. Finally, I establish methodologies for single-cell and single-nucleus transcriptomics in E. scolopes, providing a foundation for investigating symbiosis at cellular resolution.</p>\r\n\r\n<p>Together, these findings show that host responses to symbiotic bacteria emerge from interconnected processes spanning bacterial genetics, host tissues, and individual cell types. This work advances our understanding of how beneficial host\u2013microbe partnerships are established and maintained and provides new tools for studying symbiosis in cephalopods and other animal systems.</p>",
        "doi": "10.7907/q2b1-xn16",
        "publication_date": "2027",
        "thesis_type": "phd",
        "thesis_year": "2027"
    },
    {
        "id": "thesis:17880",
        "collection": "thesis",
        "collection_id": "17880",
        "cite_using_url": "https://resolver.caltech.edu/CaltechThesis:02102026-091429391",
        "primary_object_url": {
            "basename": "Thesis_final_CF.pdf",
            "content": "final",
            "filesize": 8081793,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17880/1/Thesis_final_CF.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Biophysical Modeling for Gene Expression and Evolution",
        "author": [
            {
                "family_name": "Felce",
                "given_name": "Catherine E.",
                "orcid": "0009-0009-9909-6711",
                "clpid": "Felce-Catherine-E"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Pennell",
                "given_name": "Matthew",
                "orcid": "0000-0002-2886-3970",
                "clpid": "Pennell-Matthew"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_pma"
            }
        ],
        "abstract": "Principled biophysical modeling is a necessary foundation for analyzing RNA sequencing data. In recent years, higher quality data for other data modalities at single-cell resolution have become available. I present joint biophysical models combining two of these modalities, chromatin accessibility measurements (ATAC-seq) and protein counts, individually with single-cell transcriptomic data, and give preliminary data results. I consider the extension of biophysically motivated models to the field of phylogenetics. I present competing mechanistic hypotheses for gene expression evolution and test them via parametrized single-cell cross-species data. I also consider a physics-inspired model for population-level evolution via maternal effects and interacting subpopulations.",
        "doi": "10.7907/chmp-kt37",
        "publication_date": "2026",
        "thesis_type": "phd",
        "thesis_year": "2026"
    },
    {
        "id": "thesis:18729",
        "collection": "thesis",
        "collection_id": "18729",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06012026-033232974",
        "primary_object_url": {
            "basename": "Thesis_Carilli_Maria.pdf",
            "content": "final",
            "filesize": 40294216,
            "license": "cc_by_nc_nd",
            "mime_type": "application/pdf",
            "url": "/18729/2/Thesis_Carilli_Maria.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Genetic Interrogation of Expression Regulation",
        "author": [
            {
                "family_name": "Carilli",
                "given_name": "Maria Theresa Natalina",
                "orcid": "0000-0002-8977-7224",
                "clpid": "Carilli-Maria-Theresa-Natalina"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            },
            {
                "family_name": "Yue",
                "given_name": "Yisong",
                "orcid": "0000-0001-9127-1989",
                "clpid": "Yue-Yisong"
            },
            {
                "family_name": "Engelhardt",
                "given_name": "Barbara E.",
                "orcid": "0000-0002-6139-7334",
                "clpid": "Engelhardt-Barbara-E"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Understanding the effects of variation in the genome on organisms' phenotypes is a central goal of biology. For the past several decades, this has been done by performing large-scale association tests to identify links between genetic variants across the genome and particular diseases or traits. As most variants are in non-coding regions of the genome and act only in specific contexts, there arose the intermediate step of associating variants with gene expression patterns in particular tissues using bulk RNA-sequencing data and, with the recent widespread adoption of single-cell RNA-sequencing, particular cell types. However, associative tests are generally restricted to variants within some sequence distance of the gene, as testing tens of millions of distal variants against tens of thousands of genes in hundreds of contexts is computationally and statistically burdensome. Moreover, associating these proximal variants to changes in average gene expression using scRNA-seq data is far from pinpointing the cellular process they may affect. To move beyond analysis of averages, recent work has advocated for a mechanistic approach to analyzing scRNA-seq by modeling the underlying biological processes of transcription, splicing, and degradation.</p>\r\n\r\n<p>In this thesis, we first show how biophysical models are useful for identifying changes in biophysical processes not accessible at the level of mean expression. We next develop accelerated inference procedures for these models to make feasible their application at the scale required for genetic association tests. We then propose a framework for testing for the presence or absence of proximal or distal gene regulation using homozygous crosses. Finally, we couple the genetic testing framework and biophysical models to identify regulatory strategies of biophysical processes across cell types in eight tissues of eight genetically diverse mouse strains.</p>",
        "doi": "10.7907/qnzh-rf18",
        "publication_date": "2026",
        "thesis_type": "phd",
        "thesis_year": "2026"
    },
    {
        "id": "thesis:18394",
        "collection": "thesis",
        "collection_id": "18394",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:02262026-191401810",
        "primary_object_url": {
            "basename": "markarian_nicholas_thesis.pdf",
            "content": "final",
            "filesize": 55194097,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/18394/3/markarian_nicholas_thesis.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Mixtures of Latent Variable Models for Interpreting Gene Expression Covariation from Pathways to Transcriptomes",
        "author": [
            {
                "family_name": "Markarian",
                "given_name": "Nicholas",
                "orcid": "0000-0003-1347-2392",
                "clpid": "Markarian-Nicholas"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Sternberg",
                "given_name": "Paul W.",
                "orcid": "0000-0002-7699-0173",
                "clpid": "Sternberg-P-W"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Sternberg",
                "given_name": "Paul W.",
                "orcid": "0000-0002-7699-0173",
                "clpid": "Sternberg-P-W"
            },
            {
                "family_name": "Bois",
                "given_name": "Justin S.",
                "orcid": "0000-0001-7137-8746",
                "clpid": "Bois-J-S"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "This thesis centers on interpretable subspace learning and latent variable models for characterizing covariation modulated by categorical variables in the context of biology. First, it introduces a probabilistic model with ties to Principal Component Analysis and k-means clustering, k-spaces, which has implications across different biological analyses through its interpretations as a subspace learning technique, a latent variable model, and a dimension reduction technique. Second, it establishes the problem of simultaneously characterizing gene covariation and expression level in known pathways in human tissue samples and applies k-spaces to GTEx data to lay the foundations for this line of research. Finally, it outlines a path forward to being able to use such data as references for clinical samples from patients.",
        "doi": "10.7907/dcbd-wy35",
        "publication_date": "2026",
        "thesis_type": "phd",
        "thesis_year": "2026"
    },
    {
        "id": "thesis:17005",
        "collection": "thesis",
        "collection_id": "17005",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:02172025-204256436",
        "primary_object_url": {
            "basename": "CanLi2025thesis.pdf",
            "content": "final",
            "filesize": 41960912,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17005/1/CanLi2025thesis.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Gene Regulatory Analysis of the Developing Enteric Nervous System of Zebrafish (Danio rerio)",
        "author": [
            {
                "family_name": "Li",
                "given_name": "Can",
                "orcid": "0000-0002-5352-6212",
                "clpid": "Li-Can"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Bronner",
                "given_name": "Marianne E.",
                "orcid": "0000-0003-4274-1862",
                "clpid": "Bronner-M-E"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Stathopoulos",
                "given_name": "Angelike",
                "orcid": "0000-0001-6597-2036",
                "clpid": "Stathopoulos-A"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Prober",
                "given_name": "David A.",
                "orcid": "0000-0002-7371-4675",
                "clpid": "Prober-D-A"
            },
            {
                "family_name": "Bronner",
                "given_name": "Marianne E.",
                "orcid": "0000-0003-4274-1862",
                "clpid": "Bronner-M-E"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "Neural crest cells give rise to the neurons of the enteric nervous system (ENS) that innervate the gastrointestinal tract to regulate gut motility.  The immense size and distinct subregions of the gut present a challenge to understanding the spatial organization and sequential differentiation of different neuronal subtypes. To tackle this, we profile enteric neurons and progenitors at single cell resolution during zebrafish embryonic and larval development to provide a near complete picture of transcriptional changes that accompany emergence of ENS neurons throughout the gastrointestinal tract. Multiplex spatial RNA transcript analysis was then used to reveal the temporal order and distinct localization patterns of neuronal subtypes along the length of the gut. Next, we show that functional perturbation of select transcription factors Ebf1a, Gata3 and Satb2 alters the cell fate choice, respectively, of inhibitory, excitatory and serotonergic neuronal subtypes in the developing ENS. To decipher the molecular mechanism underlying the development of ENS, we further performed single cell ATAC-seq to profile the epigenetic landscape of the developing ENS. Together with CUT&amp;RUN results, we found the master regulator Phox2bb harbors extensive binding sites throughout the genome and plays versatile roles in neuronal differentiation, including regulating progenitor gene Sox10, activating transcription factors Phox2a and Insm1b for early neural development and regulating genes Etv1 and Hmx3a for neuronal differentiation. Integrated with single cell RNA-seq analysis, we further reconstruct a putative gene regulatory circuit involving in the specification and maturation of ENS neurons.",
        "doi": "10.7907/mc6z-3k03",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:17383",
        "collection": "thesis",
        "collection_id": "17383",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06022025-234901151",
        "type": "thesis",
        "title": "The Topology of Cellular Ontogeny",
        "author": [
            {
                "family_name": "Flores-Bautista",
                "given_name": "Emanuel",
                "orcid": "0000-0002-2810-1757",
                "clpid": "Flores-Bautista-Emanuel"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Prober",
                "given_name": "David A.",
                "orcid": "0000-0002-7371-4675",
                "clpid": "Prober-D-A"
            },
            {
                "family_name": "Marcolli",
                "given_name": "Matilde",
                "orcid": "0000-0002-2045-2907",
                "clpid": "Marcolli-M"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "A fundamental goal of modern biology is to build global, predictive models of gene regulation that encompass diverse physiological contexts. Single-cell transcriptomics has enabled the creation of developmental cell atlases--detailed catalogs of gene expression patterns and differentiation trajectories at an organismal scale. The widespread availability of  cell atlases across metazoan model organisms presents an opportunity to construct global theories of cell-state control. In this thesis, we introduce a framework that uses persistent homology to decompose cell atlases into topological structures that provide signatures of gene regulation at the scale of an organism. Using this framework, we found that the topological structure of a broad set of developmental atlases contains only a discrete set of topological structures\u2014such as clusters, trees, and loops\u2014-revealing the recurrent use of global gene regulatory strategies. Our analysis revealed that the tree topology, while predominant, is not universal. Indeed, we identified non-trivial topologies containing loops in the development of human immune cells, seam-hypodermal cells in \\textit{C. elegans}, and the cnidocytes of multiple cnidarians. Analysis of cell-state manifolds with non-trivial topology demonstrated an important role of convergent structures in increasing cellular diversity along paths to a common cell fate, and of cyclic structures in self-renewal of progenitor-like states. Together, this work provides a global perspective on principles of cell-state regulation, and suggests that loops are important organizing structures for controlling cell differentiation.",
        "doi": "10.7907/t8hc-yq15",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:17389",
        "collection": "thesis",
        "collection_id": "17389",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06032025-002120461",
        "primary_object_url": {
            "basename": "Thesis.pdf",
            "content": "final",
            "filesize": 25197916,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17389/1/Thesis.pdf",
            "version": "v5.0.0"
        },
        "type": "thesis",
        "title": "A Biophysical Approach to Normalization and Trajectory Inference in Single-Cell RNA Sequencing Data Analysis",
        "author": [
            {
                "family_name": "Fang",
                "given_name": "Meichen",
                "orcid": "0000-0002-8217-0710",
                "clpid": "Fang-Meichen"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Bois",
                "given_name": "Justin S.",
                "orcid": "0000-0001-7137-8746",
                "clpid": "Bois-J-S"
            },
            {
                "family_name": "Chong",
                "given_name": "Shasha",
                "orcid": "0000-0002-5372-311X",
                "clpid": "Chong-Shasha"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Single-cell genomics assays, particularly single-cell RNA sequencing that enables genome-wide profiling of gene expression, have been driven forward by a combination of technological and computational advances. While producing extraordinary large amounts of data for biological discovery, methods for mining results currently rely heavily on heuristics and lack of modeling has resulted in limited mechanistic biological insight. This thesis presents two models for normalization and trajectory inference in single-cell RNA sequencing analysis to demonstrate how biophysical modeling, when combined with principled statistical inference, can yield interpretable insights grounded in rigorous theoretical frameworks.</p>\r\n\r\n<p>We begin by explaining the two cultures in single-cell RNA sequencing analysis. Next, we present the chemical master equation, which forms the theoretical foundation for biophysically informed stochastic models of gene expression, and explore an existing gap in developing uniform approximations over time under the large-volume limit. Returning to single-cell RNA sequencing data analysis, we introduce two mechanistic models for normalization and trajectory inference, which are essential components of single-cell RNA sequencing analysis.</p>",
        "doi": "10.7907/asek-t904",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:17344",
        "collection": "thesis",
        "collection_id": "17344",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06012025-204136107",
        "primary_object_url": {
            "basename": "Methods_for_long_read_RNA_seq_transcriptomics_Loving_Rebekah_2025.pdf",
            "content": "final",
            "filesize": 51348381,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17344/1/Methods_for_long_read_RNA_seq_transcriptomics_Loving_Rebekah_2025.pdf",
            "version": "v5.0.0"
        },
        "type": "thesis",
        "title": "Methods for Long Read RNA-Seq Transcriptomics",
        "author": [
            {
                "family_name": "Loving Ngo",
                "given_name": "Rebekah Kiana",
                "orcid": "0000-0001-8725-0376",
                "clpid": "Loving-Ngo-Rebekah-Kiana"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            },
            {
                "family_name": "Perona",
                "given_name": "Pietro",
                "orcid": "0000-0002-7583-5809",
                "clpid": "Perona-P"
            },
            {
                "family_name": "Mortazavi",
                "given_name": "Ali",
                "orcid": "0000-0002-4259-6362",
                "clpid": "Mortazavi-Ali"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "While short read RNA-seq dominated the field for decades, long read RNA-seq is particularly useful for isoform-level expression analysis, genome annotation, detecting novelly splicing transcripts, identifying exact breakpoints in gene fusions, and discovering chimeric RNAs. Long read RNA-seq has rapidly scaled to the point of producing terabytes of data from a single set of experiments. Technological advances in RNA and DNA sequencing library preparation, chemistry used in the Oxford nanopores, and basecalling algorithms have reduced long read sequencing error rates to sub-1% error. Further, the cost of long read sequencing has dropped to about one hundred US dollars per human genome. These two factors have lead to the mass production of high-throughput, long read, and single-cell RNA-seq data. While recent tools for long read RNA-seq have been developed, they have not kept pace in scalability and accuracy with long read RNA-seq in the fashion that short read RNA-seq tools have met computational scalability and accuracy challenges. To address this, in this thesis, we leverage long k-mers and pseudoalignment for mapping and quantifying long reads in the novel algorithm implemented within lr-kallisto, which yields both efficiency and higher accuracy for long read mapping and quantification than previous tools. We demonstrate that long read RNA-seq has reached sufficient depth and accuracy to yield accurate quantification of isoform-level expression for differential expression analysis. Furthermore, we explore the feasibilty of also utilizing long k-mers and pseudoalignment in both transcript discovery in dn-kallisto and gene fusion and immune receptor sequence discovery with fugi with measured success. Thus, our tools will enable a more complete, accurate, and scalable analysis of single-cell and bulk RNA-seq than has hitherto been possible in both quantifications and differential expression analysis as well as investigation of gene fusions, chimeric RNAs, and immune receptor sequences without bias.",
        "doi": "10.7907/3nz8-3c83",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:17266",
        "collection": "thesis",
        "collection_id": "17266",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05232025-094543832",
        "primary_object_url": {
            "basename": "caltech-thesis-final-delaney-sullivan.pdf",
            "content": "final",
            "filesize": 19916923,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17266/3/caltech-thesis-final-delaney-sullivan.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Software, Tools, and Methods Development for Single-Cell Transcriptomics",
        "author": [
            {
                "family_name": "Sullivan",
                "given_name": "Delaney Kalcey",
                "orcid": "0000-0002-8359-6705",
                "clpid": "Sullivan-Delaney-Kalcey"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Guttman",
                "given_name": "Mitchell",
                "orcid": "0000-0003-4748-9352",
                "clpid": "Guttman-M"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            },
            {
                "family_name": "Pimentel",
                "given_name": "Harold",
                "orcid": "0000-0001-8556-2499",
                "clpid": "Pimentel-Harold"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Guttman",
                "given_name": "Mitchell",
                "orcid": "0000-0003-4748-9352",
                "clpid": "Guttman-M"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Advances in transcriptomics have transformed the study of gene expression, enabling a shift from low-throughput bulk RNA measurements to high-resolution, large-scale single-cell RNA-sequencing (scRNA-seq). This work refines existing methodologies and introduces new strategies for achieving precise, versatile, and scalable transcriptomic analyses across a broad spectrum of assays and biological contexts.</p>\r\n\r\n<p>On the computational front, this dissertation introduces new methods for adaptable preprocessing of sequencing reads, enabling the handling of very complex read structures. It refines existing strategies for efficiently querying large-scale transcriptomic datasets and enhances approaches for quantifying nascent and mature RNA species. A general framework is introduced for discovering and organizing biologically informative sequences directly from raw sequencing data, facilitating the detection of sample-specific or condition-specific variation. On the experimental front, a novel single-cell RNA sequencing method is presented that is cost-effective, open source, and scalable, supporting large-scale studies with substantial cell numbers and high per-cell resolution.</p>\r\n\r\n<p>These developments collectively expand the toolkit for transcriptomics, enabling more efficient and comprehensive exploration of RNA biology.</p>",
        "doi": "10.7907/kee5-ty36",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:17185",
        "collection": "thesis",
        "collection_id": "17185",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:04292025-233607669",
        "primary_object_url": {
            "basename": "KilYeokyoung2025Thesis.pdf",
            "content": "final",
            "filesize": 8595745,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/17185/1/KilYeokyoung2025Thesis.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Engineering and Computational Tools for Salivary Biomedicine",
        "author": [
            {
                "family_name": "Kil",
                "given_name": "Yeokyoung (Anne)",
                "orcid": "0000-0002-1235-7379",
                "clpid": "Kil-Yeokyoung-Anne"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Gao",
                "given_name": "Wei",
                "orcid": "0000-0002-8503-4562",
                "clpid": "Gao-Wei"
            },
            {
                "family_name": "Burdick",
                "given_name": "Joel Wakeman",
                "orcid": "0000-0002-3091-540X",
                "clpid": "Burdick-J-W"
            },
            {
                "family_name": "Wyllie",
                "given_name": "Anne L",
                "orcid": "0000-0001-6015-0279",
                "clpid": "Wyllie-A-L"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_eng"
            }
        ],
        "abstract": "Saliva is emerging as a powerful biofluid for noninvasive diagnostics, offering a window into human health through its diverse biomolecular composition. This dissertation advances the field of salivary biomedicine by addressing critical challenges in saliva collection, processing, and analysis. First, a comparative analysis of five saliva collection devices highlighted key usability factors, informing the development of SalivaStraw--a novel device designed to improve collection efficiency and minimize leakage. Next, colosseum, a low-cost, open-source fraction collector, was designed and developed to facilitate scalable saliva processing and improve biomarker isolation. Finally, a computational framework leveraging spline regression was applied to longitudinal salivary transcriptomic data, enabling the identification of temporally regulated genes and underscoring saliva\u2019s potential for dynamic health monitoring. Collectively, this work contributes new tools and methodologies that strengthen the foundation of saliva-based diagnostics, broadening its applications in precision medicine and beyond.",
        "doi": "10.7907/8wdh-5v36",
        "publication_date": "2025",
        "thesis_type": "phd",
        "thesis_year": "2025"
    },
    {
        "id": "thesis:16486",
        "collection": "thesis",
        "collection_id": "16486",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06032024-182223499",
        "primary_object_url": {
            "basename": "Thesis_Draft_final_final.pdf",
            "content": "final",
            "filesize": 21944874,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/16486/1/Thesis_Draft_final_final.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Revealing Regulatory Network Organization Through Single-Cell Perturbation Profiling and Maximum Entropy Models",
        "author": [
            {
                "family_name": "Jiang",
                "given_name": "Jialong",
                "orcid": "0000-0001-8560-8397",
                "clpid": "Jiang-Jialong"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Elowitz",
                "given_name": "Michael B.",
                "orcid": "0000-0002-1221-0967",
                "clpid": "Elowitz-M-B"
            },
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "Gene regulatory networks within cells modulate the expression of the genome in response to signals and changing environmental conditions. Reconstructions of gene regulatory networks can reveal the information processing and control principles used by cells to maintain homeostasis and execute cell-state transitions. In this thesis, we introduce a computational framework, D-SPIN, that generates quantitative models of gene regulatory networks from single-cell mRNA-seq datasets collected across thousands of distinct perturbation conditions. D-SPIN models the cell as a collection of interacting gene-expression programs, and constructs a probabilistic model to infer regulatory interactions between gene-expression programs and external perturbations. Using large Perturb-seq and drug-response datasets, we demonstrate that D-SPIN models reveal the organization of cellular pathways, sub-functions of macromolecular complexes, and the logic of cellular regulation of transcription, translation, metabolism, and protein degradation in response to gene knockdown perturbations. D-SPIN can also be applied to dissect drug response mechanisms in heterogeneous cell populations, elucidating how combinations of immunomodulatory drugs can induce novel cell states through additive recruitment of gene expression programs. D-SPIN provides a computational framework for constructing interpretable models of gene-regulatory networks to reveal principles of cellular information processing and physiological control.",
        "doi": "10.7907/5zta-9818",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16438",
        "collection": "thesis",
        "collection_id": "16438",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05292024-221741183",
        "primary_object_url": {
            "basename": "Thesis_Rev_TaraChari.pdf",
            "content": "final",
            "filesize": 11821855,
            "license": "cc_by_nc_nd",
            "mime_type": "application/pdf",
            "url": "/16438/1/Thesis_Rev_TaraChari.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Perturbing the Genome: From Bench to Biophysics",
        "author": [
            {
                "family_name": "Chari",
                "given_name": "Tara Varada",
                "orcid": "0000-0002-6953-4313",
                "clpid": "Chari-Tara-Varada"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Qian",
                "given_name": "Lulu",
                "orcid": "0000-0003-4115-2409",
                "clpid": "Qian-Lulu"
            },
            {
                "family_name": "Murray",
                "given_name": "Richard M.",
                "orcid": "0000-0002-5785-7481",
                "clpid": "Murray-R-M"
            },
            {
                "family_name": "Anderson",
                "given_name": "David J.",
                "orcid": "0000-0001-6175-3872",
                "clpid": "Anderson-D-J"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>In single-cell genomics, we can simultaneously assay hundreds of thousands of cells, their molecular contents, and how they respond to perturbation, from genetic knockouts to environmental changes. This thesis focuses on how to merge experimental and computational techniques to generate and analyze large-scale perturbation data for high-resolution systems biology. Beginning at the bench, we demonstrate how combining large-scale cell atlas surveys with multi-condition experimentation can illuminate the diversity of cell types across whole organisms and cellular strategies in response to environmental changes and perturbations. We then investigate the limitations of current practice in exploratory analysis, and strategies for determining preservation or distortion of biological insight by these data transformation and dimensionality reduction techniques. To address these limitations, we demonstrate how stochastic biophysical models can rewrite the way we interpret complex perturbation data, taking greater advantage of the diverse molecular measurements to develop biological hypotheses about DNA and RNA regulation in cellular function, development, and disease.</p>",
        "doi": "10.7907/5drv-ma07",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16147",
        "collection": "thesis",
        "collection_id": "16147",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:07272023-175309910",
        "primary_object_url": {
            "basename": "Goronzy_CalTechThesis_Final2.pdf",
            "content": "final",
            "filesize": 5387002,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/16147/2/Goronzy_CalTechThesis_Final2.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Higher-Order Chromatin States and Nuclear Structures Regulating Gene Expression",
        "author": [
            {
                "family_name": "Goronzy",
                "given_name": "Isabel Nadine",
                "orcid": "0000-0002-6713-9192",
                "clpid": "Goronzy-Isabel-Nadine"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Guttman",
                "given_name": "Mitchell",
                "orcid": "0000-0003-4748-9352",
                "clpid": "Guttman-M"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Elowitz",
                "given_name": "Michael B.",
                "orcid": "0000-0002-1221-0967",
                "clpid": "Elowitz-M-B"
            },
            {
                "family_name": "Guttman",
                "given_name": "Mitchell",
                "orcid": "0000-0003-4748-9352",
                "clpid": "Guttman-M"
            },
            {
                "family_name": "Chong",
                "given_name": "Shasha",
                "orcid": "0000-0002-5372-311X",
                "clpid": "Chong-Shasha"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_chem"
            }
        ],
        "abstract": "<p>Although the same genome is present in every cell, each cell type orchestrates a distinct gene expression program, which can be rapidly adapted in response to stimuli. Accordingly, gene regulation is a highly complex, context-specific process that involves the dynamic interplay between numerous regulatory factors. Most methods to study these regulatory factors only measure pairwise interactions between molecules and are limited to mapping one regulatory protein at a time. Consequently, the combinatorial complexity of gene regulation at individual genomic loci and the functional consequence of many regulatory factors remain underexplored. To address this, we have developed new sequencing-based approaches and computational analyses to comprehensively profile, at unprecedented scale, the diverse gene regulatory landscape and directly establish the link between regulatory factors and transcriptional outcomes. In Chapter 2, we present Chromatin Immunoprecipitation Done-In-Parallel (ChIP-DIP), a highly multiplexed method for mapping hundreds of proteins to DNA within a single sample. ChIP-DIP increases the throughput of existing methods by &gt; 100-fold and enables the production of consortium-scale, cell type-specific data within a single lab. Capitalizing on the scale and diversity provided by ChIP-DIP, we uncover unique quantitative combinations of histone modifications that define distinctive classes of regulatory elements. Specifically, we find features distinguishing classes of promoters that correspond to different polymerase activity, transcriptional levels, and gene types and find acetylation patterns distinguishing classes of enhancers that exhibit distinct activity states, induction potential, and regulatory potential. Next, in Chapter 3, we apply RNA-DNA SPRITE (RD-SPRITE), a method for simultaneous measurement of RNA and DNA organization, to investigate the functional relationship between genome structure and transcription. We demonstrate that RD-SPRITE precisely detects individual, nascent pre-mRNAs at their transcriptional locus and, as a result, can be used to assess the 3D genome structure present during active transcription. We find that RNA polymerase II transcription occurs within genomic structures previously thought to be inactive, such as the B compartment and DNA regions near the nucleolus. This suggests that active transcription can occur throughout the nucleus and argues against structural domains that preclude transcription. Overall, our findings highlight the ability of RD-SPRITE to establish a structure-function link. Finally, in Chapter 4, we apply RD-SPRITE to study the transcriptional dependence of nuclear organization. We demonstrate that transcriptional inhibition leads to the loss of high-order structure around multiple RNA-processing bodies \u2014 the nucleolus, the scaRNA hub and the histone locus body \u2014 that are responsible for essential nuclear functions such as RNA processing and gene regulation. These findings suggest a role for RNA and nascent transcription in the formation and maintenance of long-range 3D contacts and critical nuclear compartments. In summary, we have developed new approaches to explore epigenomic and organizational complexity within the mammalian nucleus and have uncovered genome-wide principles of gene regulation.</p>",
        "doi": "10.7907/8gm2-jn84",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16280",
        "collection": "thesis",
        "collection_id": "16280",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:01162024-161326655",
        "primary_object_url": {
            "basename": "TDilanyan_Thesis.pdf",
            "content": "final",
            "filesize": 7266593,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/16280/4/TDilanyan_Thesis.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Open-Source Custom Beads for Single-Cell Transcriptomics",
        "author": [
            {
                "family_name": "Dilanyan",
                "given_name": "Taleen Gaied",
                "orcid": "0000-0002-3131-3259",
                "clpid": "Dilanyan-Taleen-Gaied"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Rees",
                "given_name": "Douglas C.",
                "orcid": "0000-0003-4073-1185",
                "clpid": "Rees-D-C"
            },
            {
                "family_name": "Shapiro",
                "given_name": "Mikhail G.",
                "orcid": "0000-0002-0291-4215",
                "clpid": "Shapiro-M-G"
            },
            {
                "family_name": "Hsieh-Wilson",
                "given_name": "Linda C.",
                "orcid": "0000-0001-5661-1714",
                "clpid": "Hsieh-Wilson-L-C"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_chem"
            }
        ],
        "abstract": "Open-source single-cell genomics technologies have helped democratize single-cell genomics and expedite method development. Methods such as inDrops and Drop-seq for single-cell RNA-seq preceded popular technologies such as the 10x Genomics\u2019 Chromium platform, however despite initial enthusiasm for open-source methods, their popularity has waned. A major reason has been the lack of availability of low-cost, customizable beads, which are essential for microfluidics based single-cell RNA-seq. We address this challenge by introducing a new method for producing barcoded hydrogel beads for single-cell RNA-seq called HiPER (High-throughput PER-barcoded hydrogel beads) that allows for increasing the diversity of barcode sequences, reducing manufacturing cost, and that can be readily adapted to custom applications. HiPER barcodes are decoupled from the capture sequences and can therefore be configured to capture RNA, DNA, or tailored for specific-gene enrichment.",
        "doi": "10.7907/v52p-gf80",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16368",
        "collection": "thesis",
        "collection_id": "16368",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05042024-011724418",
        "type": "thesis",
        "title": "Complexity of Transcriptomic Data Analysis and Implications for Biological Discovery",
        "author": [
            {
                "family_name": "Luebbert",
                "given_name": "Laura",
                "orcid": "0000-0003-1379-2927",
                "clpid": "Luebbert-Laura"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Van Valen",
                "given_name": "David A.",
                "orcid": "0000-0001-7534-7621",
                "clpid": "Van-Valen-D"
            },
            {
                "family_name": "Murray",
                "given_name": "Richard M.",
                "orcid": "0000-0002-5785-7481",
                "clpid": "Murray-R-M"
            },
            {
                "family_name": "Bjorkman",
                "given_name": "Pamela J.",
                "orcid": "0000-0002-2277-3990",
                "clpid": "Bjorkman-P-J"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "3MT Competition (Caltech)"
            },
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Over the past decade, the advancement of \u2018omics\u2019 technologies has ushered in a new era for the life sciences. Given the high-throughput nature of omics technologies, this era is characterized by unique computational challenges pertaining to data size and dimensionality, and technical and biological noise. Concurrently, it offers opportunities, as global, untargeted, and parallel measurement of large amounts of information often captures unexpected insights.</p> \r\n\r\n<p>This thesis describes challenges inherent to the omics era of life sciences, particularly highlighting the increasing importance of merging expertise in biology and computer science. It describes the development of multiple software tools designed to address several of these challenges, which were immediately adopted and widely implemented in transcriptomics and proteomics research. Additionally, it contains three chapters focused on unraveling previously unquantifiable information, including the interpretation of sequencing data from organisms with low-quality reference genome assemblies and workflows for identifying novel viruses using single-cell RNA sequencing data already massively generated in research, healthcare, and agriculture.</p>",
        "doi": "10.7907/xnw5-v914",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16399",
        "collection": "thesis",
        "collection_id": "16399",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05202024-215640209",
        "primary_object_url": {
            "basename": "20240524_caltech_thesis_julian_wagner.pdf",
            "content": "final",
            "filesize": 8359789,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/16399/1/20240524_caltech_thesis_julian_wagner.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Uncovering Mechanisms of Host Recognition, Host Finding and Host Specificity",
        "author": [
            {
                "family_name": "Wagner",
                "given_name": "Julian Morgan",
                "orcid": "0000-0003-3406-0450",
                "clpid": "Wagner-Julian-Morgan"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Parker",
                "given_name": "Joseph",
                "orcid": "0000-0001-9598-2454",
                "clpid": "Parker-J"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Dickinson",
                "given_name": "Michael H.",
                "orcid": "0000-0002-8587-9936",
                "clpid": "Dickinson-M-H"
            },
            {
                "family_name": "Hong",
                "given_name": "Elizabeth J.",
                "orcid": "0000-0003-3866-418X",
                "clpid": "Hong-Elizabeth-J"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Parker",
                "given_name": "Joseph",
                "orcid": "0000-0001-9598-2454",
                "clpid": "Parker-J"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "Insect diversification is thought to have been catalyzed by widespread specialization on novel hosts\u2014a process underlying exceptional radiations of phytophagous beetles, lepidopterans, parasitoid wasps, and inordinate lineages of symbionts, predators, and other trophic specialists. The fidelity of such interspecies partnerships is often posited to arise from sensory tuning to host-derived cues, a model supported by studies of neural function in host-specific model species. Abundant literature on parasites also suggest that extrinsic factors, namely dispersal mechanisms and aggressiveness/acceptance from novel hosts, externally enforces host specificity. Here, I first review what is known about host specificity, why it arises and how it is controlled, and then explore how these factors influence the biology of myrmecophiles, the intimate symbiotic associates of ants. I then test the mechanisms of host specificity by investigating the chemosensory basis of symbiotic interactions between a myrmecophile rove beetle and its single, natural host ant species. I show that host cues trigger analogous behaviors in both the ant and myrmecophile. Cuticular hydrocarbons\u2014the ant's nestmate recognition pheromones\u2014elicit partner recognition in the myrmecophile and execution of ant grooming behavior that achieves chemical mimicry. The myrmecophile also follows host trail pheromones, permitting inter-colony dispersal. Remarkably, however, the myrmecophile performs these same adaptive behaviors with non-host ants separated by up to ~100-million years and shows minimal preference for its natural host over non-host ant species. Experimentally validated agent-based modelling supports a scenario in which specificity is enforced by physiological constraints on dispersal, and negative fitness interactions with alternative hosts, rather than via sensory tuning. Infrequent realization of latent compatibilities of specialists with alternative hosts may facilitate host switching, and the persistence and diversification of seemingly specialized clades over deep time.",
        "doi": "10.7907/k4b2-2n58",
        "publication_date": "2024",
        "thesis_type": "phd",
        "thesis_year": "2024"
    },
    {
        "id": "thesis:16062",
        "collection": "thesis",
        "collection_id": "16062",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06022023-194724728",
        "type": "thesis",
        "title": "Stochastic Foundations for Single-Cell RNA Sequencing",
        "author": [
            {
                "family_name": "Gorin",
                "given_name": "Gennady",
                "orcid": "0000-0001-6097-2029",
                "clpid": "Gorin-Gennady"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Shapiro",
                "given_name": "Mikhail G.",
                "orcid": "0000-0002-0291-4215",
                "clpid": "Shapiro-M-G"
            },
            {
                "family_name": "Wang",
                "given_name": "Zhen-Gang",
                "orcid": "0000-0002-3361-6114",
                "clpid": "Wang-Zhen-Gang"
            },
            {
                "family_name": "Chong",
                "given_name": "Shasha",
                "orcid": "0000-0002-5372-311X",
                "clpid": "Chong-Shasha"
            },
            {
                "family_name": "Ismagilov",
                "given_name": "Rustem F.",
                "orcid": "0000-0002-3680-4399",
                "clpid": "Ismagilov-R-F"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_chem"
            }
        ],
        "abstract": "<p>Single-cell RNA sequencing, which quantifies cell transcriptomes, has seen widespread adoption, accompanied by proliferation of analysis methods. However, there has been relatively little systematic investigation of its best practices and their underlying assumptions, leading to challenges and discrepancies in interpretation. I present a set of generic, principled strategies for modeling the biological and technical components of sequencing experiments and use case studies to motivate their application to sequencing data.</p>",
        "doi": "10.7907/jn6n-x368",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:15235",
        "collection": "thesis",
        "collection_id": "15235",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05302023-215054202",
        "type": "thesis",
        "title": "Diversity in Notch Ligand-Receptor Signaling Interactions",
        "author": [
            {
                "family_name": "Kuintzle",
                "given_name": "Rachael Christine",
                "orcid": "0000-0002-1035-4983",
                "clpid": "Kuintzle-Rachael-Christine"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Elowitz",
                "given_name": "Michael B.",
                "orcid": "0000-0002-1221-0967",
                "clpid": "Elowitz-M-B"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Bronner",
                "given_name": "Marianne E.",
                "orcid": "0000-0003-4274-1862",
                "clpid": "Bronner-M-E"
            },
            {
                "family_name": "Hay",
                "given_name": "Bruce A.",
                "orcid": "0000-0002-5486-0482",
                "clpid": "Hay-B-A"
            },
            {
                "family_name": "Elowitz",
                "given_name": "Michael B.",
                "orcid": "0000-0002-1221-0967",
                "clpid": "Elowitz-M-B"
            }
        ],
        "local_group": [
            {
                "literal": "div_chem"
            }
        ],
        "abstract": "The ability to understand and predict signaling between different cell types is a major challenge in biology. The Notch pathway enables direct signaling through membrane-bound ligands and receptors, and is used in diverse contexts. While its canonical molecular signaling mechanism is well characterized, its many-to-many interacting pathway components, the complexity of their expression patterns, and the presence of same-cell (cis) as well as inter-cellular (trans) receptor-ligand interactions, have made it difficult to predict how a given cell will signal to others. Here, we use a cell-based approach, with Chinese hamster ovary (CHO-K1) cells and C2C12 mouse myoblasts, to systematically characterize trans-activation, cis-inhibition, and cis-activation efficiencies for the essential receptors (Notch1 and Notch2) and activating ligands (Dll1, Dll4, Jag1, and Jag2), in the presence of Lunatic Fringe (Lfng) or the enzymatically dead Lfng D289E mutant. All ligands trans-activate Notch1 and Notch2, except for Jag1, which competitively inhibits Notch1 signaling, and whose Notch1 binding strength is potentiated by Lfng. For Notch1, cis-activation is generally weaker than trans-activation, but for Notch2, cis-activation by Delta ligands is much stronger than trans-activation, and Notch2 cis-activation by Jag1 is similar in strength to trans-activation. Cis-inhibition is associated with weak cis-activation, as Dll1 and Dll4 do not cis-inhibit Notch2. Lfng expression potentiates trans-activation of both Notch1 and Notch2 by the Delta ligands and weakens trans-activation of both receptors by the Jagged ligands. The map of receptor-ligand-Fringe interaction outcomes revealed here should help guide rational perturbation and control of the Notch pathway.",
        "doi": "10.7907/w8gj-jb92",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:15247",
        "collection": "thesis",
        "collection_id": "15247",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05312023-213322223",
        "primary_object_url": {
            "basename": "moses_lambda_2023.pdf",
            "content": "final",
            "filesize": 48841902,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/15247/1/moses_lambda_2023.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Computation Foundations of Spatial Transcriptomics",
        "author": [
            {
                "family_name": "Moses",
                "given_name": "Lambda",
                "orcid": "0000-0002-7092-9427",
                "clpid": "Moses-Lambda"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Van Valen",
                "given_name": "David A.",
                "orcid": "0000-0001-7534-7621",
                "clpid": "Van-Valen-D"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            },
            {
                "family_name": "Pimentel",
                "given_name": "Harold",
                "orcid": "0000-0001-8556-2499",
                "clpid": "Pimentel-Harold"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Single-cell and spatial transcriptomics have come of age in the past few years; datasets and data analysis software packages have proliferated. With the increasing sizes of datasets, proliferating new data collection technologies, and mainstreaming of high-throughput technologies, the software can be improved for better speed and memory efficiency, standardized and consistent user interface for multiple technologies, and in documentation to onboard new users. First, I collected a database of spatial transcriptomics literature and analyzed the data on trends and sociology in this field. Based on the database and data analyses, I wrote a comprehensive book both qualitatively and quantitatively documenting the history of the field since the 1960s and reviewing more recent developments, which informed the software and methods I later developed. Then, to address the challenges with the pre-processing large datasets, we developed \\texttt{kallisto} \\texttt{bustools}  for fast and modular pseudoalignment of sequencing reads to the transcriptome in single-cell RNA-seq (scRNA-seq), giving consistent results with the established and much more computationally demanding alignment method Cell Ranger. Briefly summarized are my attempt to map dissociated cells in scRNA-seq to a spatial gene expression reference and to build a image processing pipeline for image based spatial transcriptomics data analysis. Finally, to address the challenges in downstream analyses of spatial -omics data, I first wrote the new \\texttt{SpatialFeatureExperiment} (SFE) data structure to represent and operate on geometries in spatial transcriptomics data and to organize results from spatial analyses. Based on SFE, I wrote Voyager, which brings decades of research in geospatial data analysis to spatial transcriptomics, to better utilize the opportunities from spatial information to gain novel biological insights. To reduce user learning curve, Voyager conforms to SCE styles and conventions and has a comprehensive documentation website and consistent user interface to many geospatial methods.</p>",
        "doi": "10.7907/rt24-pq60",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:15105",
        "collection": "thesis",
        "collection_id": "15105",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:02122023-103759689",
        "primary_object_url": {
            "basename": "Ph_D__thesis.pdf",
            "content": "final",
            "filesize": 11301290,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/15105/3/Ph_D__thesis.pdf",
            "version": "v8.0.0"
        },
        "type": "thesis",
        "title": "Graph Modeling for Genomics and Epidemiology",
        "author": [
            {
                "family_name": "Eldjarn Hjoerleifsson",
                "given_name": "Kristjan",
                "orcid": "0000-0002-7851-1818",
                "clpid": "Eldjarn-Hjoerleifsson-Kristjan"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            },
            {
                "family_name": "Wierman",
                "given_name": "Adam C.",
                "orcid": "0000-0002-5923-0199",
                "clpid": "Wierman-A-C"
            },
            {
                "family_name": "Melsted",
                "given_name": "Pall",
                "orcid": "0000-0002-8418-6724",
                "clpid": "Melsted-Pall"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_eng"
            }
        ],
        "abstract": "The last decades have seen great leaps made in the development of RNA sequencing technologies, yielding lower cost and greater throughput of experiments, to the point where the scale of the data produced on a daily basis is staggering. While computational hardware is also continuously improving, famously (or perhaps infamously) described by Gordon Moore (Moore, 1965), the rate at which data are produced eclipses advances on the hardware front. Over the last few years, many new methods have been proposed for bridging that ever-widening chasm, more than a few of which harness the latent graphical structure of genomic data to reduce the number of calculations required and pack the data tighter in memory. This body of work continues this development on three different, but related, fronts. Firstly, I present developments that greatly improve upon the efficiency of state-of-the-art methods for the quantification of RNA-seq reads, and describe a method that improves the accuracy of quantification without substantially increasing the computational over- head. Secondly, I introduce a procedure for the discovery of associations between novel gene isoforms and phenotypes, without prior knowledge of those isoforms. Lastly, I present the largest reconstruction of the transmission tree of a viral outbreak to date, modeled from viral genome sequences, contact tracing, and symptom data. I then use the reconstructed transmission tree to assess the efficacy of different vaccination strategies.",
        "doi": "10.7907/s32c-a211",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:16081",
        "collection": "thesis",
        "collection_id": "16081",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06042023-195408313",
        "primary_object_url": {
            "basename": "galvezmerchan_angel_2023_thesis.pdf",
            "content": "final",
            "filesize": 15283976,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/16081/1/galvezmerchan_angel_2023_thesis.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Studies of mRNA Expression and Degradation",
        "author": [
            {
                "family_name": "G\u00e1lvez Merch\u00e1n",
                "given_name": "\u00c1ngel",
                "orcid": "0000-0001-7420-8697",
                "clpid": "G\u00e1lvez-Merch\u00e1n-\u00c1ngel"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Voorhees",
                "given_name": "Rebecca M.",
                "orcid": "0000-0003-1640-2293",
                "clpid": "Voorhees-R-M"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Aravin",
                "given_name": "Alexei A.",
                "orcid": "0000-0002-6956-8257",
                "clpid": "Aravin-A-A"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Voorhees",
                "given_name": "Rebecca M.",
                "orcid": "0000-0003-1640-2293",
                "clpid": "Voorhees-R-M"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Part 1: Protein degradation coupled to Nonsense-mediated mRNA decay</p>\r\n\r\n<p>Translation of mRNAs containing premature termination codons (PTCs) results in truncated protein products with deleterious effects. Nonsense-mediated decay (NMD) is a surveillance pathway responsible for detecting PTC containing transcripts. While the molecular mechanisms governing mRNA degradation have been extensively studied, the fate of the nascent protein product remains largely uncharacterized. In part 1 of this thesis, we use a fluorescent reporter system in mammalian cells to reveal a selective degradation pathway specifically targeting the protein product of an NMD mRNA. We show that this process is post-translational, and dependent on the ubiquitin proteasome system. To systematically uncover factors involved in NMD-linked protein quality control, we conducted genome-wide flow cytometry-based screens. Our screens recovered known NMD factors, but suggested protein degradation did not depend on the canonical ribosome-quality control (RQC) pathway. A subsequent arrayed screen demonstrated that protein and mRNA branches of NMD rely on a shared recognition event. Our results establish the existence of a targeted pathway for nascent protein degradation from PTC containing mRNAs, and provides a reference for the field to identify and characterize required factors.</p>\r\n\r\n<p>Part 2: The Commons Cell Atlas</p>\r\n\r\n<p>Current cell atlas projects aim to curate representative datasets, cell-types, and marker genes for tissues across an organism. Despite their ubiquity, atlas projects rely on duplicated and manual effort to curate marker genes and annotate cell-types. Importantly, the lack of data-compatible tools and a fixed representation of the atlas make their reanalysis near-impossible. To overcome these challenges, we present a collection of data, algorithms, and tools to automate cataloging and analyzing cell-types across all tissues in an organism. We leveraged this work to build a Human Commons Cell Atlas comprising 2.9 million cells across 27 tissues that can be easily updated and that is structured to facilitate custom analyses. To showcase the flexibility of the atlas, we demonstrate that it can be used for isoform analyses. In particular, we study cell-type specificity of isoforms of OAS1, which has recently been shown to offer SARS-CoV-2 protection in certain individuals that display higher expression of the p46 isoform. Using our Commons Cell Atlas, we localize the OAS1 p44b isoform to the testis, and find that it is specific to germ line cells. By virtue of enabling customized analyses via a modular and dynamic atlas structure, the Commons Cell Atlas should be useful for exploratory analyses that are intractable within the rigid framework of current gene-centric static atlases.</p>",
        "doi": "10.7907/esxk-ch24",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:15040",
        "collection": "thesis",
        "collection_id": "15040",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:10122022-003401041",
        "primary_object_url": {
            "basename": "ZS_T_thesis_final_revised.pdf",
            "content": "final",
            "filesize": 44886817,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/15040/1/ZS_T_thesis_final_revised.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Resilience of a Precise Motor Behavior",
        "author": [
            {
                "family_name": "Torok",
                "given_name": "Zsofia Erzsebet",
                "orcid": "ORCID: 0000-0001-6298-9940",
                "clpid": "Torok-Zsofia-Erzsebet"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Lois",
                "given_name": "Carlos",
                "orcid": "0000-0002-7305-2317",
                "clpid": "Lois-Carlos"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Bronner",
                "given_name": "Marianne E.",
                "orcid": "0000-0003-4274-1862",
                "clpid": "Bronner-M-E"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Siapas",
                "given_name": "Athanassios G.",
                "orcid": "0000-0001-8837-678X",
                "clpid": "Siapas-A-G"
            },
            {
                "family_name": "Lois",
                "given_name": "Carlos",
                "orcid": "0000-0002-7305-2317",
                "clpid": "Lois-Carlos"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "Motor memory retention is an essential part of survival and reproduction of most species. However, these behaviors are variable and hard to measure. The zebra finch provides a great model organism to study motor behavior on a fine scale and ask fundamentally important questions. Zebra finch males learn their song from their father and once learnt this song remains unchanged for the remainder of the animals\u2019 life. This highly stereotypic and precise motor function engages a handful of motor nuclei organized in a spatially spread out manner that allows for precise targeting of each key circuit participant for the production of the behavior. In my studies, I focus on better understanding the role of excitatory and inhibitory neurons in the pre-motor nucleus of the song production system. The goal was to perturb the precision of behavioral execution by collapsing the neuronal circuit responsible for sequential activity. Then, to study if the behavior could re-establish in an adult less plastic state of neuronal organization. After I have shown that motor function recovers to produce the same song post disruption, I investigated the large and small scale changes in neuronal activity and transcriptomics accompanying this degradation and recovery trajectory. I have learned that loss of inhibition leads to hyperactivation which eventually leads to a circuit level homeostatic compensation to shut down the pathological activity level. In addition, the upregulation of MHC1 receptors and microglia points to a homeostatic mechanism for synaptic reorganization and re-establishment. Now that we have the means to execute precise cell-type specific manipulations that are reversible and that we understand the underlying phenomenology of perturbation and recovery, we can ask many questions about the architecture of a highly resilient motor pathway. This could shine light on specific electrophysiological and molecular candidates to study for brain damage repair and neurodegenerative research.",
        "doi": "10.7907/2ff5-e145",
        "publication_date": "2023",
        "thesis_type": "phd",
        "thesis_year": "2023"
    },
    {
        "id": "thesis:14338",
        "collection": "thesis",
        "collection_id": "14338",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:08242021-212609828",
        "primary_object_url": {
            "basename": "thesis.pdf",
            "content": "final",
            "filesize": 36577433,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/14338/1/thesis.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Physical Biology of Cellular Information Processing",
        "author": [
            {
                "family_name": "Razo-Mejia",
                "given_name": "Manuel",
                "orcid": "0000-0002-9510-0527",
                "clpid": "Razo-Mejia-Manuel"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Newman",
                "given_name": "Dianne K.",
                "orcid": "0000-0003-1647-1918",
                "clpid": "Newman-D-K"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Goentoro",
                "given_name": "Lea A.",
                "orcid": "0000-0002-3904-0195",
                "clpid": "Goentoro-L-A"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            }
        ],
        "local_group": [
            {
                "literal": "div_chem"
            }
        ],
        "abstract": "<p>The state of matter that we define as <em>life</em> is different from anything else we have encountered so far in the universe. Living systems not only perpetuate their existence out of equilibrium against the will of the second law of thermodynamics, but they do so while keeping up with an ever-changing environment. A key part of this capacity to adapt to environmental changes is the ability of organisms to gather information from their surroundings to put together an adequate response to the challenges presented to them. This thesis presents an effort to understand, from first principles, this fundamental feature of information gathering that all life on earth shares. We dig into the physics behind one of the most pervasive mechanisms through which living systems sense and respond to the environment\u2013the ability to turn <em>on</em> and <em>off</em> genes. In doing so, we hope to uncover general principles of how organisms deal with the problem of collecting information about the world that surrounds them.</p>\r\n\r\n<p>In Chapter 1, we develop the theoretical and conceptual tools to navigate the rest of the thesis. I introduce the idea of gene regulation, as well as different theoretical models of this pervasive biological phenomenon. We also delve into the realm of information theory and learn how the plastic concept of information can be mathematically defined and quantified.</p>\r\n\r\n<p>The second stop in our exploration (Chapter 2) asks the following question: can we understand, from first principles, how it is that proteins allow cells to regulate their genes on-demand upon sensing environmental cues? For this, we explore the physics behind transcriptional control due to allosteric transcription factors. Using simple quasi-equilibrium models of the two processes involved in this type of regulation\u2014the regulation of the gene by the binding and unbinding of the transcription factor, and the regulation of the activity of the transcription factor itself by the binding and unbinding of an effector molecule\u2014we are able to predict the input-output function of a simple genetic circuit, and compare such predictions with experimental determinations of the mean response of a population of bacterial cells.</p>\r\n\r\n<p>We then expand on these insights to ask questions about the inescapable cell-to-cell variability that isogenic cells encounter. For this, we have to leave behind the pure thermodynamic framework and work in the language of chemical kinetics. This allows us to make predictions beyond the mean input-output gene expression response of cells by reconstructing full gene expression distributions. With these probabilistic input-output functions, in Chapter 3 we formalize the question of the <em>amount of information</em> that cells can gather from the environment. For this, we turn to information-theoretic concepts of maximal mutual information (otherwise known as channel capacity) between the state of the environment and the gene expression response from bacterial cells. Finally, we compare our predictions of the maximum amount of information\u2014measured in bits\u2014that cells can gather with single-cell inferences of this quantity.</p>",
        "doi": "10.7907/kpc2-b345",
        "publication_date": "2022",
        "thesis_type": "phd",
        "thesis_year": "2022"
    },
    {
        "id": "thesis:14652",
        "collection": "thesis",
        "collection_id": "14652",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05302022-235732314",
        "primary_object_url": {
            "basename": "2022_XFZ_PhD_Thesis_final.pdf",
            "content": "final",
            "filesize": 24249102,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/14652/1/2022_XFZ_PhD_Thesis_final.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Biocontrol of Biomolecular Systems: Polyhedral Constraints on Binding's Regulation of Catalysis from Biocircuits to Metabolism",
        "author": [
            {
                "family_name": "Xiao",
                "given_name": "Fangzhou",
                "orcid": "0000-0002-5001-5644",
                "clpid": "Xiao-Fangzhou"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Doyle",
                "given_name": "John Comstock",
                "orcid": "0000-0002-1828-2486",
                "clpid": "Doyle-J-C"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Murray",
                "given_name": "Richard M.",
                "orcid": "0000-0002-5785-7481",
                "clpid": "Murray-R-M"
            },
            {
                "family_name": "Winfree",
                "given_name": "Erik",
                "orcid": "0000-0002-5899-7523",
                "clpid": "Winfree-E"
            },
            {
                "family_name": "Phillips",
                "given_name": "Robert B.",
                "orcid": "0000-0003-3082-2809",
                "clpid": "Phillips-R"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Doyle",
                "given_name": "John Comstock",
                "orcid": "0000-0002-1828-2486",
                "clpid": "Doyle-J-C"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>One eventual goal of bioengineering is to build complex biological machines that fully realize the unique potential of biotechnology, namely adaptation, survival, growth, and dominance. In order to do so, not only do we need theoretical understanding and reliable manufacturing of biological parts and components, we also need a systems theory that captures fundamental structures to obtain insight about the space of all possible behaviors when parts are put together. This enables us to understand what can and cannot be achieved. Examples from other engineering disciplines are Turing machines for computers, information channels for communication networks, linear input output systems for electrical circuits, and thermodynamics for heat engines. This work is an attempt at developing a systems theory tailored to biomolecular systems in cells. The results form the following statements.</p>\r\n\r\n<p>Biomolecular systems are binding and catalysis reactions. Catalysis determines the direction of change, while binding regulates how the catalysis rates vary with reactant concentrations. Given a binding reaction network, the full range of regulatory profiles can be captured by the reaction orders of catalysis, which in turn is constrained in polyhedral sets determined by the stoichiometry of binding. This constitute a rule, that since cells control catalysis by binding, cells control catalysis rates by regulating reaction orders constrained in polyhedral sets. This rule has ramifications in several directions. On metabolism, by incorporating the constraint that reaction orders of metabolic fluxes, not the fluxes themselves, are controlled, we can predict metabolism dynamics directly from network stoichiometry, e.g. glycolytic oscillations and growth arrests. This is a fully dynamic upgrade of flux balance analysis, a popular constraint-based method to model metabolism. On systems biology, this rule derives a method of biocircuit analysis based on the full range of values that reaction orders can take. This allows discovery of necessary and sufficient conditions for a circuit to achieve a certain function, thus revealing regimes hidden by traditional methods of analysis. It also promotes holistic comparisons of different circuit implementations, e.g. activating versus repressing, thur enabling biocircuit design where we know when a design will work, and when a design will fail. On dynamics and control of biocircuits, reaction order can work as a robust basis for stability, perfect adaptation, multistability, and oscillations. Lyapunov functions and dissipative control theory tailored for biomolecular systems are constructed based on reaction orders. On the mathematics of biology, it relates bioregulation to convex polyhedra, log derivative operator decompositions, and fundamental rules of calculus for positive variables.</p>",
        "doi": "10.7907/rtwq-v497",
        "publication_date": "2022",
        "thesis_type": "phd",
        "thesis_year": "2022"
    },
    {
        "id": "thesis:14649",
        "collection": "thesis",
        "collection_id": "14649",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05292022-204424650",
        "type": "thesis",
        "title": "Foundations and Applications of Single-Cell RNA Sequencing",
        "author": [
            {
                "family_name": "Booeshaghi",
                "given_name": "Ali Sina",
                "orcid": "0000-0002-6442-4502",
                "clpid": "Booeshaghi-Ali-Sina"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Greer",
                "given_name": "Julia R.",
                "orcid": "0000-0002-9675-1508",
                "clpid": "Greer-J-R"
            },
            {
                "family_name": "Colonius",
                "given_name": "Tim",
                "orcid": "0000-0003-0326-3909",
                "clpid": "Colonius-T"
            },
            {
                "family_name": "Melsted",
                "given_name": "P\u00e1ll",
                "orcid": "0000-0002-8418-6724",
                "clpid": "Melsted-P\u00e1ll"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "local_group": [
            {
                "literal": "div_eng"
            }
        ],
        "abstract": "<p>Single-cell RNA-sequencing is an experimental technique for studying cellular gene expression, with a multitude of engineering challenges. These challenges transcend the boundaries of traditional academic disciplines and the field of mechanical engineering, that aims to address roadblocks in critical technologies towards engineering our environment, is central to this endeavor.</p> \r\n\r\n<p>This thesis addresses three engineering challenges that must be met in order to realize the goal of bringing single-cell RNA sequencing to the clinic. The first is scalable cellular isolation and sampling. Chapter 2 describes the <i>poseidon</i> and <i>colosseum</i> instruments that enable massive scale single-cell isolation and collection. They each have novel design elements that reduce cost and enable modularity, at a similar accuracy to expensive commercial alternatives.</p>\r\n\r\n<p>The second challenge is the rapid preprocessing of single-cell RNA-sequencing data. Chapter 3 describes the <i>kallisto</i> | <i>bustools</i> command-line tools that make scalable scRNAseq analysis fast and efficient. These tools implement novel algorithms for sequence read-alignment, barcode error correction, and molecular counting that helps resolve ambiguities in sequence mapping.</p> \r\n\r\n<p>The third challenge is refining gene expression data to the isoform level. This refinement is crucial for understanding transcriptional regulation and the effects of alternative splicing in biological processes. Towards that end, I have extended the <i>kallisto</i> | <i>bustools</i> workflow to process full-length scRNAseq data taking advantage of expectation maximization algorithm to disambiguate sequence alignments. Chapter four describes how I used these tools to assemble the first ever spatially-resolved single-cell isoform atlas, and in particular one of great interest in the neuroscience community (the mouse primary motor cortex) with data generated with three RNA-sequencing assays.</p>",
        "doi": "10.7907/ptbp-a779",
        "publication_date": "2022",
        "thesis_type": "phd",
        "thesis_year": "2022"
    },
    {
        "id": "thesis:14631",
        "collection": "thesis",
        "collection_id": "14631",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:05262022-234214451",
        "primary_object_url": {
            "basename": "WittmannBruce_Thesis.pdf",
            "content": "final",
            "filesize": 10219666,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/14631/5/WittmannBruce_Thesis.pdf",
            "version": "v5.0.0"
        },
        "type": "thesis",
        "title": "Strategies and Tools for Machine Learning-Assisted Protein Engineering",
        "author": [
            {
                "family_name": "Wittmann",
                "given_name": "Bruce James",
                "orcid": "0000-0001-8144-9157",
                "clpid": "Wittmann-Bruce-James"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Arnold",
                "given_name": "Frances Hamilton",
                "orcid": "0000-0002-4027-364X",
                "clpid": "Arnold-F-H"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Reisman",
                "given_name": "Sarah E.",
                "orcid": "0000-0001-8244-9300",
                "clpid": "Reisman-S-E"
            },
            {
                "family_name": "Mayo",
                "given_name": "Stephen L.",
                "orcid": "0000-0002-9785-5018",
                "clpid": "Mayo-S-L"
            },
            {
                "family_name": "Arnold",
                "given_name": "Frances Hamilton",
                "orcid": "0000-0002-4027-364X",
                "clpid": "Arnold-F-H"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "Proteins perform critical roles in a growing list of human-devised applications, and as demands for new applications arise, new proteins must be engineered to meet them. Machine learning-assisted protein engineering (MLPE) has recently arisen as a new philosophy of protein engineering, promising to overcome many of the limitations of existing engineering strategies. Despite its promise, however, as a relatively new approach to protein engineering, MLPE faces many challenges that hinder its routine application. This thesis is focused on addressing a number of them. Chapter 1 provides a theoretical overview of protein engineering, introduces the core steps of a typical MLPE pipeline, and discusses the challenges that currently hinder MLPE\u2019s advancement. This chapter is written to be accessible to all members of the highly multidisciplinary audience that either use or develop MLPE tools, in turn providing a resource that eliminates the steep barrier to entry that can hinder broader participation in the field. Chapter 2 provides a solution to the challenge of applying MLPE to proteins whose fitness landscapes are dominated by \u201choles\u201d (protein variants with zero or extremely low fitness). Using my development of the strategy \u201cfocused training machine learning-assisted directed evolution (ftMLDE)\u201d as an example, I demonstrate how auxiliary information from protein sequence and structure can be used to navigate landscapes despite holes, in turn dramatically improving the efficiency of MLPE. Chapter 3 explores strategies for reducing the amount of sequence-fitness data needed for building MLPE models. Specifically, I detail the motivation behind and development of a new model designed to augment limited protein sequence-fitness datasets with information extracted from raw protein sequence and structure data. Finally, chapter 4 introduces \u201cevery variant sequencing\u201d (evSeq), a collection of tools and protocols that enables extremely low-cost, routine collection of large protein sequence-fitness datasets. Not only does this technology drastically improve the financial feasibility of numerous MLPE applications, but it also potentiates the construction of a massive database of diverse protein sequence-fitness data, the likes of which would revolutionize our ability to engineer proteins with data-driven methods. Overall, the work described in this thesis advances both our understanding of MLPE and our ability to engineer proteins using it.",
        "doi": "10.7907/azzt-0q97",
        "publication_date": "2022",
        "thesis_type": "phd",
        "thesis_year": "2022"
    },
    {
        "id": "thesis:14938",
        "collection": "thesis",
        "collection_id": "14938",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06022022-201232129",
        "primary_object_url": {
            "basename": "Final Thesis version Eduardo da Veiga Beltrame.pdf",
            "content": "final",
            "filesize": 17501518,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/14938/1/Final Thesis version Eduardo da Veiga Beltrame.pdf",
            "version": "v4.0.0"
        },
        "type": "thesis",
        "title": "Stories in Single Cell RNA Sequencing",
        "author": [
            {
                "family_name": "da Veiga Beltrame",
                "given_name": "Eduardo",
                "orcid": "0000-0002-1529-9207",
                "clpid": "da-Veiga-Beltrame-Eduardo"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Sternberg",
                "given_name": "Paul W.",
                "orcid": "0000-0002-7699-0173",
                "clpid": "Sternberg-P-W"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Van Valen",
                "given_name": "David A.",
                "orcid": "0000-0001-7534-7621",
                "clpid": "Van-Valen-D"
            },
            {
                "family_name": "Sternberg",
                "given_name": "Paul W.",
                "orcid": "0000-0002-7699-0173",
                "clpid": "Sternberg-P-W"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>This thesis describes the projects I have worked on since starting the Caltech bioengineering program in fall 2017. The general theme of my projects is that they are all about single cell RNA sequencing (scRNA-seq), spanning the experimental and computational realms.</p> \r\n\r\n<p>Chapter 1 is an introduction explaining the essential concepts and is meant to be readable by a wide audience. For the other chapters, each one describes a separate project in a succinct manner, including links to the related preprint, published paper or code repositories at the start of each chapter.</p>\r\n\r\n<p>Chapter 2 describes the scVI generative model for scRNA-seq data and the scvi-tools framework, which forms the basis of many of my computational projects.</p> \r\n\r\n<p>Chapter 3 describes an open source 3D printable syringe pump system that was developed envisioning facilitating many kinds of experiments, in particular droplet based scRNA-seq.</p> \r\n\r\n<p>Chapter 4 describes a new way of fabricating hydrogel beads with unique DNA barcodes that are used for scRNA-seq experiments.</p> \r\n\r\n<p>Chapter 5 describes a database listing most published scRNA-seq studies that I helped create, and provides a useful overview of the state of the field.</p> \r\n\r\n<p>Chapter 6 describes the kallisto bus workflow, which is used for pre-processing scRNA-seq data, going from FASTQ file to gene count matrix in a very efficient manner.</p> \r\n\r\n<p>Chapter 7 describes a new way of using scVI to quantify the trade- off in the quality of scRNA-seq of a given dataset when surveying more cells or sequencing more reads per cell.</p> \r\n\r\n<p>Chapter 8 describes tools developed for the WormBase users to leverage scRNA-seq data on <i>C. elegans</i>, and which can be deployed with any other scRNA-seq dataset.</p> \r\n\r\n<p>Chapter 9 describes a remarkably successful offshoot of the devel- opment of these tools: a simple scVI based analysis and visualization strategy for finding candidate marker genes using <i>C. elegans</i> scRNA-seq data, which was experimentally validated by members of the Sternberg lab.</p>",
        "doi": "10.7907/4kgh-8420",
        "publication_date": "2022",
        "thesis_type": "phd",
        "thesis_year": "2022"
    },
    {
        "id": "thesis:11799",
        "collection": "thesis",
        "collection_id": "11799",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:09222019-132051506",
        "primary_object_url": {
            "basename": "Taeb_thesis.pdf",
            "content": "final",
            "filesize": 10371630,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/11799/1/Taeb_thesis.pdf",
            "version": "v5.0.0"
        },
        "type": "thesis",
        "title": "Latent-Variable Modeling: Algorithms, Inference, and Applications",
        "author": [
            {
                "family_name": "Taeb",
                "given_name": "Armeen",
                "orcid": "0000-0002-5647-3160",
                "clpid": "Taeb-Armeen"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Chandrasekaran",
                "given_name": "Venkat",
                "clpid": "Chandrasekaran-V"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Hassibi",
                "given_name": "Babak",
                "orcid": "0000-0002-1375-5838",
                "clpid": "Hassibi-B"
            },
            {
                "family_name": "Stuart",
                "given_name": "Andrew M.",
                "orcid": "0000-0001-9091-7266",
                "clpid": "Stuart-A-M"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Doyle",
                "given_name": "John Comstock",
                "orcid": "0000-0002-1828-2486",
                "clpid": "Doyle-J-C"
            },
            {
                "family_name": "Chandrasekaran",
                "given_name": "Venkat",
                "clpid": "Chandrasekaran-V"
            }
        ],
        "local_group": [
            {
                "literal": "Resnick Sustainability Institute"
            },
            {
                "literal": "div_eng"
            }
        ],
        "abstract": "<p>Many driving factors of physical systems are often latent or unobserved. Thus, understanding such systems crucially relies on accounting for the influence of the latent structure. This thesis makes advances in three aspects of latent-variable modeling: inference, algorithms, and applications. Specifically, we develop and explore latent-variable techniques that a) ensure interpretable and statistically significant models, b) can be efficiently optimized to identify best fit to data, and c) provide useful insights in real-world applications. The specific contributions of this thesis are:</p>\r\n\r\n<p>1. We employ a latent-variable graphical modeling technique to develop the first state-wide statistical model of the California reservoir network. With this model, we precisely characterize the system-wide behavior of the network to hypothetical drought conditions, and proposed guidelines for more sustainable reservoir management.</p>\r\n\r\n<p>2. Motivated by the previous application, we provide a geometric framework to assess the extent to which our latent variable model has learned true or false discoveries about the relevant physical phenomena. Our approach generalizes the classical notions of true and false discoveries in mathematical statistics that rely on the discrete structure of the decision space to settings where the decision space is continuous and more complicated. We highlight the utility of this viewpoint in problems involving subspace selection and low-rank estimation.</p>\r\n\r\n<p>3. We propose a convex optimization procedure to fit a latent-variable graphical model for generalized linear models. This framework provides a flexible approach to model non-Gaussian variables including Poisson, Bernoulli, and exponential variables. A particularly novel aspect of our formulation is that it incorporates regularizers that are tailored to the type of latent variables.</p>\r\n\r\n<p>4. We describe a computationally efficient framework to learn a latent-variable model with high-dimensional and non-iid data. This framework is based on factoriable precision operators that decouple the component associated with the observational dependencies and the component associated to interdependencies among the variables.</p>\r\n\r\n<p>5. We propose a convex optimization technique to provide semantics to latent variables of a factor model. This approach is based on linking auxiliary variables -- chosen based on domain expertise -- to these latent variables.</p>",
        "doi": "10.7907/YRF1-7W29",
        "publication_date": "2020",
        "thesis_type": "phd",
        "thesis_year": "2020"
    },
    {
        "id": "thesis:13815",
        "collection": "thesis",
        "collection_id": "13815",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:06112020-001257003",
        "primary_object_url": {
            "basename": "JohnThompson_Thesis_Final.pdf",
            "content": "final",
            "filesize": 14848581,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/13815/1/JohnThompson_Thesis_Final.pdf",
            "version": "v6.0.0"
        },
        "type": "thesis",
        "title": "Chemical Tools for Studying O-GlcNAc Glycosylation at the Systems Level",
        "author": [
            {
                "family_name": "Thompson",
                "given_name": "John Warren Lenzi",
                "orcid": "0000-0003-0061-4996",
                "clpid": "Thompson-John-Warren-Lenzi"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Hsieh-Wilson",
                "given_name": "Linda C.",
                "orcid": "0000-0001-5661-1714",
                "clpid": "Hsieh-Wilson-L-C"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Hoelz",
                "given_name": "Andre",
                "orcid": "0000-0003-0923-3284",
                "clpid": "Hoelz-A"
            },
            {
                "family_name": "Chan",
                "given_name": "David C.",
                "orcid": "0000-0002-0191-2154",
                "clpid": "Chan-D-C"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Hsieh-Wilson",
                "given_name": "Linda C.",
                "orcid": "0000-0001-5661-1714",
                "clpid": "Hsieh-Wilson-L-C"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>The addition of O-linked \u03b2-N-acetylglucosamine (O-GlcNAc) to intracellular serine and threonine residues is a ubiquitous post-translational modification (PTM) found in all higher eukaryotes. Like other PTMs, it is finely regulated in response to stimuli and dysregulated in multiple diseases. However, unlike other PTMs, methods to detect and profile the dynamics of O-GlcNAc glycosylation are still in their infancy. Herein, we discuss the background, development, and application of new chemical tools that have allowed for some of the first systems-level investigations of O-GlcNAcylation in different cells, organ systems, and disease states. We also significantly advance established techniques for the detection and monitoring of O-GlcNAc on proteins of interest. Using these new techniques, we first uncover a novel O-GlcNAcylation site on Cdk5 and show that this site can dynamically regulate Cdk5 activity in the context of neurodegenerative disease. Next, we apply novel chemical, mass spectrometric, and computational tools to, for the first time, uncover cellular networks engaged by O-GlcNAcylation in vivo. Finally, we undertake the systematic optimization of mass spectrometry based O-GlcNAcomics and use these new insights to significantly advance our understanding of O-GlcNAcylation dynamics in metabolic diseases of the liver. Overall, the techniques developed and data generated herein are closing the methodological and intellectual gaps between the study of O-GlcNAc glycosylation and that of other PTMs.</p>",
        "doi": "10.7907/gx3z-k069",
        "publication_date": "2020",
        "thesis_type": "phd",
        "thesis_year": "2020"
    },
    {
        "id": "thesis:13609",
        "collection": "thesis",
        "collection_id": "13609",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:12162019-183140887",
        "primary_object_url": {
            "basename": "Thesis_Dong-Wook_Kim_v3.pdf",
            "content": "final",
            "filesize": 17443112,
            "license": "other",
            "mime_type": "application/pdf",
            "url": "/13609/1/Thesis_Dong-Wook_Kim_v3.pdf",
            "version": "v5.0.0"
        },
        "type": "thesis",
        "title": "Multimodal Analysis of Cell Types in a Hypothalamic Node Controlling Social Behavior in Mice",
        "author": [
            {
                "family_name": "Kim",
                "given_name": "Dong-Wook",
                "orcid": "0000-0002-5497-5853",
                "clpid": "Kim-Dong-Wook"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Anderson",
                "given_name": "David J.",
                "orcid": "0000-0001-6175-3872",
                "clpid": "Anderson-D-J"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Anderson",
                "given_name": "David J.",
                "orcid": "0000-0001-6175-3872",
                "clpid": "Anderson-D-J"
            },
            {
                "family_name": "Oka",
                "given_name": "Yuki",
                "orcid": "0000-0003-2686-0677",
                "clpid": "Oka-Yuki"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>The advent and recent advances of single-cell RNA sequencing (scRNA-seq) have yielded transformative insights into our understanding of cellular diversity in the central nervous system (CNS) with unprecedented detail. However, due to current experimental and computational limitations on defining transcriptomic cell types (T-types) and the multiple phenotypic features of cell types in the CNS, an integrative and multimodal approach should be required for the comprehensive classification of cell types.</p>\r\n\r\n<p>To this end, performing multimodal analysis of scRNA-seq in hypothalamus would be very beneficial in that hypothalamus, controlling homeostatic and innate survival behaviors which known to be highly conserved across a wide range of species and encoded in hard-wired brain circuits, is likely to display the more straightforward relationship between transcriptomic identity, axonal projections, and behavioral activation, respectively. In my dissertation, I have been focused on the cell type characterizations of a hypothalamic node controlling innate social behavior in mice, the ventrolateral subdivision of the ventromedial hypothalamus (VMHvl). VMHvl only contains ~4,000 neurons per hemisphere in mice but due to its behavioral, anatomical, and molecular heterogeneity, which T-types in VMHvl are related to connectivity and behavioral function is largely unknown.</p>\r\n\r\n<p>In Chapter II, I described my main thesis work to perform scRNA-seq in VMHvl using two independent platforms: SMART-seq2 (~4,500 neurons sequenced) and 10x (~78,000 neurons sequenced). Specifically, 17 joint VMHvl T-types including several sexually dimorphic clusters were identified by canonical correlation analysis (CCA) in Seurat, and the majority of them were validated by multiplexed single-molecule FISH (seqFISH). Correspondence between transcriptomic identity, and axonal projections or behavioral activation, respectively, was also investigated. Immediate early gene analysis identified T-types exhibiting preferential responses to intruder males versus females but only rare examples of behavior-specific activation. Unexpectedly, many VMHvl T-types comprise a mixed population of neurons with different projection target preferences. Overall our analysis revealed that, surprisingly, few VMHvl T-types exhibit a clear correspondence with behavior-specific activation and connectivity.</p>\r\n\r\n<p>In Chapter III, I will discuss about future directions for a deeper and better understanding of VMHvl cell types. Briefly, my previous data from whole-cell patch clamp recording in VMHvl slices suggested that there were at least 4 distinct electrophysiological cell types (E-types). Additionally, two distinct neuromodulatory effects on VMHvl were observed (persistently activated by vasopressin/oxytocin vs. silenced by nitric oxide) by monitoring populational activities using two-photon Ca2+ imaging in slices. Based on the results from the first part and combined with advanced molecular techniques (e.g. Patch-seq and CRISPR-Cas9), we can further dissect out the cellular diversity in VMHvl and their functional implications.</p>",
        "doi": "10.7907/RGVK-9962",
        "publication_date": "2020",
        "thesis_type": "phd",
        "thesis_year": "2020"
    },
    {
        "id": "thesis:11226",
        "collection": "thesis",
        "collection_id": "11226",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:10102018-143313907",
        "type": "thesis",
        "title": "Statistical Methods for Gene Differential Expression Analysis of RNA-Sequencing",
        "author": [
            {
                "family_name": "Yi",
                "given_name": "Lynn Donglin",
                "orcid": "0000-0003-4575-0158",
                "clpid": "Yi-Lynn-Donglin"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Chan",
                "given_name": "David C.",
                "orcid": "0000-0002-0191-2154",
                "clpid": "Chan-D-C"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Chandrasekaran",
                "given_name": "Venkat",
                "clpid": "Chandrasekaran-V"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>RNA-Sequencing (\"RNA-Seq\") is performed to measure gene expression, often to ask the question of what genes are differentially expressed across various biological conditions. Statistical methods have been used to model RNA-Seq quantifications in order to determine differential expression, and have traditionally be divided into gene-level methods and transcript-level methods. There has been little attempt to connect the statistical divide, although transcript expression and gene expression are biologically inextricably linked. In this thesis, we provide a case study of a comparative differential expression analysis, demonstrating that many differential expression events happen on the isoform-level, and that performing an analysis using only summarized gene quantifications would fail to capture these events. Furthermore, we develop statistical methods that unify the transcript-level and gene-level analysis. In bulk RNA-Seq, by using p-value aggregation methods, we are able to translate transcript-level results into gene-level results under a unified framework. For single cell RNA-Seq, we propose using multiple logistic regression, leveraging the high dimensionality of the data in order to determine if the transcript quantifications pertaining to a gene are able to constitute a linear discriminant for cell type. This method combines differential transcript expression analysis and differential gene expression analysis into a unified framework which we call \u201cgene differential expression.\u201d Lastly, we demonstrate that our methods could be used on transcript compatibility counts instead of transcript quantifications in order to bypass ambiguous read assignment and improve accuracy. We show that transcript compatibility counts obtained via transcriptome pseudoalignment are comparable in quantification accuracy to quantifications from genome alignment methods.</p>",
        "doi": "10.7907/0YE6-2217",
        "publication_date": "2019",
        "thesis_type": "phd",
        "thesis_year": "2019"
    },
    {
        "id": "thesis:11396",
        "collection": "thesis",
        "collection_id": "11396",
        "cite_using_url": "https://resolver.caltech.edu/CaltechTHESIS:02192019-004236200",
        "type": "thesis",
        "title": "The Changing Mouse Embryo Transcriptome at Whole Tissue and Single-Cell Resolution",
        "author": [
            {
                "family_name": "He",
                "given_name": "Peng",
                "orcid": "0000-0002-2457-3554",
                "clpid": "He-Peng"
            }
        ],
        "thesis_advisor": [
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            }
        ],
        "thesis_committee": [
            {
                "family_name": "Pachter",
                "given_name": "Lior S.",
                "orcid": "0000-0002-9164-6231",
                "clpid": "Pachter-L"
            },
            {
                "family_name": "Sternberg",
                "given_name": "Paul W.",
                "orcid": "0000-0002-7699-0173",
                "clpid": "Sternberg-P-W"
            },
            {
                "family_name": "Fejes Toth",
                "given_name": "Katalin",
                "orcid": "0000-0001-6558-2636",
                "clpid": "Fejes-Toth-K"
            },
            {
                "family_name": "Thomson",
                "given_name": "Matthew",
                "orcid": "0000-0003-1021-1234",
                "clpid": "Thomson-M-W"
            },
            {
                "family_name": "Guttman",
                "given_name": "Mitchell",
                "orcid": "0000-0003-4748-9352",
                "clpid": "Guttman-M"
            },
            {
                "family_name": "Wold",
                "given_name": "Barbara J.",
                "orcid": "0000-0003-3235-8130",
                "clpid": "Wold-B-J"
            }
        ],
        "local_group": [
            {
                "literal": "div_bbe"
            }
        ],
        "abstract": "<p>Mammalian histogenesis is a sophisticated process of coordinated changes of cellular composition governed by selective gene expression. This thesis focuses on the systematic application of modern RNA-seq methods to histogenesis processes in developing mouse embryos. Most of the work presented here is conducted as part of the ENCODE (ENCyclopedia Of DNA Elements) Project. Chapter 1 introduces the current advances of transcriptome studies on tissue development. Chapter 2 discusses a large-scale study on the whole-tissue transcriptome of 12 embryonic tissues at up to 8 time points and 5 additional perinatal tissues. Coherent themes of biological function and underlying regulatory mechanisms are revealed from the large-scale analysis. Chapter 3 presents a high-resolution single-cell RNAseq study focused on the developing forelimb of the mouse embryo. This approach enables the assignment of differential genes to corresponding lineages and provides an even more accurate picture of RNA level patterns and regulatory modes. Finally, whole-tissue and single-cell methods are compared, contrasted, and integrated in Chapter 4 to extrapolate from the main discoveries of this thesis.</p>",
        "doi": "10.7907/35S4-HG18",
        "publication_date": "2019",
        "thesis_type": "phd",
        "thesis_year": "2019"
    }
]