BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://live-hu-cmsa-222.pantheonsite.io
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260903T090000
DTEND;TZID=America/New_York:20260904T170000
DTSTAMP:20260902T213438Z
CREATED:20260217T174509Z
LAST-MODIFIED:20260902T213438Z
UID:10003846-1788426000-1788541200@live-hu-cmsa-222.pantheonsite.io
SUMMARY:Big Data Conference 2026
DESCRIPTION:Big Data Conference 2026 \nDates: Sep. 3–4\, 2026 \nLocation: Harvard University CMSA\, 20 Garden Street\, Cambridge MA & via Zoom \nThe Big Data Conference features speakers from the Harvard community as well as scholars from across the globe\, with talks focusing on computer science\, statistics\, math and physics\, and economics. \nRegister to attend in person \nRegister for Zoom Webinar \n  \nConfirmed Speakers \n\nSueYeon Chung\, Harvard\nBailey Flanigan\, MIT\nSergey Ovchinnikov\, MIT\nAriel Procaccia\, Harvard\nAdit Radhakrishnan\, MIT\nAndrew Sutherland\, MIT\nChris Wiggins\, Columbia\nRex Ying\, Yale\nWei Zhou\, Harvard\n\n  \nOrganizers \n\nMichael Desai\, Harvard\nMichael R. Douglas\, CMSA\nYannai Gonczarowski\, Harvard\nMelanie Weber\, Harvard\n\n  \nSchedule (download pdf) \nThursday\, Sep. 3\, 2026 \n8:45–9:10 am\nBreakfast \n9:10–9:15 am\nIntroductions \n9:15–10:15 am\nRex Ying\, Yale\nA Riemannian Geometry Perspective on Foundation Models\nAbstract: In development of foundation models and LLMs\, Euclidean space has been the de facto geometric setting for machine learning architectures. However\, at a large scale\, real-world data often exhibit inherently non-Euclidean structures\, such as multi-way relationships\, hierarchies\, symmetries\, and non-isotropic scaling\, in a variety of domains\, such as languages\, vision\, and the natural sciences. It is challenging to effectively capture these structures within the constraints of Euclidean spaces. This talk will cover my research that moves beyond Euclidean geometry to maintain the scaling law for the next generation of foundation models. By adopting these geometries\, foundation models could more efficiently leverage the structures. Task-aware adaptability that dynamically reconfigures embeddings to match the geometry of downstream applications\, could further enhance efficiency and expressivity. This talk will demonstrate the advantages of hyperbolic geometry in building foundation models from individual non-Euclidean components\, LLM pre-training / fine-tuning\, retrieval-augmented generation\, and mixture of expert models. We further present a comprehensive library that greatly simplifies the process of building and adapting foundation models in non-Euclidean. Finally\, the talk will enumerate a few promising directions that are being actively explored in this domain. \n10:15–10:30 am\nBreak \n10:30–11:30 am\nAdit Radhakrishnan\, MIT\nToward universal steering and monitoring of AI models\nAbstract: Artificial intelligence (AI) models contain much of human knowledge. Understanding the representation of this knowledge will lead to improvements in model capabilities and safeguards. Building on advances in feature learning\, we developed an approach for extracting linear representations of semantic notions or concepts in AI models. We showed how these representations enabled model steering\, through which we exposed vulnerabilities and improved model capabilities. We demonstrated that concept representations were transferable across languages and enabled multiconcept steering. Across hundreds of concepts\, we found that larger models were more steerable and that steering improved model capabilities beyond prompting. We showed that concept representations were more effective for monitoring misaligned content than for using judge models. Our results illustrate the power of internal representations for advancing AI safety and model capabilities. \n11:30 am–12:45 pm\nLunch \n12:45–1:45 pm\nAndrew Sutherland\, MIT\nCrowdsourcing mathematical databases\nAbstract: The L-functions and Modular Forms Database (LMFDB) is one of the largest online repositories of mathematical research data. It contains comprehensive catalogs of mathematical objects that arise in the context of the Langlands program\, including number fields\, elliptic curves\, modular forms\, and higher-dimensional analogs of these objects. In cooperation with the Foundation for Science and AI Research (SAIR)\, we recently ran a competition aimed at solving the inverse Galois problem over Q in degree 24 (IGP24). This goal was achieved by collecting more than 50 million candidate number fields submitted by more than 100 teams\, including examples that realize all 25\,000 transitive permutation groups of degree 24 as Galois groups. AI tools and computer algebra systems played a central role in this project\, both in building the infrastructure to run the contest and in enhancing the capabilities of participants. \n1:45–2:00 pm\nBreak \n2:00–3:00 pm\nBailey Flanigan\, MIT\nAlgorithmic Tools for Trading Off Sortition Ideals\nAbstract: Citizens’ assemblies and other deliberative minipublics — representative groups of everyday people convened to deliberate on a policy issue and then make recommendations — are now used by governments around the world. Choosing who sits on these panels is the problem of sortition: randomly selecting a small group of citizens that represents the broader population. Sortition has been the subject of substantial computer science research in recent years\, and the resulting algorithms are now widely used in practice. This talk will describe the key challenges that arise in the practice of sortition\, and the algorithmic tools that have been developed to navigate them optimally. \n3:00–3:15 pm\nBreak \n3:15–4:15 pm\nChris Wiggins\, Columbia\nPrescriptive Learning at Scale: Applied and Deployed Decision-Making\nAbstract: Many applications in health and industry require making decisions while learning how the world responds to them: e.g.\, which article to recommend\, which treatment to assign\, or when to show a paywall. This problem is prescriptive\, what to do rather than what is true. Statistical decision theory has taken up the problem repeatedly since Neyman’s interwar work; recent decades have brought a new wave of results\, motivated in large part by the opportunity to make and evaluate decisions at scale online. I will highlight a thread of 21st-century results running through adaptive experimentation\, contextual bandits\, and off-policy evaluation: they combine advances in supervised learning with fundamental ideas from statistical sampling\, and now shape digital products\, from content recommendation to marketing\, as well as adaptive interventions in health. \n  \nFriday\, Sep. 4\, 2026 \n8:45–9:15 am\nBreakfast \n9:15–10:15 am\nAriel Procaccia\, Harvard\nNo Generation Without Representation\nAbstract: AI systems and democratic processes are confronting similar challenges around representation. I examine two related questions that cut across both domains. First\, how can AI enable democratic processes that handle vast spaces of opinions or statements while ensuring proportional representation of a population’s views? Second\, when AI systems themselves provide normative guidance\, whose viewpoints do they reflect\, and can we make this precise? Drawing on social choice theory\, I present formal frameworks and algorithms for both problems\, showing that meaningful representation guarantees are feasible and practical. \n10:15–10:30 am\nBreak \n10:30–11:30 am\nWei Zhou\, Harvard\nBiobank-scale genetic discovery: from association testing to global meta-analysis\nAbstract: Biobanks linking genomic data with electronic health records provide unprecedented opportunities for genetic discovery for complex human diseases\, but they also pose analytical challenges that extend well beyond sample size. Within a biobank\, association studies must account for population structure and relatedness\, highly unbalanced case–control ratios\, rare genetic variants\, longitudinal and censored outcomes\, and the computational demands of analyzing hundreds of thousands of individuals and millions of genetic variants. Across biobanks\, additional challenges arise from differences in ancestry\, phenotype definitions\, recruitment strategies and genetic effects.In this talk\, I will discuss statistical and computational methods developed to address these challenges at successive stages of biobank analysis. These include scalable generalized linear mixed models for binary traits\, survival mixed models for censored time-to-event outcomes\, and gene- and region-based tests that aggregate rare variants. I will describe statistical approximations and computational strategies that make these analyses feasible at biobank scale while maintaining calibration in the presence of relatedness and highly unbalanced phenotypes. I will then introduce the Global Biobank Meta-analysis Initiative (GBMI) and describe how genetic evidence can be combined across biobanks without sharing individual-level data. Examples from GBMI will illustrate how combining evidence across biobanks can increase statistical power through larger sample sizes and broaden genetic discovery through greater ancestral diversity. \n11:30 am–12:45 pm\nLunch \n12:45–1:45 pm\nSergey Ovchinnikov\, MIT\nUsing AI for protein structure modeling and design\nAbstract: In this talk\, I’ll describe recent work in model interpretability\, focusing on decomposing protein language models and structure prediction models like AlphaFold. The goal is to understand what they are learning\, their limitations and how we can use this information to develop better models and design proteins. \n1:45–2:00 pm\nBreak \n2:00–3:00 pm\nSueYeon Chung\, Harvard/Flatiron Institute\nComputing with Neural Manifolds: A Multi-Scale Framework for Understanding Biological and Artificial Neural Networks\nRecent breakthroughs in experimental neuroscience and machine learning have opened new frontiers in understanding the computational principles governing neural circuits and artificial neural networks (ANNs). Both biological and artificial systems exhibit an astonishing degree of orchestrated information processing capabilities across multiple scales – from the microscopic responses of individual neurons to the emergent macroscopic phenomena of cognition and task functions. At the mesoscopic scale\, the structures of neuron population activities manifest themselves as neural representations. Neural computation can be viewed as a series of transformations of these representations through various processing stages of the brain. The primary focus of my lab’s research is to develop theories of neural representations that describe the principles of neural coding and\, importantly\, capture the complex structure of real data from both biological and artificial systems. \nIn this talk\, I will present three related approaches that leverage techniques from statistical physics\, machine learning\, and geometry to study the multi-scale nature of neural computation. First\, I will introduce new theories based on statistical physics and convex geometry that connect complex geometric structures that arise from neural responses (i.e.\, neural manifolds) to the efficiency of neural representations in implementing a task. Second\, I will employ these theories to analyze how these representations evolve across scales\, shaped by the properties of single neurons\, learning dynamics\, and the transformations across distinct brain regions. Finally\, I will show how these insights extend efficient coding principles beyond early sensory stages\, linking representational geometry to efficient task implementations. This framework not only help interpret and compare models of brain data but also offers a principled approach to designing ANN models for higher-level vision. This perspective opens new opportunities for using neuroscience-inspired principles to guide the development of intelligent systems. \n\n 
URL:https://live-hu-cmsa-222.pantheonsite.io/event/bigdata_2026/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:Big Data Conference,Conference,Event
ATTACH;FMTTYPE=image/png:https://live-hu-cmsa-222.pantheonsite.io/media/Big-Data-2026_ad.crop_.png
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260908T090000
DTEND;TZID=America/New_York:20260911T170000
DTSTAMP:20260902T170935Z
CREATED:20260217T174544Z
LAST-MODIFIED:20260902T170935Z
UID:10003847-1788858000-1789146000@live-hu-cmsa-222.pantheonsite.io
SUMMARY:The Geometry of Machine Learning 2026
DESCRIPTION:The Geometry of Machine Learning 2026 \nDates: September 8–11\, 2026 \nLocation: Harvard CMSA\, Room G10\, 20 Garden Street\, Cambridge MA 02138 & via Zoom Webinar \nRegister to attend in person \nRegister for Zoom Webinar \nLarge language models are presently\, and will increasingly\, be complemented by other dimensions of intelligence: formal verification and energy-based optimizers\, becoming parts of larger ecosystems. Can AIs reason geometrically and can we use geometry to reveal how data is currently processed in NNs? Can AIs reveal the geometry of mathematics\, as well as studying geometry as a subject within math. This conference is intended to continue the discussion of these topics. \nConfirmed Speakers: \n\nNada Amin\, Harvard\nRandall Balestriero\, Brown\nMichael Brenner\, Harvard and Google\nBennet Chow\, UCSD\nSurya Ganguli\, Stanford\nBoris Hanin\, Princeton\nRoi Holtzman\, Oxford\nRobert Koirala\, UCSD\nDmitry Krotov\, Dynamical Mind\nSlava Krushkal\, Virginia\nJared Duker Lichtman\, Stanford\nMike Mulligan\, UCR\, Logical Intelligence\nLuca Pesce\, Harvard\nGabriel Poesia\, U Michigan (via Zoom)\nMathew Vanherreweghe\, Logical Intelligence\nSean Welleck\, CMU (via Zoom)\nMattiew Wyart\, JHU (via Zoom)\n\nOrganizers: Michael R. Douglas (CMSA) and Mike Freedman (CMSA) \n  \nSchedule \nTuesday\, Sep. 8\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nMike Mulligan\, UCR\, Logical Intelligence\nCompression is all you need: Modeling mathematics\nAbstract: The mathematics humans discover and value (“human math”) is a vanishingly small subset of all valid deductions (“formal math”). I’ll argue that human math is distinguished by its compressibility through hierarchically nested definitions and theorems\, like a polynomial-growth space rather than the exponential-growth space one might expect when proofs are viewed as strings of symbols. The argument combines toy monoid models with an empirical analysis of MathLib\, a large Lean library of formalized mathematics we treat as a proxy for human math. I’ll close with how compression itself can serve as a measure of mathematical interest\, giving agents a sense of direction toward where human math lives. \n9:45–10:30 am\nBoris Hanin\, Princeton \n10:30–11:00 am\nBreak \n11:00–11:45 am\nSlava Krushkal\, University of Virginia \n12:00–12:45 pm\nRobert Koirala and Bennett Chow\, UCSD\nAI for Ricci flow: Discovery and Formalization \n  \nWednesday\, Sep. 9\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nSurya Ganguli\, Stanford \n9:45–10:30 am\nMathew Vanherreweghe\, Logical Intelligence \n10:30–11:00 am\nBreak \n11:00–11:45 am\nMichael Brenner\, Harvard and Google \n12:00–12:45 pm\nMattiew Wyart\, JHU (via Zoom)\nLearn from your own latents\, not from tokens\nAbstract: Language models need more than a hundred thousand times the data a child does. One explanation is that predicting raw tokens is simply the wrong level: methods like data2vec and JEPA instead train a network to predict its own internal representations\, with strong empirical results but no theory of why. Using a hierarchical grammar that models that language and images have a hidden hierarchical structure\, we quantify the gain exactly. Token-level learning needs a number of examples growing exponentially with the depth of the hierarchy; latent prediction needs a number independent of it. We also show data2vec performs this hierarchical prediction implicitly\, which suggests that explicitly stacking levels – as in H-JEPA – yields little. \n  \nThursday\, Sep. 10\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nDmitry Krotov\, Dynamical Mind\nDense Associative Memory: Physical systems for novel AI architectures\nAbstract: Dense Associative Memories are recurrent neural networks with fixed-point attractor states that are described by an energy function. In contrast to conventional Hopfield Networks\, which were popular in the 1980s\, Dense Associative Memories have a very large information storage capacity\, making them appealing tools for many problems in AI. In this talk\, I will provide an intuitive understanding and mathematical framework for this class of models and give examples of problems in AI that can be tackled using these new ideas. Specifically\, I will explore the relationship between Dense Associative Memories and transformers. I will present a neural network called the Energy Transformer\, which unifies energy-based modeling\, associative memories\, and transformers in a single architecture. I will demonstrate how Energy Transformers can be used for challenging tasks in image processing\, solve partial differential equations\, and serve as computational modules for energy-based language modeling. I will also discuss an exciting possibility of mapping these models onto analog hardware accelerators\, which could enable much more energy-efficient inference compared to GPUs. \n9:45–10:30 am\nRoi Holtzman\, Oxford \n10:30–11:00 am\nBreak \n11:00–11:45 am\nSean Welleck\, CMU (via Zoom) \n12:00–12:45 pm\nGabriel Poesia\, University of Michigan (via Zoom) \nFriday\, Sep. 11\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nNada Amin\, Harvard \n9:45–10:30 am\nRandall Balestriero\, Brown University \n10:30–11:00 am\nBreak \n11:00–11:45 am\nLuca Pesce\, Harvard CMSA\nA spiked perspective on feature learning with gradient-based methods\nAbstract: Depth is widely believed to give neural networks a clear computational advantage over shallow models\, and making this belief precise is a central problem in learning theory. We study a controlled high-dimensional setting where this can be done. The targets are hierarchical: the relevant structure is distributed across latent subspaces of decreasing dimension\, so that a shallow model must resolve all of it simultaneously\, while a deep network does not. Each layer forms an intermediate representation\, allowing learning to proceed in stages\, with every stage reducing the effective dimension of the remaining problem. We analyze this staged mechanism through the gradient descent dynamics of a deep network yielding a sharp separation in sample complexity between shallow and deep architectures. The result is a concrete account of why depth allows such functions to be learned from substantially fewer samples than shallow methods require. \n12:00–12:45 pm\nJared Duker Lichtman\, Stanford \n  \nSupport provided by Logical Intelligence. \n \n  \n 
URL:https://live-hu-cmsa-222.pantheonsite.io/event/gml_2026/
LOCATION:CMSA 20 Garden Street Cambridge\, Massachusetts 02138 United States
CATEGORIES:Conference,Event
ATTACH;FMTTYPE=image/jpeg:https://live-hu-cmsa-222.pantheonsite.io/media/GML2026-Poster.4.jpg
END:VEVENT
END:VCALENDAR