BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CMSA - ECPv6.17.4//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CMSA
X-ORIGINAL-URL:https://live-hu-cmsa-222.pantheonsite.io
X-WR-CALDESC:Events for CMSA
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20250309T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20251102T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20260308T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20261101T060000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20270314T070000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20271107T060000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260908T090000
DTEND;TZID=America/New_York:20260911T170000
DTSTAMP:20260910T123647Z
CREATED:20260217T174544Z
LAST-MODIFIED:20260910T123647Z
UID:10003847-1788858000-1789146000@live-hu-cmsa-222.pantheonsite.io
SUMMARY:The Geometry of Machine Learning 2026
DESCRIPTION:The Geometry of Machine Learning 2026 \nDates: September 8–11\, 2026 \nLocation: Harvard CMSA\, Room G10\, 20 Garden Street\, Cambridge MA 02138 & via Zoom Webinar \nRegister to attend in person \nRegister for Zoom Webinar \nLarge language models are presently\, and will increasingly\, be complemented by other dimensions of intelligence: formal verification and energy-based optimizers\, becoming parts of larger ecosystems. Can AIs reason geometrically and can we use geometry to reveal how data is currently processed in NNs? Can AIs reveal the geometry of mathematics\, as well as studying geometry as a subject within math. This conference is intended to continue the discussion of these topics. \nConfirmed Speakers: \n\nNada Amin\, Harvard\nRandall Balestriero\, Brown\nMichael Brenner\, Harvard and Google\nBennet Chow\, UCSD\nSurya Ganguli\, Stanford\nBoris Hanin\, Princeton\nRoi Holtzman\, Oxford\nRobert Koirala\, UCSD\nDmitry Krotov\, Dynamical Mind\nSlava Krushkal\, Virginia\nJared Duker Lichtman\, Stanford\nMike Mulligan\, UCR\, Logical Intelligence\nLuca Pesce\, Harvard\nGabriel Poesia\, U Michigan (via Zoom)\nZiyang Qin\, Cornell\nMathew Vanherreweghe\, Logical Intelligence\nSean Welleck\, CMU (via Zoom)\nMattiew Wyart\, JHU (via Zoom)\n\nOrganizers: Michael R. Douglas (CMSA) and Mike Freedman (CMSA) \n  \nSchedule \nTuesday\, Sep. 8\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nMike Mulligan\, UCR\, Logical Intelligence\nCompression is all you need: Modeling mathematics\nAbstract: The mathematics humans discover and value (“human math”) is a vanishingly small subset of all valid deductions (“formal math”). I’ll argue that human math is distinguished by its compressibility through hierarchically nested definitions and theorems\, like a polynomial-growth space rather than the exponential-growth space one might expect when proofs are viewed as strings of symbols. The argument combines toy monoid models with an empirical analysis of MathLib\, a large Lean library of formalized mathematics we treat as a proxy for human math. I’ll close with how compression itself can serve as a measure of mathematical interest\, giving agents a sense of direction toward where human math lives. \n9:30–9:45 am\nBreak \n9:45–10:30 am\nBoris Hanin\, Princeton\nThe Score Hamiltonian: Diffusion Models via Adiabatic Transport  \n10:30–11:00 am\nBreak \n11:00–11:45 am\nSlava Krushkal\, University of Virginia\nUsing AI to study 4-manifold topology\nAbstract: Central open questions in geometric classification theory of topological 4-manifolds have a reformulation in terms of the Round Handle Problem. It asks whether a given link in the 3-sphere is slice (bounds disjoint disks) in the 4-manifold obtained by attaching certain round handles to the 4-ball. I will discuss an algebraic-combinatorial formulation of the problem\, ongoing AI-assisted work on it\, and the results obtained to date. \n11:45 am–12:00 pm\nBreak \n12:00–12:45 pm\nRobert Koirala\, UCSD; Ziyang Qin\, Cornell; and Bennett Chow\, UCSD\nAI for Ricci flow: Discovery and Formalization\nWe discuss the use of AI in mathematical exploration and proof formalization through projects in Ricci flow and Riemannian geometry. The geometric part begins with the question of what information the heat kernel retains about the underlying space. We introduce the Fisher information metric associated with the conjugate heat kernel of a Ricci flow and explain its monotonicity in scale\, its relation to pointed Nash entropy\, and its connection to Euclidean splitting. We also describe a related heat-kernel approach to volume-growth estimates under curvature assumptions\, including the setting of Gromov’s volume conjecture.\nThe formalization part concerns Hamilton’s theorem that a closed three-manifold with positive Ricci curvature admits a metric of constant positive curvature. We describe its AI-assisted formalization in Lean\, including the process of identifying and proving the ingredients needed for the final theorem. Throughout\, we discuss the respective roles of AI-generated calculations and arguments\, human mathematical judgment\, and formal proof checking. Particular attention is given to the distinction between checking a formal proof and checking that its definitions and statements express the intended mathematics. \n  \nWednesday\, Sep. 9\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nSurya Ganguli\, Stanford\nTowards understanding the geometry of high dimensional nonlinear maps \n9:30–9:45 am\nBreak \n9:45–10:30 am\nMathew Vanherreweghe\, Logical Intelligence\nSparsity Before Averaging: Kolmogorov–Arnold Geometry in Language Models\nAbstract: The Kolmogorov–Arnold theorem writes any continuous multivariate function as a composition of one-dimensional functions and addition. Freedman and Mulligan recently showed that ordinary neural networks\, trained by gradient descent\, rediscover the geometry of that construction on their own\, first in regression\, then in vision. This talk asks the same question of large language models\, through the Jacobian that connects each prediction back to the model’s internal state. Measured carefully\, per context and before any averaging\, the newest language models turn out to have grown this geometry on their own to a degree. We then show that this geometry can be installed intentionally\, cheaply\, and without significant cost on downstream tasks\, and examine what this offers for interpretability. Joint work with Michael Freedman and Michael Mulligan. \n10:30–11:00 am\nBreak \n11:00–11:45 am\nMichael Brenner\, Harvard and Google\nBuilding a Science Assistant \n11:45 am–12:00 pm\nBreak \n12:00–12:45 pm\nMattiew Wyart\, JHU (via Zoom)\nLearn from your own latents\, not from tokens\nAbstract: Language models need more than a hundred thousand times the data a child does. One explanation is that predicting raw tokens is simply the wrong level: methods like data2vec and JEPA instead train a network to predict its own internal representations\, with strong empirical results but no theory of why. Using a hierarchical grammar that models that language and images have a hidden hierarchical structure\, we quantify the gain exactly. Token-level learning needs a number of examples growing exponentially with the depth of the hierarchy; latent prediction needs a number independent of it. We also show data2vec performs this hierarchical prediction implicitly\, which suggests that explicitly stacking levels – as in H-JEPA – yields little. \n  \nThursday\, Sep. 10\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nDmitry Krotov\, Dynamical Mind\nDense Associative Memory: Physical systems for novel AI architectures\nAbstract: Dense Associative Memories are recurrent neural networks with fixed-point attractor states that are described by an energy function. In contrast to conventional Hopfield Networks\, which were popular in the 1980s\, Dense Associative Memories have a very large information storage capacity\, making them appealing tools for many problems in AI. In this talk\, I will provide an intuitive understanding and mathematical framework for this class of models and give examples of problems in AI that can be tackled using these new ideas. Specifically\, I will explore the relationship between Dense Associative Memories and transformers. I will present a neural network called the Energy Transformer\, which unifies energy-based modeling\, associative memories\, and transformers in a single architecture. I will demonstrate how Energy Transformers can be used for challenging tasks in image processing\, solve partial differential equations\, and serve as computational modules for energy-based language modeling. I will also discuss an exciting possibility of mapping these models onto analog hardware accelerators\, which could enable much more energy-efficient inference compared to GPUs. \n9:30–9:45 am\nBreak \n9:45–10:30 am\nRoi Holtzman\, Oxford\nHyperparameter Transfer for Dense Associative Memories\nAbstract: Dense Associative Memories are energy-based neural networks that generalize Hopfield networks and underlie architectures such as Energy Transformers. Their tied weights and strongly nonlinear activations make standard hyperparameter-scaling prescriptions difficult to apply. \nI will discuss how to define an analogue of a thermodynamic limit for these models\, in which the input dimension\, hidden width\, dataset size\, and batch size grow together while the training dynamics remain well defined. We derive parameterizations that lead to hyperparameter transfer and even collapse of the full training dynamics across scale. Nonlinear activations reveal additional phenomena\, including a spectral instability removed by centering and an optimizer-dependent localization instability for softmax. \nThese results provide a first step toward extending muP style scaling ideas to energy-based architectures. \n10:30–11:00 am\nBreak \n11:00–11:45 am\nSean Welleck\, CMU (via Zoom)\nThe Problem is the Problem: Towards Scalable Mathematical Discovery\nAbstract: If we give AI a mathematical problem\, it can often help us find a solution. However\, research and discovery also involve choosing which problems to solve in the first place. In this talk\, I will describe Find\, Attempt\, and Recommend\, an agentic pipeline that finds open problems in the literature\, attempts to solve them\, and recommends promising problem-resolution pairs for human review. I will discuss a pilot study in combinatorics that found resolutions to several open conjectures\, along with strategies for allocating a budget of model attempts in order to maximize different discovery objectives. \n11:45 am–12:00 pm\nBreak \n12:00–12:45 pm\nGabriel Poesia\, University of Michigan (via Zoom)\nMaking Trouble: Creating Problems with LLMs for Reasoning Evaluation and Verified Programming\nAbstract: AI research most often focuses on solving challenging problems across diverse domains. Here\, we explore the complementary direction of using AI to create problems: a task that humans routinely engage in both for ourselves (e.g.\, when authoring educational material or creating olympiad competition problems) and for training and evaluating AI systems. First\, I will present The Token Games (TTG)\, an evaluation framework inspired by Renaissance-era mathematical duels\, where LLMs compete by both posing and solving programming puzzles between themselves. TTG allows us to produce Elo-style rankings that strongly correlate with expert reasoning benchmarks (like HLE and GPQA)\, despite being cheap to run and in principle avoiding saturation. In the second part\, I will present work on Formal Disco\, an open-ended system using LLM-based agents that synthesizes formally verified programs (in Dafny\, Verus and Frama-C) at scale\, from ideation to specification\, implementation and proofs\, taking seed ideas from random GitHub READMEs. The system uses its own data to improve not only at producing correct programs\, but also in producing increasingly diverse programs and specifications according to user-defined program features. Throughout\, we discuss standing challenges in understanding what makes good problems. \nFriday\, Sep. 11\, 2026 \n8:15–8:45 am\nBreakfast \n8:45–9:30 am\nNada Amin\, Harvard\nCompiling Programs to Neurons\nAbstract: Neural networks are ordinarily programmed indirectly: we specify architectures\, objectives\, and data\, and rely on learning to discover a computation. I will focus on Cajal\, a typed\, higher-order linear programming language whose programs compile correctly to linear neurons and\, with iteration\, to recurrent neurons. This allows discrete programming structures such as conditionals and iteration to coexist with gradient-based learning. Experiments show that connecting compiled neurons with learned networks can improve learning speed and data efficiency. This work is led by PhD student Joey Velez-Ginorio and is joint with his advisors at UPenn\, Konrad Kording and Steve Zdancewic. I will close with a complementary direction from my work: using language models together with formal verification to generate programs and proofs with machine-checkable guarantees. Together\, these projects explore how programming languages can provide structure and control at the boundary between programs and learned systems. \n9:30–9:45 am\nBreak \n9:45–10:30 am\nRandall Balestriero\, Brown University\nCounterfactual World Models for Real World Deployment \n10:30–11:00 am\nBreak \n11:00–11:45 am\nLuca Pesce\, Harvard CMSA\nA spiked perspective on feature learning with gradient-based methods\nAbstract: Depth is widely believed to give neural networks a clear computational advantage over shallow models\, and making this belief precise is a central problem in learning theory. We study a controlled high-dimensional setting where this can be done. The targets are hierarchical: the relevant structure is distributed across latent subspaces of decreasing dimension\, so that a shallow model must resolve all of it simultaneously\, while a deep network does not. Each layer forms an intermediate representation\, allowing learning to proceed in stages\, with every stage reducing the effective dimension of the remaining problem. We analyze this staged mechanism through the gradient descent dynamics of a deep network yielding a sharp separation in sample complexity between shallow and deep architectures. The result is a concrete account of why depth allows such functions to be learned from substantially fewer samples than shallow methods require. \n11:45 am–12:00 pm\nBreak \n12:00–12:45 pm\nJared Duker Lichtman\, Stanford\nThe future of mathematics: Erdős #1196 and beyond\nAbstract: First\, we report on a recent success in human-AI collaboration\, in particular the solution of Erdős Problem #1196 and a subsequent cluster of conjectures in number theory. Joint work with Boris Alexeev\, Kevin Barreto\, Yanyang Li\, Liam Price\, Jibran Iqbal Shah\, Quanyu Tang\, and Terence Tao. Then\, time allowing\, we outline a viewpoint on the future of mathematics from the perspective of expansion and compression \nSupport provided by Logical Intelligence. \n \n  \n 
URL:https://live-hu-cmsa-222.pantheonsite.io/event/gml_2026/
LOCATION:CMSA 20 Garden Street Cambridge\, Massachusetts 02138 United States
CATEGORIES:Conference,Event
ATTACH;FMTTYPE=image/jpeg:https://live-hu-cmsa-222.pantheonsite.io/media/GML2026-Poster.4.jpg
END:VEVENT
BEGIN:VEVENT
DTSTART;TZID=America/New_York:20260910T163000
DTEND;TZID=America/New_York:20260910T173000
DTSTAMP:20260908T205008Z
CREATED:20260729T182154Z
LAST-MODIFIED:20260908T205008Z
UID:10003979-1789057800-1789061400@live-hu-cmsa-222.pantheonsite.io
SUMMARY:Higher Virasoro algebras
DESCRIPTION:Geometry and Mathematical Physics Seminar \nSpeaker: Brian Williams\, Boston University \nTitle: Higher Virasoro algebras \nAbstract: I will introduce and classify central extensions of the dg Lie algebra of derived global sections of the tangent sheaf on the punctured d-disk for d > 1. Then\, I will formulate and sketch the proof of a formal and universal version of the Grothendieck–Riemann–Roch theorem. Time permitting I will discuss how to generalize this work to complements of other algebraic subvarities of affine space. This is joint work with Zhengping Gui.
URL:https://live-hu-cmsa-222.pantheonsite.io/event/dgphys_91026/
LOCATION:CMSA Room G10\, CMSA\, 20 Garden Street\, Cambridge\, MA\, 02138\, United States
CATEGORIES:Geometry and Mathematical Physics Seminar
ATTACH;FMTTYPE=image/png:https://live-hu-cmsa-222.pantheonsite.io/media/CMSA-Geometry-Math-Physics-Seminar-9.10.26-scaled.png
END:VEVENT
END:VCALENDAR