Big Data Conference 2026

Big Data Conference 2026
Dates: Sep. 3–4, 2026
Location: Harvard University CMSA, 20 Garden Street, Cambridge MA & via Zoom
The Big Data Conference features speakers from the Harvard community as well as scholars from across the globe, with talks focusing on computer science, statistics, math and physics, and economics.
Confirmed Speakers
- SueYeon Chung, Harvard
- Bailey Flanigan, MIT
- Sergey Ovchinnikov, MIT
- Ariel Procaccia, Harvard
- Adit Radhakrishnan, MIT
- Andrew Sutherland, MIT
- Chris Wiggins, Columbia
- Rex Ying, Yale
- Wei Zhou, Harvard
Organizers
Thursday, Sep. 3, 2026
8:45–9:10 am
Breakfast
9:10–9:15 am
Introductions
9:15–10:15 am
Rex Ying, Yale
A Riemannian Geometry Perspective on Foundation Models
Abstract: In development of foundation models and LLMs, Euclidean space has been the de facto geometric setting for machine learning architectures. However, at a large scale, real-world data often exhibit inherently non-Euclidean structures, such as multi-way relationships, hierarchies, symmetries, and non-isotropic scaling, in a variety of domains, such as languages, vision, and the natural sciences. It is challenging to effectively capture these structures within the constraints of Euclidean spaces. This talk will cover my research that moves beyond Euclidean geometry to maintain the scaling law for the next generation of foundation models. By adopting these geometries, foundation models could more efficiently leverage the structures. Task-aware adaptability that dynamically reconfigures embeddings to match the geometry of downstream applications, could further enhance efficiency and expressivity. This talk will demonstrate the advantages of hyperbolic geometry in building foundation models from individual non-Euclidean components, LLM pre-training / fine-tuning, retrieval-augmented generation, and mixture of expert models. We further present a comprehensive library that greatly simplifies the process of building and adapting foundation models in non-Euclidean. Finally, the talk will enumerate a few promising directions that are being actively explored in this domain.
10:15–10:30 am
Break
10:30–11:30 am
Adit Radhakrishnan, MIT
Toward universal steering and monitoring of AI models
Abstract: Artificial intelligence (AI) models contain much of human knowledge. Understanding the representation of this knowledge will lead to improvements in model capabilities and safeguards. Building on advances in feature learning, we developed an approach for extracting linear representations of semantic notions or concepts in AI models. We showed how these representations enabled model steering, through which we exposed vulnerabilities and improved model capabilities. We demonstrated that concept representations were transferable across languages and enabled multiconcept steering. Across hundreds of concepts, we found that larger models were more steerable and that steering improved model capabilities beyond prompting. We showed that concept representations were more effective for monitoring misaligned content than for using judge models. Our results illustrate the power of internal representations for advancing AI safety and model capabilities.
11:30 am–12:45 pm
Lunch
12:45–1:45 pm
Andrew Sutherland, MIT
Crowdsourcing mathematical databases
Abstract: The L-functions and Modular Forms Database (LMFDB) is one of the largest online repositories of mathematical research data. It contains comprehensive catalogs of mathematical objects that arise in the context of the Langlands program, including number fields, elliptic curves, modular forms, and higher-dimensional analogs of these objects. In cooperation with the Foundation for Science and AI Research (SAIR), we recently ran a competition aimed at solving the inverse Galois problem over Q in degree 24 (IGP24). This goal was achieved by collecting more than 50 million candidate number fields submitted by more than 100 teams, including examples that realize all 25,000 transitive permutation groups of degree 24 as Galois groups. AI tools and computer algebra systems played a central role in this project, both in building the infrastructure to run the contest and in enhancing the capabilities of participants.
1:45–2:00 pm
Break
2:00–3:00 pm
Bailey Flanigan, MIT
Algorithmic Tools for Trading Off Sortition Ideals
Abstract: Citizens’ assemblies and other deliberative minipublics — representative groups of everyday people convened to deliberate on a policy issue and then make recommendations — are now used by governments around the world. Choosing who sits on these panels is the problem of sortition: randomly selecting a small group of citizens that represents the broader population. Sortition has been the subject of substantial computer science research in recent years, and the resulting algorithms are now widely used in practice. This talk will describe the key challenges that arise in the practice of sortition, and the algorithmic tools that have been developed to navigate them optimally.
3:00–3:15 pm
Break
3:15–4:15 pm
Chris Wiggins, Columbia
Prescriptive Learning at Scale: Applied and Deployed Decision-Making
Abstract: Many applications in health and industry require making decisions while learning how the world responds to them: e.g., which article to recommend, which treatment to assign, or when to show a paywall. This problem is prescriptive, what to do rather than what is true. Statistical decision theory has taken up the problem repeatedly since Neyman’s interwar work; recent decades have brought a new wave of results, motivated in large part by the opportunity to make and evaluate decisions at scale online. I will highlight a thread of 21st-century results running through adaptive experimentation, contextual bandits, and off-policy evaluation: they combine advances in supervised learning with fundamental ideas from statistical sampling, and now shape digital products, from content recommendation to marketing, as well as adaptive interventions in health.
Friday, Sep. 4, 2026
8:45–9:15 am
Breakfast
9:15–10:15 am
Ariel Procaccia, Harvard
No Generation Without Representation
Abstract: AI systems and democratic processes are confronting similar challenges around representation. I examine two related questions that cut across both domains. First, how can AI enable democratic processes that handle vast spaces of opinions or statements while ensuring proportional representation of a population’s views? Second, when AI systems themselves provide normative guidance, whose viewpoints do they reflect, and can we make this precise? Drawing on social choice theory, I present formal frameworks and algorithms for both problems, showing that meaningful representation guarantees are feasible and practical.
10:15–10:30 am
Break
10:30–11:30 am
Wei Zhou, Harvard
Biobank-scale genetic discovery: from association testing to global meta-analysis
Abstract: Biobanks linking genomic data with electronic health records provide unprecedented opportunities for genetic discovery for complex human diseases, but they also pose analytical challenges that extend well beyond sample size. Within a biobank, association studies must account for population structure and relatedness, highly unbalanced case–control ratios, rare genetic variants, longitudinal and censored outcomes, and the computational demands of analyzing hundreds of thousands of individuals and millions of genetic variants. Across biobanks, additional challenges arise from differences in ancestry, phenotype definitions, recruitment strategies and genetic effects.In this talk, I will discuss statistical and computational methods developed to address these challenges at successive stages of biobank analysis. These include scalable generalized linear mixed models for binary traits, survival mixed models for censored time-to-event outcomes, and gene- and region-based tests that aggregate rare variants. I will describe statistical approximations and computational strategies that make these analyses feasible at biobank scale while maintaining calibration in the presence of relatedness and highly unbalanced phenotypes. I will then introduce the Global Biobank Meta-analysis Initiative (GBMI) and describe how genetic evidence can be combined across biobanks without sharing individual-level data. Examples from GBMI will illustrate how combining evidence across biobanks can increase statistical power through larger sample sizes and broaden genetic discovery through greater ancestral diversity.
11:30 am–12:45 pm
Lunch
12:45–1:45 pm
Sergey Ovchinnikov, MIT
Using AI for protein structure modeling and design
Abstract: In this talk, I’ll describe recent work in model interpretability, focusing on decomposing protein language models and structure prediction models like AlphaFold. The goal is to understand what they are learning, their limitations and how we can use this information to develop better models and design proteins.
1:45–2:00 pm
Break
2:00–3:00 pm
SueYeon Chung, Harvard/Flatiron Institute
Computing with Neural Manifolds: A Multi-Scale Framework for Understanding Biological and Artificial Neural Networks
Recent breakthroughs in experimental neuroscience and machine learning have opened new frontiers in understanding the computational principles governing neural circuits and artificial neural networks (ANNs). Both biological and artificial systems exhibit an astonishing degree of orchestrated information processing capabilities across multiple scales – from the microscopic responses of individual neurons to the emergent macroscopic phenomena of cognition and task functions. At the mesoscopic scale, the structures of neuron population activities manifest themselves as neural representations. Neural computation can be viewed as a series of transformations of these representations through various processing stages of the brain. The primary focus of my lab’s research is to develop theories of neural representations that describe the principles of neural coding and, importantly, capture the complex structure of real data from both biological and artificial systems.
In this talk, I will present three related approaches that leverage techniques from statistical physics, machine learning, and geometry to study the multi-scale nature of neural computation. First, I will introduce new theories based on statistical physics and convex geometry that connect complex geometric structures that arise from neural responses (i.e., neural manifolds) to the efficiency of neural representations in implementing a task. Second, I will employ these theories to analyze how these representations evolve across scales, shaped by the properties of single neurons, learning dynamics, and the transformations across distinct brain regions. Finally, I will show how these insights extend efficient coding principles beyond early sensory stages, linking representational geometry to efficient task implementations. This framework not only help interpret and compare models of brain data but also offers a principled approach to designing ANN models for higher-level vision. This perspective opens new opportunities for using neuroscience-inspired principles to guide the development of intelligent systems.