Overview

Why This Workshop Now

Deep generative models now power image synthesis, video generation, language modeling, and scientific discovery, but their most important capabilities are still poorly understood.

Diffusion models, flow-based models, and autoregressive models have delivered impressive empirical gains. At the same time, major open questions remain around reliability, interpretability, privacy, and scientific use. This workshop creates a focused venue for theory, empirical analysis, and domain-driven applications to meet.

The program is designed around a practical scientific question: when a model appears capable, is it reproducing training data, capturing a real distributional structure, or performing a stronger form of compositional reasoning that transfers beyond what it has seen?

Memorization

Understand how over-parameterized DGMs retain, reproduce, or expose training data, and what that means for privacy, robustness, and trustworthy deployment.

Generalization

Characterize when generated samples reflect learned structure rather than template matching, and how model size, data complexity, and training dynamics shape that boundary.

Reasoning

Evaluate whether DGMs can support compositional, causal, or structured inference that matters for multi-step generation and scientific workflows.

Topics of Interest

Topics of Interest

The workshop welcomes work on foundational, empirical, and application-driven questions in deep generative models.

This workshop aims to bring together researchers working on the foundations of diffusion models, flow-based models, autoregressive models, and related generative learning frameworks. We welcome work that sharpens our understanding of what DGMs learn, how they behave under scale, and how they can be evaluated reliably.

The topics below reflect the research directions prioritized for the workshop and are intended to guide the scope of submissions and discussion during the workshop.

Memorization and Generalization

Empirical and theoretical studies of memorization, generalization, regime transitions, and the roles of capacity, data complexity, and scaling. Examples of related work can be found here.

Reasoning and Compositionality

Mechanisms for compositional, causal, or structured inference, including in-context learning, chain-of-thought, and multi-step generation.

Optimization and Inductive Bias

Learning dynamics, architecture, and implicit regularization in shaping memorization, generalization, and reasoning behavior.

Evaluation and Benchmarking

Metrics and diagnostic frameworks for distinguishing memorization from genuine generalization, plus robustness, privacy, and extrapolation benchmarks.

Scientific Discovery with DGMs

Applications in scientific machine learning, healthcare, protein design, and molecular discovery where interpretability and reasoning matter as much as raw sample quality.

Call for Papers

Call for Papers

This workshop is non-archival, and accepted papers will not appear in official proceedings. Accepted submissions will appear on OpenReview, but authors remain free to submit and publish their work elsewhere in the future. Accepted papers will be presented as talks or posters during the workshop. The workshop will select the best paper to recognize outstanding contributions in the field.

Important Dates

  • Paper Submission Deadline: 11:59 PM AoE on May 8, 2026
  • Notification of Acceptance: 11:59 PM AoE on May 23, 2026
  • Camera-Ready Deadline: 11:59 PM AoE on June 15, 2026
  • Workshop Date: 8:00 AM - 5:00 PM KST on July 10, 2026

Submission Instructions

  • Submit via OpenReview.
  • Contributed papers are expected to align with the workshop scope described in Topics of Interest.

Formatting Instructions

  • Please prepare submissions using the workshop Submission Style Template.
  • Please prepare the camera ready version using the workshop Camera Ready Style Template.
  • Papers should be a maximum of 8 pages, excluding references and appendices.
  • We recommend using the camera-ready template linked above, which uses a single-column layout. However, if you prefer to continue using the two-column template, please ensure that the author information is clearly included and that the workshop name is clearly indicated in the paper.
  • To address reviewer comments, the camera-ready version may be extended by up to two additional pages, for a maximum length of 10 pages when using the recommended single-column camera-ready template. Authors who choose to retain the two-column format may extend their paper by up to one additional page, for a maximum length of 9 pages.
  • Submission is double-blind, and authors must anonymize their manuscripts.

Schedule

Workshop Schedule

This workshop combines invited talks, contributed oral presentations, poster sessions, a panel discussion, and closing awards. Note that this schedule is subject to change.

Morning

  • 8:00 AM - 8:10 AMOpening Remarks - Qing Qu
  • 8:10 AM - 8:40 AM Ge Liu

    Structured Generative Modeling for Scientific Data: Geometry, Symmetry, and Variable Length

    Show Abstract

    Modern diffusion and flow-based generative models have become powerful tools for scientific discovery, yet many scientific problems go beyond probability paths defined in fixed-dimensional Euclidean spaces. Biological trajectories may lie on unknown manifolds, physical systems may obey Lie group symmetries, and protein or molecular design may require variable-length generation. This talk presents recent work from our group on generative modeling methods that respect these structures.

    First, I will present a flow-matching method for trajectory inference and generation on unknown manifolds. We learn a geodesically convex latent space that enables meaningful interpolation, and then pull the learned transport dynamics back to the data space. This allows the model to discover the geometry of a general scientific data manifold and use it for more effective transport and generation. Second, I will describe Trivialized Generative Models, a new family of generative models for Lie groups. TGM learns endpoint-constrained paths in a fixed Lie algebra and lifts them to the Lie group, providing a simple and powerful way to define valid group-valued generative dynamics. This endpoint formulation enables flexible path design beyond standard geodesic or exponential interpolation, avoids key complications of general Riemannian generative models, naturally supports non-compact groups, and extends to few-step consistency models and momentum-based higher-order dynamics. Finally, I will introduce Generalized Poisson Flow for variable-length protein design, where the model jointly learns the evolution of length and the generation of length-dependent multimodal features. Our method achieves state-of-the-art empirical performance across unconditional structure and sequence generation, motif scaffolding, and peptide co-design, validating a new paradigm for flexible-length protein design.

    Together, these works suggest a broader foundation for scientific generative modeling: designing probability paths that respect the geometry, symmetry, and variable-dimensional structure of scientific data.

  • 8:40 AM - 9:10 AM Taiji Suzuki

    Statistical optimality theory and adaptation method of diffusion models

    Show Abstract

    This talk explores recent theoretical and methodological advancements in diffusion models. On the theoretical side, we demonstrate that diffusion models can circumvent the curse of dimensionality by discovering intrinsic low-dimensional structures within data distributions. We discuss this capability for both continuous and discrete variables and establish their theoretical optimality, showing that these advantages are driven by their feature learning abilities.

    Methodologically, we introduce novel inference-time-adaptation/post-training techniques for discrete diffusion models. Leveraging Doob's h-transform as our primary technical tool, we demonstrate that integrating this transform with the efficient sampling capabilities of diffusion models facilitates effective inference time adaptation without explicit parameter updates. Specifically, our approach enables reinforcement learning on the unmasking order distribution without requiring model updates, yielding substantial performance improvements, and we show the unmasking order adaptation has a significant impact on the performance improvement.

  • 9:10 AM - 10:00 AMPoster & Break
  • 10:00 AM - 10:30 AM Yi Ma

    Pursuing the Nature of Intelligence

    Show Abstract

    In this talk, we will try to clarify different levels and mechanisms of intelligence from historical, scientific, mathematical, and computational perspective. From the evolution of intelligence in nature, from phylogenetic, to ontogenetic, societal, and to scientific intelligence, we will try to shed light on how to understand the true nature of the seemingly dramatic advancements in the technologies of machine intelligence in the past decade. We achieve this goal by developing a principled theoretical framework to explain deductively the practices of deep representation learning from the first principle of pursuing low-dimensional structures in data distributions. This framework not only reveals true nature hence both capabilities and limitations of the current deep architectures, and but also provides principled guidelines to develop more complete and more efficient learning architectures and systems. Eventually, we will clarify the difference and relationship between Knowledge and Intelligence, which may guide us to pursue the goal of developing systems with true intelligence, at least at the level for a predictive and generative memory. If time permits, we will also showcase some of the ongoing new technological developments towards realizing intelligence within an open real physical world.

  • 10:30 AM - 11:00 AM Surya Ganguli

    An exact information theoretic phase transition boundary between memorization and generalization in Bayesian diffusion models

    Show Abstract

    How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. We show that spatially local BIRD models closely approximate trained diffusion models early in training, across different architectures such as UNets and DiTs. We identify an exact information-theoretic phase transition boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-trained diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality. This is joint work with Henry Hunt and Mason Kamb.

  • 11:00 AM - 11:45 AM 3 Oral Presentations
  • 11:45 AM - 1:00 PMLunch Break

Afternoon

  • 1:00 PM - 1:30 PMMengdi Wang
  • 1:30 PM - 2:00 PM Cengiz Pehlevan

    Solvable Models of Test-Time Scaling

    Show Abstract

    Modern generative models increasingly spend compute at inference time. When does this extra compute improve performance, when does it saturate, and when can it backfire? This talk introduces a family of exactly solvable high-dimensional models, analyzed using tools from random matrix theory, that make these questions analytically tractable and offer a quantitative theory of test-time scaling.

  • 2:00 PM - 2:50 PMPoster & Break
  • 2:50 PM - 3:20 PM Kenji Fukumizu

    Analysis of OT-Flow Matching under Manifold Hypothesis

    Show Abstract

    To understand why generative models such as flow matching and diffusion models succeed on high-dimensional data, it is essential to analyze their behavior under the manifold hypothesis, which posits that the data distribution is supported on a low-dimensional submanifold. In this talk, we focus on Optimal Transport Conditional Flow Matching and show an exact proximal form via an extended Brenier potential under the manifold hypothesis. More precisely, the mapping to recover the target point, or the "decoder", is expressed by a proximal operator, which yields an explicit expression of the vector field. Using mathematical tools from convex analysis, we analyze the local behavior of the vector field around the manifold and show the stability of the manifold structure against perturbation to the dynamics: the dynamics does not expand or shrink exponentially along manifold directions.

  • 3:20 PM - 3:50 PM 2 Oral Presentations
  • 4:00 PM - 4:50 PMPanel Discussion
  • 4:50 PM - 5:00 PMAwards & Closing

Speakers

Confirmed Invited Speakers (Alphabetical Order)

All seven invited speakers listed here are confirmed. The speaker slate spans theory, reasoning, optimization, and scientific applications of deep generative models across multiple continents and career stages.

Ge Liu

Ge Liu

Assistant Professor, University of Illinois Urbana-Champaign

Yi Ma

Yi Ma

Chair Professor, University of Hong Kong

Organizers

Organizing Team (Alphabetical Order)

The team combines expertise across diffusion models, deep learning theory, optimization, sampling, and applications.

Wei Huang

Wei Huang

Research Scientist, RIKEN Center for Advanced Intelligence Project

Qing Qu

Qing Qu

Assistant Professor, University of Michigan

Molei Tao

Molei Tao

Professor, Georgia Institute of Technology

Peng Wang

Peng Wang

Assistant Professor, University of Macau

Renyuan Xu

Renyuan Xu

Assistant Professor, Stanford University

Student Organizers

Student Organizers (Alphabetical Order)

Justin Lee

Justin Lee

Ph.D. Student, University of Michigan

Xiao Li

Xiao Li

Postdoctoral Researcher, University of Hong Kong

Awards

Awards

Congratulations to the authors recognized for outstanding workshop submissions.

Contact

Workshop Information

For questions about participation, contributions, or logistics, contact the workshop leads below.