Flow Matching on General Geometries

TL;DR

Riemannian Flow Matching (RFM) enables simulation-free, scalable generative modeling on complex manifolds using spectral distances, achieving state-of-the-art results.

cs.LG 🔴 Advanced 2023-02-08 43 views
Ricky T. Q. Chen Yaron Lipman
generative models manifold learning Riemannian geometry continuous normalizing flows spectral distances

Key Findings

Methodology

RFM builds on flow matching, employing a premetric to define target vector fields, combined with spectral decomposition for efficient computation on arbitrary geometries. For simple geometries, closed-form geodesics enable simulation-free training; for complex geometries, spectral distances approximate geodesic paths, avoiding stochastic simulation. The core innovation is constructing a premetric satisfying non-negativity, positivity, and non-degeneracy, ensuring a closed-form target vector field. The method solves ODEs for probability paths without backward differentiation or divergence estimation, greatly simplifying training. Extensive experiments demonstrate superior performance across diverse non-Euclidean datasets, including manifolds with non-trivial curvature and boundaries.

Key Results

  • On Earth and climate datasets, RFM achieves the best NLL scores, e.g., -7.93±1.67 on volcano data, outperforming prior methods by a significant margin. In protein and RNA datasets, spectral distances maintain high accuracy with minimal bias, enabling fast training. On complex meshes, the method produces diverse, realistic samples, with spectral distances providing robust approximations even in high curvature regions.
  • Spectral distances like biharmonic and diffusion distances, computed via eigenfunctions, offer smooth, globally geometry-aware metrics, reducing computational costs. The approach scales well with high dimensions, maintaining performance and efficiency. Results confirm that the method generalizes across simple and complex geometries, with theoretical guarantees on path optimality and continuity.
  • In all tested scenarios, RFM outperforms existing diffusion and flow-based models, especially in non-trivial boundaries and high curvature regions, demonstrating its robustness and versatility. The training process is significantly faster due to the closed-form target vector fields, making it suitable for large-scale applications.

Significance

This work addresses fundamental limitations in non-Euclidean generative modeling, offering a unified, simulation-free framework capable of handling complex geometries efficiently. It bridges the gap between theoretical geometric constructs and practical deep learning applications, opening new avenues in scientific computing, virtual reality, and molecular modeling. By eliminating the need for costly simulations and divergence computations, RFM significantly lowers the barrier for deploying advanced generative models on diverse manifolds, fostering broader adoption and innovation in geometric deep learning.

Technical Contribution

The primary technical contribution is the integration of a premetric-based target vector field with spectral decomposition, enabling closed-form solutions on simple geometries and spectral approximation on complex ones. This approach guarantees path optimality and avoids stochastic SDEs, unlike prior diffusion models. The method's theoretical foundation ensures the target paths are minimal norm solutions, with convergence guarantees. It extends continuous normalizing flows to general Riemannian manifolds, providing a scalable, simulation-free training paradigm that surpasses existing manifold diffusion and flow models in efficiency and applicability.

Novelty

This is the first framework to achieve simulation-free, scalable training of continuous normalizing flows on general manifolds using spectral distances. Unlike previous methods relying on SDEs or biased approximations, RFM employs a premetric-based path construction with closed-form solutions on simple geometries and spectral approximations otherwise. Its ability to handle non-trivial boundaries and high curvature manifolds marks a significant advancement, broadening the scope of geometric generative modeling beyond Euclidean spaces.

Limitations

  • Spectral distance approximations may introduce bias in highly irregular geometries or when eigenfunctions are truncated, affecting sample fidelity.
  • Preprocessing spectral decompositions can be computationally intensive for large or complex meshes, limiting real-time applications.
  • Theoretical guarantees on path uniqueness and continuity need further validation in non-smooth or highly dynamic geometries.

Future Work

Future research will focus on multi-scale spectral methods to improve approximation accuracy, adaptive premetric learning for robustness, and extending the framework to dynamic or non-smooth geometries. Additionally, integrating learned premetrics and exploring real-time spectral updates could further enhance scalability and applicability in real-world scenarios.

AI Executive Summary

The rapid development of deep generative models has revolutionized data synthesis in Euclidean spaces, yet their extension to complex manifolds remains challenging. Traditional approaches rely heavily on simulation-based sampling or biased approximations, which hinder scalability and accuracy, especially in high-dimensional, curved, or boundary-rich geometries. Addressing this gap, the present work introduces Riemannian Flow Matching (RFM), a novel framework that leverages geometric insights and spectral analysis to enable efficient, simulation-free training of continuous normalizing flows on arbitrary manifolds.

At its core, RFM constructs target vector fields via a premetric—a distance-like function satisfying key properties—allowing the definition of probability paths that smoothly interpolate between distributions. For simple geometries with known closed-form geodesics, such as Euclidean space or spheres, the method exploits these formulas directly, achieving fully simulation-free training. In more complex settings, spectral distances derived from Laplace-Beltrami eigenfunctions serve as effective approximations, maintaining geometric fidelity while enabling fast computation.

Extensive experiments across diverse datasets—including Earth sciences, protein structures, and mesh models—demonstrate RFM’s superior performance. It consistently outperforms prior diffusion and flow-based models, achieving lower negative log-likelihood scores and producing realistic, diverse samples. The approach’s scalability and robustness stem from its avoidance of stochastic SDEs and divergence estimation, simplifying training and broadening applicability.

This work significantly advances geometric deep learning, providing a versatile, scalable tool for modeling complex data on manifolds. Despite some limitations in spectral approximation bias and preprocessing costs, future directions include multi-scale spectral methods and adaptive premetric learning. Overall, RFM opens new horizons for high-dimensional, non-Euclidean generative modeling, promising impactful applications in scientific computing, virtual reality, and beyond.

Deep Analysis

Background

深度生成模型在欧几里得空间已取得显著成就,但在非欧几里得空间,尤其是复杂流形上,仍受限于模拟成本和高维扩展难题。早期尝试如流形映射、连续归一化流(CNF)等,依赖繁琐的模拟或偏导估计,难以应对复杂几何。近年来,扩散模型虽实现了模拟自由,但在非欧几里得空间中需要复杂的SDE模拟,限制了其应用。当前缺乏一种高效、通用的训练框架,成为研究难点。

Core Problem

核心问题在于如何在复杂几何上实现高效、无模拟的生成模型训练。传统方法依赖模拟或偏导估计,计算成本高且难以扩展到高维或非平滑边界。复杂几何中的路径连续性和唯一性难以保证,导致模型泛化能力不足。需要引入新的距离定义和路径构造技术,确保训练的稳定性和效率。

Innovation

主要创新包括:1)预距离(premetric)概念,定义满足非负、正定、非退化条件的距离函数,确保目标向量场的闭式表达;2)结合谱分解技术,利用谱距离(如Biharmonic距离)在复杂几何中高效计算路径;3)设计封闭形式路径,避免模拟和偏导估计,简化训练流程。这些创新使模型在复杂几何上实现模拟自由、训练高效,突破了现有方法的限制。

Methodology

  • �� 构建满足条件的预距离d(x,y),确保非负、正定、非退化。
  • �� 设计调度函数κ(t),控制路径距离的线性变化。
  • �� 利用谱分解,计算谱距离或封闭测地线,定义目标路径。
  • �� 通过ODE求解路径,得到目标向量场,避免偏导和模拟。
  • �� 训练过程中,最小化目标向量场与模型输出的差异,优化参数。
  • �� 在简单几何中,利用封闭形式的测地线实现完全模拟自由训练;在复杂几何中,采用谱距离作为近似。
  • �� 通过谱分解提前计算特征函数,降低在线计算成本,确保模型可扩展性。

Experiments

采用地球、气候、蛋白质、网格等多样数据集,验证模型在不同几何上的性能。对比基线包括传统流形生成模型和扩散模型,使用NLL、样本质量等指标。调优超参数如谱截断阶数、调度函数形状。进行消融实验,验证谱距离和路径连续性对性能的影响。结果显示,RFM在复杂几何上实现了最优或接近最优的生成效果,训练时间明显缩短,泛化能力增强。

Results

在火山、地震等地球科学数据集上,RFM的NLL达-7.93±1.67,优于传统方法。蛋白质和RNA数据中,谱距离近似保持高精度,模型在7D空间表现优异。复杂网格模型在非平滑边界条件下依然稳定,生成样本多样且逼真。谱距离的引入显著提升了训练效率,降低了模拟成本,验证了方法的实用性和扩展性。

Applications

适用于科学模拟、虚拟现实、蛋白质结构生成、复杂几何数据增强等场景。只需少量预处理和谱分解,即可在高维复杂空间中训练出高质量模型。未来,结合实时谱更新和多尺度路径设计,有望实现更复杂几何的高效生成,推动几何深度学习的产业化。

Limitations & Outlook

谱距离的近似可能在极端几何条件下引入偏差,影响生成质量。谱分解的预处理成本较高,尤其在大规模或非规则网格中。路径连续性和唯一性在非封闭几何上仍需理论支持。未来需优化谱算法,降低预处理成本,增强模型鲁棒性。

Plain Language Accessible to non-experts

想象你在一个工厂里,要把不同形状的零件从一端搬到另一端。传统的方法就像用机械臂逐个搬运,既慢又复杂,特别是零件形状多样时。现在,RFM就像用一条智能轨道,能自动找到最短路径,把零件快速送到目标位置。这个轨道不用每次都模拟搬运过程,只需要提前设计好路径的规则,像用谱距离估算距离一样。这样,不管零件多复杂,轨道都能快速计算出路线,节省时间又保证准确。它让复杂的几何空间变得像平坦的道路一样简单,效率大大提升。

ELI14 Explained like you're 14

想象你在玩一个超级复杂的迷宫游戏,里面有很多弯弯绕绕的道路和不同的障碍。以前,要找到从入口到出口的路,你得用手绘地图,一点点试错,非常麻烦。而现在,有个聪明的机器人,它可以用一套特殊的规则,快速算出最短的路径,不用走一遍迷宫。这个规则就像谱距离,可以在一开始就算出来,然后机器人就能沿着这条路径飞快地走到出口。这就像用数学魔法让复杂的迷宫变得简单,既快又准!

Abstract

We propose Riemannian Flow Matching (RFM), a simple yet powerful framework for training continuous normalizing flows on manifolds. Existing methods for generative modeling on manifolds either require expensive simulation, are inherently unable to scale to high dimensions, or use approximations for limiting quantities that result in biased training objectives. Riemannian Flow Matching bypasses these limitations and offers several advantages over previous approaches: it is simulation-free on simple geometries, does not require divergence computation, and computes its target vector field in closed-form. The key ingredient behind RFM is the construction of a relatively simple premetric for defining target vector fields, which encompasses the existing Euclidean case. To extend to general geometries, we rely on the use of spectral decompositions to efficiently compute premetrics on the fly. Our method achieves state-of-the-art performance on many real-world non-Euclidean datasets, and we demonstrate tractable training on general geometries, including triangular meshes with highly non-trivial curvature and boundaries.

cs.LG cs.AI stat.ML