I'm grateful to my co-authors Jerry Yao-Chieh Hu, Mingcheng Lu, Maojiang Su, Weimin Wu, Minshuo Chen, and Prof. Han Liu for their collaboration and insights throughout this project.
I really enjoyed the geometric flavor of this one. It pushed me past the algebraic, statistical, and calculus-flavored symbol-pushing I was comfortable with and made me actually understand the picture driving the results — the velocity decomposition I proved came out of thinking about the geometry of conditional flows, not out of shuffling equations around until something worked.
Abstract
We study flow-matching transformers when data lie on low-dimensional manifolds. Following our NeurIPS work on diffusion transformers, we ask whether flow matching can exploit manifold structure through an analogous decomposition. We prove a velocity decomposition theorem: under a linear latent subspace assumption, the optimal velocity splits into tangent and orthogonal components that can be learned separately, which yields risk decoupling, identifiability of the on-manifold dynamics, and built-in stability.
That geometric insight motivates a two-headed transformer architecture and lets us establish intrinsic-dimension-optimal minimax rates for velocity approximation, velocity estimation, and distribution estimation — showing how flow-matching transformers escape the curse of dimensionality by using the intrinsic structure of the data.