ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
The article introduces ChainFlow-VLA, a new approach to causal flow planning that integrates vision-language models. This method addresses the limitations of current autonomous driving systems by unifying causal modeling and global optimization. Experiments show that ChainFlow-VLA achieves state-of-the-art performance, matching human-level capabilities in trajectory planning.
- ▪ChainFlow-VLA combines causal generation and global refinement in a single probabilistic framework.
- ▪The method utilizes an autoregressive generator to produce causal trajectory modes, followed by a diffusion-based refiner.
- ▪ChainFlow-VLA achieved a score of 94.85 on the NAVSIM v1 leaderboard, comparable to human-level performance.
arXiv cs.AI files mainly under ai research. We currently carry 1,128 of its stories.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Computer Vision and Pattern Recognition arXiv:2605.23270 (cs) [Submitted on 22 May 2026] Title:ChainFlow-VLA: Causal Flow Planning with Vision-Language Models Authors:Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen, Xingtai Gui, Zhi Xu, Xiaolei Wu, Feiyang Tan, Hangning Zhou, Mu Yang View a PDF of the paper titled ChainFlow-VLA: Causal Flow Planning with Vision-Language Models, by Xiyang Wang and 9 other authors View PDF HTML (experimental) Abstract:Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv cs.AI.