自动驾驶视觉-语言-动作模型综述(英)_21页_513kb
报告摘要
Summary of "A Survey on Vision-Language-Action Models for Autonomous Driving"
Introduction: This survey explores the Vision-Language-Action (VLA) models for autonomous driving (AD), outlining their evolution from end-to-end AD systems. It discusses how VLA models integrate vision, language, and action to enhance explainability, robustness, and human interaction, addressing limitations like lack of natural interfaces and black-box behavior in traditional AD.
VLA4AD Architecture: The architecture involves multimodal inputs (sensors like vision and LiDAR, language commands, and planned trajectories) and outputs (control actions and explanations). Core components include a vision encoder (e.g., self-supervised backbones), language processor (e.g., large language models), and action decoder, often with hierarchical controllers. Tri-modal data scarcity remains a key challenge.
Progress of VLA4AD Models: VLA models have evolved through four stages: (1) Pre-VLA Explainers (e.g., language for post-hoc explanations, but passive); (2) Modular Systems (e.g., integrating language for active planning with symbolic rules); (3) Unified End-to-End Models (e.g., mapping all inputs directly to actions); (4) Reasoning-Augmented Models (e.g., incorporating iterative thinking for long-term planning). These developments improve reasoning and interaction but face issues like latency and hallucinations.
Datasets and Benchmarks: High-quality datasets are crucial for training, including real-world multi-sensor data (e.g., nuScenes), safety-focused tests, and fine-grained reasoning data (e.g., Reason2Drive). Benchmarks involve metrics for both driving performance (e.g., closed-loop success rate) and language interaction (e.g., BLEU scores), with datasets and eval suites driving model advancement.
Training and Evaluation: Training typically involves pre-training vision-language models, fine-tuning with tri-modal data (vision, language, and control), and reinforcement learning for corner cases. Evaluation assesses driving, language generation, and robustness, using benchmarks like DriveBench to test against sensor noise and out-of-distribution scenarios.
Challenges and Future Directions: Key challenges include LLM hallucinations, real-time performance (e.g., achieving 30 Hz execution), formal safety verification, and data scarcity. Future work emphasizes neuro-symbolic integration for hybrid safety kernels, scalable memory for reasoning, fleet-scale continual learning for collaborative knowledge sharing, and ongoing data collection.
Conclusion: This survey provides a comprehensive overview of VLA4AD, highlighting its progression from interpretability to unified systems. Progress is accelerated by datasets and benchmarks, but challenges in robustness and real-time constraints persist. A unified evaluation framework and continued innovation are essential for practical deployment.
试读结束,高清完整版pdf/doc/ppt,请点下载