paper/arxiv/2608.26067
Robotics Dataset
StreamPI is a streaming multimodal temporal modeling framework that enhances single-frame Vision-Language-Action (VLA) models with temporal reasoning without adding parameters, using instruction-anchored temporal modeling and random-interval streaming training for robust asynchronous deployment.
Catalog metadata
- Repository owner
- Zhe Liu, Jinghua Hou, Yuxiang Lu et al.
- Source
- Research paper
- Task
- StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
- Cameras
- 0
Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.