D-Robotics/SmolVLM2-500M-Video-Instruct-GGUF-BPU

Robotics Model

This is a quantized GGUF and ONNX version of the SmolVLM2-500M-Video-Instruct model, fine-tuned for image-text-to-text tasks, supporting video and language inputs, and released under Apache-2.0.

Catalog metadata

Repository owner
D-Robotics
Source
Hugging Face Models
Task
image text to text
Cameras
0
License
apache-2.0

Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.

Explore more robotics datasets and models