CurHarsh/sft_robotics_vlm_all_task_821_Llama-3.2-11B-Vision-Instruct

Robotics Model

This is a fine-tuned vision-language model for image-text-to-text tasks in robotics, based on Llama-3.2-11B-Vision-Instruct, supporting conversational AI and text generation.

Catalog metadata

Repository owner
CurHarsh
Source
Hugging Face Models
Task
image text to text
Cameras
0

Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.

Explore more robotics datasets and models