CurHarsh/sft_robotics_vlm_all_task_821_Llama-3.2-11B-Vision-Instruct
Robotics Model
This is a fine-tuned vision-language model for image-text-to-text tasks in robotics, based on Llama-3.2-11B-Vision-Instruct, supporting conversational AI and text generation.
Catalog metadata
- Repository owner
- CurHarsh
- Source
- Hugging Face Models
- Task
- image text to text
- Cameras
- 0
Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.