✨✨Latest Advances on Multimodal Large Language Models
-
Updated
Sep 4, 2026
✨✨Latest Advances on Multimodal Large Language Models
[CVPR2024] The code for "Osprey: Pixel Understanding with Visual Instruction Tuning"
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave, llava-next-video, llava-onevision, llama-3.2-vision, qwen-vl, qwen2-vl, phi3-v etc.
A collection of visual instruction tuning datasets.
🦩 Official repository of paper "Visual Instruction Tuning with Polite Flamingo" (AAAI-24 Oral)
Gamified Adversarial Prompting (GAP): Crowdsourcing AI-weakness-targeting data through gamification. Boost model performance with community-driven, strategic data collection
[EMNLP 2024] A Video Chat Agent with Temporal Prior
Vistral-V: Visual Instruction Tuning for Vistral - Vietnamese Large Vision-Language Model.
[ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
Collections of multimodal search libraries, service and research papers
[WWW 2026] Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey
Mistral assisted visual instruction data generation by following LLaVA
Pluggable PyTorch connectors that bridge frozen vision encoders to LLMs for visual instruction tuning
To associate your repository with the visual-instruction-tuning topic, visit your repo's landing page and select "manage topics."