SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jialu, Cho, Jaemin, Sung, Yi-Lin, Yoon, Jaehong, Bansal, Mohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
von: Wang, Zun, et al.
Veröffentlicht: (2025)
von: Wang, Zun, et al.
Veröffentlicht: (2025)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
von: Zala, Abhay, et al.
Veröffentlicht: (2023)
von: Zala, Abhay, et al.
Veröffentlicht: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
von: Lin, Han, et al.
Veröffentlicht: (2023)
von: Lin, Han, et al.
Veröffentlicht: (2023)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
von: Wang, Zun, et al.
Veröffentlicht: (2026)
von: Wang, Zun, et al.
Veröffentlicht: (2026)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
Hierarchy-Aware Multimodal Unlearning for Medical AI
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
von: Pothiraj, Atin, et al.
Veröffentlicht: (2025)
von: Pothiraj, Atin, et al.
Veröffentlicht: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
von: Cho, Jaemin, et al.
Veröffentlicht: (2024)
von: Cho, Jaemin, et al.
Veröffentlicht: (2024)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
von: Wan, David, et al.
Veröffentlicht: (2024)
von: Wan, David, et al.
Veröffentlicht: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
Rethinking Interactive Image Segmentation with Low Latency, High Quality, and Diverse Prompts
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
Glider: Global and Local Instruction-Driven Expert Router
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
von: Lin, Han, et al.
Veröffentlicht: (2024)
von: Lin, Han, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025) -
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025) -
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024) -
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)