Saved in:
| Main Authors: | Liu, Qin, Cho, Jaemin, Bansal, Mohit, Niethammer, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.00741 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
by: Niu, Tianyi, et al.
Published: (2025)
by: Niu, Tianyi, et al.
Published: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
On The Robustness of Foundational 3D Medical Image Segmentation Models Against Imprecise Visual Prompts
by: Chattopadhyay, Soumitri, et al.
Published: (2026)
by: Chattopadhyay, Soumitri, et al.
Published: (2026)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
by: Zala, Abhay, et al.
Published: (2023)
by: Zala, Abhay, et al.
Published: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
by: Pothiraj, Atin, et al.
Published: (2025)
by: Pothiraj, Atin, et al.
Published: (2025)
Exploring Cycle Consistency Learning in Interactive Volume Segmentation
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
Zero-shot Domain Generalization of Foundational Models for 3D Medical Image Segmentation: An Experimental Study
by: Chattopadhyay, Soumitri, et al.
Published: (2025)
by: Chattopadhyay, Soumitri, et al.
Published: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
by: Cho, Jaemin, et al.
Published: (2024)
by: Cho, Jaemin, et al.
Published: (2024)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
by: Wang, Zun, et al.
Published: (2026)
by: Wang, Zun, et al.
Published: (2026)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
LiVOS: Light Video Object Segmentation with Gated Linear Matching
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
by: Wang, Zun, et al.
Published: (2025)
by: Wang, Zun, et al.
Published: (2025)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
by: Lin, Han, et al.
Published: (2026)
by: Lin, Han, et al.
Published: (2026)
Investigating Demographic Bias in Brain MRI Segmentation: A Comparative Study of Deep-Learning and Non-Deep-Learning Methods
by: Danaee, Ghazal, et al.
Published: (2025)
by: Danaee, Ghazal, et al.
Published: (2025)
BALD-SAM: Disagreement-based Active Prompting in Interactive Segmentation
by: Chowdhury, Prithwijit, et al.
Published: (2026)
by: Chowdhury, Prithwijit, et al.
Published: (2026)
PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications
by: Bonazzi, Pietro, et al.
Published: (2025)
by: Bonazzi, Pietro, et al.
Published: (2025)
Semantic Prompting with Image-Token for Continual Learning
by: Han, Jisu, et al.
Published: (2024)
by: Han, Jisu, et al.
Published: (2024)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
by: Lee, Daeun, et al.
Published: (2026)
by: Lee, Daeun, et al.
Published: (2026)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
by: Huang, Yidong, et al.
Published: (2026)
by: Huang, Yidong, et al.
Published: (2026)
SPT: Sequence Prompt Transformer for Interactive Image Segmentation
by: Cheng, Senlin, et al.
Published: (2024)
by: Cheng, Senlin, et al.
Published: (2024)
A Unified Model for Longitudinal Multi-Modal Multi-View Prediction with Missingness
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
ESAM++: Efficient Online 3D Perception on the Edge
by: Liu, Qin, et al.
Published: (2026)
by: Liu, Qin, et al.
Published: (2026)
Guiding Registration with Emergent Similarity from Pre-Trained Diffusion Models
by: Tursynbek, Nurislam, et al.
Published: (2025)
by: Tursynbek, Nurislam, et al.
Published: (2025)
$\texttt{NePhi}$: Neural Deformation Fields for Approximately Diffeomorphic Medical Image Registration
by: Tian, Lin, et al.
Published: (2023)
by: Tian, Lin, et al.
Published: (2023)
UnSCAR: Universal, Scalable, Controllable, and Adaptable Image Restoration
by: Mandal, Debabrata, et al.
Published: (2026)
by: Mandal, Debabrata, et al.
Published: (2026)
PA-SAM: Prompt Adapter SAM for High-Quality Image Segmentation
by: Xie, Zhaozhi, et al.
Published: (2024)
by: Xie, Zhaozhi, et al.
Published: (2024)
Benchmarking Human and Automated Prompting in the Segment Anything Model
by: Quesada, Jorge, et al.
Published: (2024)
by: Quesada, Jorge, et al.
Published: (2024)
NFL-BA: Near-Field Light Bundle Adjustment for SLAM in Dynamic Lighting
by: Beltran, Andrea Dunn, et al.
Published: (2024)
by: Beltran, Andrea Dunn, et al.
Published: (2024)
CARL: A Framework for Equivariant Image Registration
by: Greer, Hastings, et al.
Published: (2024)
by: Greer, Hastings, et al.
Published: (2024)
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
by: Chen, Zewen, et al.
Published: (2024)
by: Chen, Zewen, et al.
Published: (2024)
Similar Items
-
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
by: Lin, Han, et al.
Published: (2024) -
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
by: Niu, Tianyi, et al.
Published: (2025) -
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
by: Wan, David, et al.
Published: (2025) -
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024) -
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)