DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yuzhong, Liu, Feng, Liu, Yue, Liao, Mingxiang, Gong, Chen, Ye, Qixiang, Wan, Fang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ControlCap: Controllable Region-level Captioning
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025)
by: Yang, Wenxi, et al.
Published: (2025)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024)
by: Liao, Mingxiang, et al.
Published: (2024)
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
by: Zhang, Mu, et al.
Published: (2024)
by: Zhang, Mu, et al.
Published: (2024)
AceTone: Bridging Words and Colors for Conditional Image Grading
by: Ma, Tianren, et al.
Published: (2026)
by: Ma, Tianren, et al.
Published: (2026)
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning
by: Qian, Chengxuan, et al.
Published: (2025)
by: Qian, Chengxuan, et al.
Published: (2025)
VMamba: Visual State Space Model
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation
by: Sun, Han, et al.
Published: (2025)
by: Sun, Han, et al.
Published: (2025)
Delving into Spectral Clustering with Vision-Language Representations
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Delving into Dark Regions for Robust Shadow Detection
by: Guan, Huankang, et al.
Published: (2024)
by: Guan, Huankang, et al.
Published: (2024)
Self-supervised Feature-Gate Coupling for Dynamic Network Pruning
by: Shi, Mengnan, et al.
Published: (2021)
by: Shi, Mengnan, et al.
Published: (2021)
ChatterBox: Multi-round Multimodal Referring and Grounding
by: Tian, Yunjie, et al.
Published: (2024)
by: Tian, Yunjie, et al.
Published: (2024)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
by: Zhou, Xirui, et al.
Published: (2025)
by: Zhou, Xirui, et al.
Published: (2025)
Referring Multiple Regions with Large Multimodal Models via Contextual Latent Steering
by: Xing, Yun, et al.
Published: (2026)
by: Xing, Yun, et al.
Published: (2026)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
by: Akkerman, Rick, et al.
Published: (2024)
by: Akkerman, Rick, et al.
Published: (2024)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again
by: Li, Weize, et al.
Published: (2025)
by: Li, Weize, et al.
Published: (2025)
Delving into Mapping Uncertainty for Mapless Trajectory Prediction
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)
by: Han, Yudong, et al.
Published: (2024)
Building Vision Models upon Heat Conduction
by: Wang, Zhaozhi, et al.
Published: (2024)
by: Wang, Zhaozhi, et al.
Published: (2024)
DynVFX: Augmenting Real Videos with Dynamic Content
by: Yatim, Danah, et al.
Published: (2025)
by: Yatim, Danah, et al.
Published: (2025)
DynPoint: Dynamic Neural Point For View Synthesis
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
Region-Aware CAM: High-Resolution Weakly-Supervised Defect Segmentation via Salient Region Perception
by: Dong, Hang-Cheng, et al.
Published: (2025)
by: Dong, Hang-Cheng, et al.
Published: (2025)
ChatDyn: Language-Driven Multi-Actor Dynamics Generation in Street Scenes
by: Wei, Yuxi, et al.
Published: (2024)
by: Wei, Yuxi, et al.
Published: (2024)
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
by: Liu, Xiaolu, et al.
Published: (2026)
by: Liu, Xiaolu, et al.
Published: (2026)
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2026)
by: Shang, Shuyao, et al.
Published: (2026)
DynProto: Dynamic Prototype Evolution for Out-of-Distribution Detection
by: Wu, Yanqi, et al.
Published: (2026)
by: Wu, Yanqi, et al.
Published: (2026)
CoDynTrust: Robust Asynchronous Collaborative Perception via Dynamic Feature Trust Modulus
by: Xu, Yunjiang, et al.
Published: (2025)
by: Xu, Yunjiang, et al.
Published: (2025)
Trust but Verify: Adaptive Conditioning for Reference-Based Diffusion Super-Resolution via Implicit Reference Correlation Modeling
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
Expandable Residual Approximation for Knowledge Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
Spatial Transform Decoupling for Oriented Object Detection
by: Yu, Hongtian, et al.
Published: (2023)
by: Yu, Hongtian, et al.
Published: (2023)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
Dyn-E: Local Appearance Editing of Dynamic Neural Radiance Fields
by: Zhang, Shangzan, et al.
Published: (2023)
by: Zhang, Shangzan, et al.
Published: (2023)
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025)
by: Wu, Lin, et al.
Published: (2025)
DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair
by: Li, Weihang, et al.
Published: (2024)
by: Li, Weihang, et al.
Published: (2024)
Diff-Plugin: Revitalizing Details for Diffusion-based Low-level Tasks
by: Liu, Yuhao, et al.
Published: (2024)
by: Liu, Yuhao, et al.
Published: (2024)
CLIPSym: Delving into Symmetry Detection with CLIP
by: Yang, Tinghan, et al.
Published: (2025)
by: Yang, Tinghan, et al.
Published: (2025)
Similar Items
-
ControlCap: Controllable Region-level Captioning
by: Zhao, Yuzhong, et al.
Published: (2024) -
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025) -
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024) -
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025) -
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
by: Zhang, Mu, et al.
Published: (2024)