Saved in:
| Main Authors: | Dalal, Dwip, Mishra, Utkarsh, Ahuja, Narendra, Jojic, Nebojsa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.15933 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
by: Dalal, Dwip, et al.
Published: (2025)
by: Dalal, Dwip, et al.
Published: (2025)
Learning Informative Attention Weights for Person Re-Identification
by: Wang, Yancheng, et al.
Published: (2025)
by: Wang, Yancheng, et al.
Published: (2025)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024)
by: Liu, Xinhao, et al.
Published: (2024)
Self-Calibrating 4D Novel View Synthesis from Monocular Videos Using Gaussian Splatting
by: Li, Fang, et al.
Published: (2024)
by: Li, Fang, et al.
Published: (2024)
RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
by: Li, Fang, et al.
Published: (2025)
by: Li, Fang, et al.
Published: (2025)
Measuring the (Un)Faithfulness of Concept-Based Explanations
by: Kumar, Shubham, et al.
Published: (2025)
by: Kumar, Shubham, et al.
Published: (2025)
S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Navigating Label Ambiguity for Facial Expression Recognition in the Wild
by: Lee, JunGyu, et al.
Published: (2025)
by: Lee, JunGyu, et al.
Published: (2025)
Learning Implicit Representation for Reconstructing Articulated Objects
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
by: Zhao, Xunyi, et al.
Published: (2025)
by: Zhao, Xunyi, et al.
Published: (2025)
Piecewise-Linear Manifolds for Deep Metric Learning
by: Bhatnagar, Shubhang, et al.
Published: (2024)
by: Bhatnagar, Shubhang, et al.
Published: (2024)
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
by: Mei, Yanghong, et al.
Published: (2025)
by: Mei, Yanghong, et al.
Published: (2025)
CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
by: Lee, Jungdae, et al.
Published: (2024)
by: Lee, Jungdae, et al.
Published: (2024)
PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
by: Hu, Xiaodan, et al.
Published: (2025)
by: Hu, Xiaodan, et al.
Published: (2025)
MagicPose4D: Crafting Articulated Models with Appearance and Motion Control
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?
by: Konstantinidou, Despina, et al.
Published: (2025)
by: Konstantinidou, Despina, et al.
Published: (2025)
Potential Field Based Deep Metric Learning
by: Bhatnagar, Shubhang, et al.
Published: (2024)
by: Bhatnagar, Shubhang, et al.
Published: (2024)
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
by: Rawlekar, Samyak, et al.
Published: (2024)
by: Rawlekar, Samyak, et al.
Published: (2024)
Efficiently Disentangling CLIP for Multi-Object Perception
by: Rawlekar, Samyak, et al.
Published: (2025)
by: Rawlekar, Samyak, et al.
Published: (2025)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Scaling Vision-and-Language Navigation With Offline RL
by: Bundele, Valay, et al.
Published: (2024)
by: Bundele, Valay, et al.
Published: (2024)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
by: Han, Mingfei, et al.
Published: (2026)
by: Han, Mingfei, et al.
Published: (2026)
Enhanced Survival Prediction in Head and Neck Cancer Using Convolutional Block Attention and Multimodal Data Fusion
by: Farooq, Aiman, et al.
Published: (2024)
by: Farooq, Aiman, et al.
Published: (2024)
CountQA: How Well Do MLLMs Count in the Wild?
by: Tamarapalli, Jayant Sravan, et al.
Published: (2025)
by: Tamarapalli, Jayant Sravan, et al.
Published: (2025)
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation
by: Zhu, Junyou, et al.
Published: (2024)
by: Zhu, Junyou, et al.
Published: (2024)
Aligning Knowledge Graph with Visual Perception for Object-goal Navigation
by: Xu, Nuo, et al.
Published: (2024)
by: Xu, Nuo, et al.
Published: (2024)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
by: Nayak, Abhijeet, et al.
Published: (2025)
by: Nayak, Abhijeet, et al.
Published: (2025)
Explore and Explain: Self-supervised Navigation and Recounting
by: Bigazzi, Roberto, et al.
Published: (2020)
by: Bigazzi, Roberto, et al.
Published: (2020)
NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
by: Qin, Yiran, et al.
Published: (2025)
by: Qin, Yiran, et al.
Published: (2025)
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning
by: Zhang, Sixian, et al.
Published: (2026)
by: Zhang, Sixian, et al.
Published: (2026)
WildAvatar: Learning In-the-wild 3D Avatars from the Web
by: Huang, Zihao, et al.
Published: (2024)
by: Huang, Zihao, et al.
Published: (2024)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026)
by: Sun, Peiwen, et al.
Published: (2026)
Similar Items
-
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
by: Dalal, Dwip, et al.
Published: (2025) -
Learning Informative Attention Weights for Person Re-Identification
by: Wang, Yancheng, et al.
Published: (2025) -
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024) -
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
by: Liu, Xinhao, et al.
Published: (2024) -
Self-Calibrating 4D Novel View Synthesis from Monocular Videos Using Gaussian Splatting
by: Li, Fang, et al.
Published: (2024)