HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ke, Fucai, Cai, Zhixi, Jahangard, Simindokht, Wang, Weiqing, Haghighi, Pari Delir, Rezatofighi, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026)
by: Ke, Fucai, et al.
Published: (2026)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024)
by: Jahangard, Simindokht, et al.
Published: (2024)
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
Explain Before You Answer: A Survey on Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
by: Cai, Zhixi, et al.
Published: (2026)
by: Cai, Zhixi, et al.
Published: (2026)
AR-Facilitated Safety Inspection and Fall Hazard Detection on Construction Sites
by: Liu, Jiazhou, et al.
Published: (2024)
by: Liu, Jiazhou, et al.
Published: (2024)
Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
by: Biswas, Sandika, et al.
Published: (2025)
by: Biswas, Sandika, et al.
Published: (2025)
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Any Convex Parametric Shapes
by: Le, Duy-Tho, et al.
Published: (2025)
by: Le, Duy-Tho, et al.
Published: (2025)
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene
by: Biswas, Sandika, et al.
Published: (2024)
by: Biswas, Sandika, et al.
Published: (2024)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
Improving Visual Perception of a Social Robot for Controlled and In-the-wild Human-robot Interaction
by: Zhong, Wangjie, et al.
Published: (2024)
by: Zhong, Wangjie, et al.
Published: (2024)
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
by: Du, Wanjun, et al.
Published: (2026)
by: Du, Wanjun, et al.
Published: (2026)
ASAP-Textured Gaussians: Enhancing Textured Gaussians with Adaptive Sampling and Anisotropic Parameterization
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
by: Biswas, Sandika, et al.
Published: (2026)
by: Biswas, Sandika, et al.
Published: (2026)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
A Neurosymbolic Agent System for Compositional Visual Reasoning
by: Xu, Yichang, et al.
Published: (2025)
by: Xu, Yichang, et al.
Published: (2025)
DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
by: Lu, Jueqing, et al.
Published: (2025)
by: Lu, Jueqing, et al.
Published: (2025)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions
by: Cai, Zhixi, et al.
Published: (2024)
by: Cai, Zhixi, et al.
Published: (2024)
DrVideo: Document Retrieval Based Long Video Understanding
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition
by: Xie, Kunyuan, et al.
Published: (2026)
by: Xie, Kunyuan, et al.
Published: (2026)
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
by: Qiu, Xuerui, et al.
Published: (2026)
by: Qiu, Xuerui, et al.
Published: (2026)
CalibAnyView: Beyond Single-View Camera Calibration in the Wild
by: Li, Boying, et al.
Published: (2026)
by: Li, Boying, et al.
Published: (2026)
Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
by: Shi, Zekai, et al.
Published: (2025)
by: Shi, Zekai, et al.
Published: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
Heterogeneous Network Based Contrastive Learning Method for PolSAR Land Cover Classification
by: Cai, Jianfeng, et al.
Published: (2024)
by: Cai, Jianfeng, et al.
Published: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
by: Dutta, Saikat, et al.
Published: (2025)
by: Dutta, Saikat, et al.
Published: (2025)
HYDRA: HYbrid knowledge Distillation and spectral Reconstruction Algorithm for high channel hyperspectral camera applications
by: Thirgood, Christopher, et al.
Published: (2025)
by: Thirgood, Christopher, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models
by: Pei, Ruiguang, et al.
Published: (2025)
by: Pei, Ruiguang, et al.
Published: (2025)
Similar Items
-
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026) -
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025) -
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024) -
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025) -
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025)