Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Jingru, Yu, Huan, Jingxin, Yang, Xu, Chentianye, Biao, Yin, Sun, Yu, He, Shengfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Morpho-Aware Global Attention for Image Matting
por: Yang, Jingru, et al.
Publicado: (2024)
por: Yang, Jingru, et al.
Publicado: (2024)
CryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders
por: Xu, Chentianye, et al.
Publicado: (2024)
por: Xu, Chentianye, et al.
Publicado: (2024)
Towards General Visual-Linguistic Face Forgery Detection
por: Sun, Ke, et al.
Publicado: (2023)
por: Sun, Ke, et al.
Publicado: (2023)
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
por: Song, Zikai, et al.
Publicado: (2026)
por: Song, Zikai, et al.
Publicado: (2026)
Zero-shot Object Counting with Good Exemplars
por: Zhu, Huilin, et al.
Publicado: (2024)
por: Zhu, Huilin, et al.
Publicado: (2024)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
por: Zhu, Huilin, et al.
Publicado: (2025)
por: Zhu, Huilin, et al.
Publicado: (2025)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
por: Teoh, Benjamin, et al.
Publicado: (2025)
por: Teoh, Benjamin, et al.
Publicado: (2025)
StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
por: Yang, Haoxin, et al.
Publicado: (2025)
por: Yang, Haoxin, et al.
Publicado: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
por: Wei, Yana, et al.
Publicado: (2025)
por: Wei, Yana, et al.
Publicado: (2025)
Expanding Zero-Shot Object Counting with Rich Prompts
por: Zhu, Huilin, et al.
Publicado: (2025)
por: Zhu, Huilin, et al.
Publicado: (2025)
Zero-Shot Video Translation via Token Warping
por: Zhu, Haiming, et al.
Publicado: (2024)
por: Zhu, Haiming, et al.
Publicado: (2024)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
por: Huang, Ronggang, et al.
Publicado: (2025)
por: Huang, Ronggang, et al.
Publicado: (2025)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
por: Yu, Xinlei, et al.
Publicado: (2025)
por: Yu, Xinlei, et al.
Publicado: (2025)
MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention
por: Yang, Rui, et al.
Publicado: (2025)
por: Yang, Rui, et al.
Publicado: (2025)
Lagrangian Motion Fields for Long-term Motion Generation
por: Yang, Yifei, et al.
Publicado: (2024)
por: Yang, Yifei, et al.
Publicado: (2024)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
por: Lin, Honglin, et al.
Publicado: (2024)
por: Lin, Honglin, et al.
Publicado: (2024)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
por: Yuan, Haobo, et al.
Publicado: (2025)
por: Yuan, Haobo, et al.
Publicado: (2025)
CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model
por: Yin, Pengwei, et al.
Publicado: (2024)
por: Yin, Pengwei, et al.
Publicado: (2024)
PanopticQuery: Unified Query-Time Reasoning for 4D Scenes
por: Tang, Ruilin, et al.
Publicado: (2026)
por: Tang, Ruilin, et al.
Publicado: (2026)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
por: Jiang, Huajie, et al.
Publicado: (2025)
por: Jiang, Huajie, et al.
Publicado: (2025)
Towards General Visual-Linguistic Face Forgery Detection(V2)
por: Sun, Ke, et al.
Publicado: (2025)
por: Sun, Ke, et al.
Publicado: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
por: Li, Yian, et al.
Publicado: (2026)
por: Li, Yian, et al.
Publicado: (2026)
Marmot: Object-Level Self-Correction via Multi-Agent Reasoning
por: Sun, Jiayang, et al.
Publicado: (2025)
por: Sun, Jiayang, et al.
Publicado: (2025)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
por: Lassoued, Aymen, et al.
Publicado: (2026)
por: Lassoued, Aymen, et al.
Publicado: (2026)
Instruct2See: Learning to Remove Any Obstructions Across Distributions
por: Li, Junhang, et al.
Publicado: (2025)
por: Li, Junhang, et al.
Publicado: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
por: Ou, Linyu, et al.
Publicado: (2025)
por: Ou, Linyu, et al.
Publicado: (2025)
Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
por: Yu, Yuyang, et al.
Publicado: (2025)
por: Yu, Yuyang, et al.
Publicado: (2025)
Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection
por: Guo, Yu, et al.
Publicado: (2025)
por: Guo, Yu, et al.
Publicado: (2025)
Unifying Global-Local Representations in Salient Object Detection with Transformer
por: Ren, Sucheng, et al.
Publicado: (2021)
por: Ren, Sucheng, et al.
Publicado: (2021)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains
por: Xiong, Yuqi, et al.
Publicado: (2026)
por: Xiong, Yuqi, et al.
Publicado: (2026)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
por: Jiang, Haichao, et al.
Publicado: (2026)
por: Jiang, Haichao, et al.
Publicado: (2026)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
por: Ranasinghe, Kanchana, et al.
Publicado: (2024)
por: Ranasinghe, Kanchana, et al.
Publicado: (2024)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
por: Li, Kailing, et al.
Publicado: (2025)
por: Li, Kailing, et al.
Publicado: (2025)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
por: Wang, Shaoan, et al.
Publicado: (2026)
por: Wang, Shaoan, et al.
Publicado: (2026)
Towards Affordance-Aware Articulation Synthesis for Rigged Objects
por: Yu, Yu-Chu, et al.
Publicado: (2025)
por: Yu, Yu-Chu, et al.
Publicado: (2025)
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
por: Wang, Chuanming, et al.
Publicado: (2024)
por: Wang, Chuanming, et al.
Publicado: (2024)
Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
por: Cao, Yang, et al.
Publicado: (2024)
por: Cao, Yang, et al.
Publicado: (2024)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
por: Zhong, Zhide, et al.
Publicado: (2026)
por: Zhong, Zhide, et al.
Publicado: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
por: Ying, Kaining, et al.
Publicado: (2025)
por: Ying, Kaining, et al.
Publicado: (2025)
Ejemplares similares
-
Morpho-Aware Global Attention for Image Matting
por: Yang, Jingru, et al.
Publicado: (2024) -
CryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders
por: Xu, Chentianye, et al.
Publicado: (2024) -
Towards General Visual-Linguistic Face Forgery Detection
por: Sun, Ke, et al.
Publicado: (2023) -
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
por: Song, Zikai, et al.
Publicado: (2026) -
Zero-shot Object Counting with Good Exemplars
por: Zhu, Huilin, et al.
Publicado: (2024)