Re-Aligning Language to Visual Objects with an Agentic Workflow
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yuming, Feng, Jiangyan, Zhang, Haodong, Gong, Lijun, Zhu, Feng, Zhao, Rui, Hou, Qibin, Cheng, Ming-Ming, Song, Yibing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
di: Dang, Ronghao, et al.
Pubblicazione: (2023)
di: Dang, Ronghao, et al.
Pubblicazione: (2023)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Zone Evaluation: Revealing Spatial Bias in Object Detection
di: Zheng, Zhaohui, et al.
Pubblicazione: (2023)
di: Zheng, Zhaohui, et al.
Pubblicazione: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
di: Wang, Jiabao, et al.
Pubblicazione: (2023)
di: Wang, Jiabao, et al.
Pubblicazione: (2023)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
di: Chen, Yuming, et al.
Pubblicazione: (2023)
di: Chen, Yuming, et al.
Pubblicazione: (2023)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
di: Zhou, Yupeng, et al.
Pubblicazione: (2024)
di: Zhou, Yupeng, et al.
Pubblicazione: (2024)
Referring Camouflaged Object Detection
di: Zhang, Xuying, et al.
Pubblicazione: (2023)
di: Zhang, Xuying, et al.
Pubblicazione: (2023)
Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
di: Fan, Deng-Ping, et al.
Pubblicazione: (2019)
di: Fan, Deng-Ping, et al.
Pubblicazione: (2019)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
di: Yin, Bo-Wen, et al.
Pubblicazione: (2025)
di: Yin, Bo-Wen, et al.
Pubblicazione: (2025)
Towards Stable 3D Object Detection
di: Wang, Jiabao, et al.
Pubblicazione: (2024)
di: Wang, Jiabao, et al.
Pubblicazione: (2024)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2025)
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2025)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
di: Li, Yunheng, et al.
Pubblicazione: (2024)
di: Li, Yunheng, et al.
Pubblicazione: (2024)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
di: Li, Yunheng, et al.
Pubblicazione: (2024)
di: Li, Yunheng, et al.
Pubblicazione: (2024)
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
di: Yuan, Xinbin, et al.
Pubblicazione: (2025)
di: Yuan, Xinbin, et al.
Pubblicazione: (2025)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
di: Yin, Bo-Wen, et al.
Pubblicazione: (2025)
di: Yin, Bo-Wen, et al.
Pubblicazione: (2025)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
di: Ouyang, Ziheng, et al.
Pubblicazione: (2025)
di: Ouyang, Ziheng, et al.
Pubblicazione: (2025)
Contrastive Masked Autoencoders are Stronger Vision Learners
di: Huang, Zhicheng, et al.
Pubblicazione: (2022)
di: Huang, Zhicheng, et al.
Pubblicazione: (2022)
Advancing Textual Prompt Learning with Anchored Attributes
di: Li, Zheng, et al.
Pubblicazione: (2024)
di: Li, Zheng, et al.
Pubblicazione: (2024)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
di: Li, Yuxuan, et al.
Pubblicazione: (2024)
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
di: Zhang, Xuying, et al.
Pubblicazione: (2025)
di: Zhang, Xuying, et al.
Pubblicazione: (2025)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
di: Yin, Bowen, et al.
Pubblicazione: (2023)
di: Yin, Bowen, et al.
Pubblicazione: (2023)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
di: Zhang, Shi-Chen, et al.
Pubblicazione: (2025)
di: Zhang, Shi-Chen, et al.
Pubblicazione: (2025)
Sora Generates Videos with Stunning Geometrical Consistency
di: Li, Xuanyi, et al.
Pubblicazione: (2024)
di: Li, Xuanyi, et al.
Pubblicazione: (2024)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
di: Li, Yunheng, et al.
Pubblicazione: (2025)
di: Li, Yunheng, et al.
Pubblicazione: (2025)
Re-thinking Co-Salient Object Detection
di: Fan, Deng-Ping, et al.
Pubblicazione: (2020)
di: Fan, Deng-Ping, et al.
Pubblicazione: (2020)
Traffic Scene Parsing through the TSP6K Dataset
di: Jiang, Peng-Tao, et al.
Pubblicazione: (2023)
di: Jiang, Peng-Tao, et al.
Pubblicazione: (2023)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
di: Zhou, Yupeng, et al.
Pubblicazione: (2023)
di: Zhou, Yupeng, et al.
Pubblicazione: (2023)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2024)
di: Zeng, Quan-Sheng, et al.
Pubblicazione: (2024)
AAformer: Auto-Aligned Transformer for Person Re-Identification
di: Zhu, Kuan, et al.
Pubblicazione: (2021)
di: Zhu, Kuan, et al.
Pubblicazione: (2021)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
Mixture of Style Experts for Diverse Image Stylization
di: Zhu, Shihao, et al.
Pubblicazione: (2026)
di: Zhu, Shihao, et al.
Pubblicazione: (2026)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
di: Li, Yunheng, et al.
Pubblicazione: (2026)
di: Li, Yunheng, et al.
Pubblicazione: (2026)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
di: Zhou, Yupeng, et al.
Pubblicazione: (2025)
di: Zhou, Yupeng, et al.
Pubblicazione: (2025)
AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
di: Cao, Guiming, et al.
Pubblicazione: (2025)
di: Cao, Guiming, et al.
Pubblicazione: (2025)
KAC: Kolmogorov-Arnold Classifier for Continual Learning
di: Hu, Yusong, et al.
Pubblicazione: (2025)
di: Hu, Yusong, et al.
Pubblicazione: (2025)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
di: Wan, Yuhao, et al.
Pubblicazione: (2024)
di: Wan, Yuhao, et al.
Pubblicazione: (2024)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
Described Object Detection: Liberating Object Detection with Flexible Expressions
di: Xie, Chi, et al.
Pubblicazione: (2023)
di: Xie, Chi, et al.
Pubblicazione: (2023)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
di: Dang, Ronghao, et al.
Pubblicazione: (2023) -
Aligning Audio-Visual Joint Representations with an Agentic Workflow
di: Mo, Shentong, et al.
Pubblicazione: (2024) -
Zone Evaluation: Revealing Spatial Bias in Object Detection
di: Zheng, Zhaohui, et al.
Pubblicazione: (2023) -
CrossKD: Cross-Head Knowledge Distillation for Object Detection
di: Wang, Jiabao, et al.
Pubblicazione: (2023) -
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
di: Chen, Yuming, et al.
Pubblicazione: (2023)