Augmented Commonsense Knowledge for Remote Object Grounding
Fuente:
arXiv
Salvato in:
| Autori principali: | Mohammadi, Bahram, Hong, Yicong, Qi, Yuankai, Wu, Qi, Pan, Shirui, Shi, Javen Qinfeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
di: Mohammadi, Bahram, et al.
Pubblicazione: (2025)
di: Mohammadi, Bahram, et al.
Pubblicazione: (2025)
Towards Commonsense Knowledge based Fuzzy Systems for Supporting Size-Related Fine-Grained Object Detection
di: Zhang, Pu, et al.
Pubblicazione: (2023)
di: Zhang, Pu, et al.
Pubblicazione: (2023)
CLAP: Isolating Content from Style through Contrastive Learning with Augmented Prompts
di: Cai, Yichao, et al.
Pubblicazione: (2023)
di: Cai, Yichao, et al.
Pubblicazione: (2023)
A Study of Commonsense Reasoning over Visual Object Properties
di: Kolari, Abhishek, et al.
Pubblicazione: (2025)
di: Kolari, Abhishek, et al.
Pubblicazione: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)
di: Liu, Huabin, et al.
Pubblicazione: (2025)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
di: Wang, Eileen, et al.
Pubblicazione: (2024)
di: Wang, Eileen, et al.
Pubblicazione: (2024)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
Causality Model for Semantic Understanding on Videos
di: Yicong, Li
Pubblicazione: (2025)
di: Yicong, Li
Pubblicazione: (2025)
Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model
di: Wen, Congcong, et al.
Pubblicazione: (2025)
di: Wen, Congcong, et al.
Pubblicazione: (2025)
VCD: A Dataset for Visual Commonsense Discovery in Images
di: Shen, Xiangqing, et al.
Pubblicazione: (2024)
di: Shen, Xiangqing, et al.
Pubblicazione: (2024)
ClassWise-CRF: Category-Specific Fusion for Enhanced Semantic Segmentation of Remote Sensing Imagery
di: Zhu, Qinfeng, et al.
Pubblicazione: (2025)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2025)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
di: Chen, Kesheng, et al.
Pubblicazione: (2026)
di: Chen, Kesheng, et al.
Pubblicazione: (2026)
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
di: Huy, Ta Duc, et al.
Pubblicazione: (2025)
di: Huy, Ta Duc, et al.
Pubblicazione: (2025)
A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
di: Cheng, Hongrong, et al.
Pubblicazione: (2023)
di: Cheng, Hongrong, et al.
Pubblicazione: (2023)
Can I Trust Your Answer? Visually Grounded Video Question Answering
di: Xiao, Junbin, et al.
Pubblicazione: (2023)
di: Xiao, Junbin, et al.
Pubblicazione: (2023)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
di: Jiang, Bowen, et al.
Pubblicazione: (2023)
di: Jiang, Bowen, et al.
Pubblicazione: (2023)
Voronoi-guided Bilateral 2D Gaussian Splatting for Arbitrary-Scale Hyperspectral Image Super-Resolution
di: Zhang, Jie, et al.
Pubblicazione: (2026)
di: Zhang, Jie, et al.
Pubblicazione: (2026)
Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray
di: Deng, Qiao, et al.
Pubblicazione: (2024)
di: Deng, Qiao, et al.
Pubblicazione: (2024)
Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
di: Xu, Qi, et al.
Pubblicazione: (2025)
di: Xu, Qi, et al.
Pubblicazione: (2025)
TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
di: Qi, Qiang, et al.
Pubblicazione: (2025)
di: Qi, Qiang, et al.
Pubblicazione: (2025)
A Simple-but-effective Baseline for Training-free Class-Agnostic Counting
di: Lin, Yuhao, et al.
Pubblicazione: (2024)
di: Lin, Yuhao, et al.
Pubblicazione: (2024)
Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
di: Jiang, Tianjiao, et al.
Pubblicazione: (2025)
di: Jiang, Tianjiao, et al.
Pubblicazione: (2025)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
di: Shang, Haoliang, et al.
Pubblicazione: (2025)
di: Shang, Haoliang, et al.
Pubblicazione: (2025)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
di: Zhang, Zhi, et al.
Pubblicazione: (2023)
di: Zhang, Zhi, et al.
Pubblicazione: (2023)
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
di: He, Allen, et al.
Pubblicazione: (2026)
di: He, Allen, et al.
Pubblicazione: (2026)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
di: Park, Jun-Hyung, et al.
Pubblicazione: (2024)
di: Park, Jun-Hyung, et al.
Pubblicazione: (2024)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
Can We Break Free from Strong Data Augmentations in Self-Supervised Learning?
di: Gowda, Shruthi, et al.
Pubblicazione: (2024)
di: Gowda, Shruthi, et al.
Pubblicazione: (2024)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
di: Wu, Wentao, et al.
Pubblicazione: (2025)
di: Wu, Wentao, et al.
Pubblicazione: (2025)
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
di: Hasan, Maram, et al.
Pubblicazione: (2026)
di: Hasan, Maram, et al.
Pubblicazione: (2026)
Efficient Adaptation For Remote Sensing Visual Grounding
di: Moughnieh, Hasan, et al.
Pubblicazione: (2025)
di: Moughnieh, Hasan, et al.
Pubblicazione: (2025)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
di: Hong, Yuyang, et al.
Pubblicazione: (2025)
di: Hong, Yuyang, et al.
Pubblicazione: (2025)
V-IRL: Grounding Virtual Intelligence in Real Life
di: Yang, Jihan, et al.
Pubblicazione: (2024)
di: Yang, Jihan, et al.
Pubblicazione: (2024)
A Novel Neural-symbolic System under Statistical Relational Learning
di: Yu, Dongran, et al.
Pubblicazione: (2023)
di: Yu, Dongran, et al.
Pubblicazione: (2023)
Investigating Long-term Training for Remote Sensing Object Detection
di: Park, JongHyun, et al.
Pubblicazione: (2024)
di: Park, JongHyun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
di: Mohammadi, Bahram, et al.
Pubblicazione: (2025) -
Towards Commonsense Knowledge based Fuzzy Systems for Supporting Size-Related Fine-Grained Object Detection
di: Zhang, Pu, et al.
Pubblicazione: (2023) -
CLAP: Isolating Content from Style through Contrastive Learning with Augmented Prompts
di: Cai, Yichao, et al.
Pubblicazione: (2023) -
A Study of Commonsense Reasoning over Visual Object Properties
di: Kolari, Abhishek, et al.
Pubblicazione: (2025) -
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)