Self-Improving Small Object Grounding in LVLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Tianze, Shi, Yucheng, Sun, Ruitong, Liu, Ninghao, Sun, Jin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
OUSAC: Optimized Guidance Scheduling with Adaptive Caching for DiT Acceleration
by: Sun, Ruitong, et al.
Published: (2025)
by: Sun, Ruitong, et al.
Published: (2025)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs
by: Che, Liwei, et al.
Published: (2025)
by: Che, Liwei, et al.
Published: (2025)
Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
SoFlow: Solution Flow Models for One-Step Generative Modeling
by: Luo, Tianze, et al.
Published: (2025)
by: Luo, Tianze, et al.
Published: (2025)
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
by: Guo, Xin, et al.
Published: (2025)
by: Guo, Xin, et al.
Published: (2025)
Self-Supervised Weight Templates for Scalable Vision Model Initialization
by: Xie, Yucheng, et al.
Published: (2026)
by: Xie, Yucheng, et al.
Published: (2026)
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
by: Tan, Xuwei, et al.
Published: (2026)
by: Tan, Xuwei, et al.
Published: (2026)
B2Net: Camouflaged Object Detection via Boundary Aware and Boundary Fusion
by: Cai, Junmin, et al.
Published: (2024)
by: Cai, Junmin, et al.
Published: (2024)
Estimating Noisy Class Posterior with Part-level Labels for Noisy Label Learning
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
Study of Dropout in PointPillars with 3D Object Detection
by: Sun, Xiaoxiang, et al.
Published: (2024)
by: Sun, Xiaoxiang, et al.
Published: (2024)
Chart Deep Research in LVLMs via Parallel Relative Policy Optimization
by: Tang, Jiajin, et al.
Published: (2026)
by: Tang, Jiajin, et al.
Published: (2026)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Improving Resnet-9 Generalization Trained on Small Datasets
by: Awad, Omar Mohamed, et al.
Published: (2023)
by: Awad, Omar Mohamed, et al.
Published: (2023)
Data Augmentation For Small Object using Fast AutoAugment
by: Yoon, DaeEun, et al.
Published: (2025)
by: Yoon, DaeEun, et al.
Published: (2025)
Improving Alignment in LVLMs with Debiased Self-Judgment
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Variance-Aware Adaptive Weighting for Diffusion Model Training
by: Sun, Nanlong, et al.
Published: (2026)
by: Sun, Nanlong, et al.
Published: (2026)
Grounded Object Centric Learning
by: Kori, Avinash, et al.
Published: (2023)
by: Kori, Avinash, et al.
Published: (2023)
Foster Adaptivity and Balance in Learning with Noisy Labels
by: Sheng, Mengmeng, et al.
Published: (2024)
by: Sheng, Mengmeng, et al.
Published: (2024)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
by: Jo, Yujin, et al.
Published: (2026)
by: Jo, Yujin, et al.
Published: (2026)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
by: Salamatian, Ali, et al.
Published: (2025)
by: Salamatian, Ali, et al.
Published: (2025)
Object-level Self-Distillation for Vision Pretraining
by: Hızlı, Çağlar, et al.
Published: (2025)
by: Hızlı, Çağlar, et al.
Published: (2025)
DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
by: Sethi, Geet, et al.
Published: (2026)
by: Sethi, Geet, et al.
Published: (2026)
Relating Events and Frames Based on Self-Supervised Learning and Uncorrelated Conditioning for Unsupervised Domain Adaptation
by: Rostami, Mohammad, et al.
Published: (2024)
by: Rostami, Mohammad, et al.
Published: (2024)
Is Bigger Always Better? Efficiency Analysis in Resource-Constrained Small Object Detection
by: Mbobda-Kuate, Kwame, et al.
Published: (2026)
by: Mbobda-Kuate, Kwame, et al.
Published: (2026)
CPPF++: Uncertainty-Aware Sim2Real Object Pose Estimation by Vote Aggregation
by: You, Yang, et al.
Published: (2022)
by: You, Yang, et al.
Published: (2022)
Interpretable Dynamic Graph Neural Networks for Small Occluded Object Detection and Tracking
by: Soudeep, Shahriar, et al.
Published: (2024)
by: Soudeep, Shahriar, et al.
Published: (2024)
A Data-Driven RetinaNet Model for Small Object Detection in Aerial Images
by: Tang, Zhicheng, et al.
Published: (2025)
by: Tang, Zhicheng, et al.
Published: (2025)
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
by: Jung, Whie, et al.
Published: (2024)
by: Jung, Whie, et al.
Published: (2024)
Improved Object-Based Style Transfer with Single Deep Network
by: Kulkarni, Harshmohan, et al.
Published: (2024)
by: Kulkarni, Harshmohan, et al.
Published: (2024)
Moving Object Proposals with Deep Learned Optical Flow for Video Object Segmentation
by: Shi, Ge, et al.
Published: (2024)
by: Shi, Ge, et al.
Published: (2024)
SMaRt: Improving GANs with Score Matching Regularity
by: Xia, Mengfei, et al.
Published: (2023)
by: Xia, Mengfei, et al.
Published: (2023)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025)
by: Hua, Zhenglin, et al.
Published: (2025)
Similar Items
-
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025) -
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
by: Yang, Tianze, et al.
Published: (2025) -
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
by: Shi, Yucheng, et al.
Published: (2025) -
OUSAC: Optimized Guidance Scheduling with Adaptive Caching for DiT Acceleration
by: Sun, Ruitong, et al.
Published: (2025) -
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)