Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Peng, Lee, Yujian, Zhang, Xiaofeng, Chen, Zailong, Zhang, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval
by: Gao, Peng, et al.
Published: (2025)
by: Gao, Peng, et al.
Published: (2025)
MeInTime: Bridging Age Gap in Identity-Preserving Face Restoration
by: Song, Teer, et al.
Published: (2026)
by: Song, Teer, et al.
Published: (2026)
Dynamic Identity-Guided Attention Network for Visible-Infrared Person Re-identification
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
by: Marouf, Imad Eddine, et al.
Published: (2025)
by: Marouf, Imad Eddine, et al.
Published: (2025)
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction
by: Dong, Jiacheng, et al.
Published: (2026)
by: Dong, Jiacheng, et al.
Published: (2026)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD
by: Ntinou, Ioanna, et al.
Published: (2024)
by: Ntinou, Ioanna, et al.
Published: (2024)
Bridging the Gap between Text, Audio, Image, and Any Sequence: A Novel Approach using Gloss-based Annotation
by: Fang, Sen, et al.
Published: (2024)
by: Fang, Sen, et al.
Published: (2024)
Three Creates All: You Only Sample 3 Steps
by: Cai, Yuren, et al.
Published: (2026)
by: Cai, Yuren, et al.
Published: (2026)
A Plug-and-Play Method for Rare Human-Object Interactions Detection by Bridging Domain Gap
by: Zhang, Lijun, et al.
Published: (2024)
by: Zhang, Lijun, et al.
Published: (2024)
MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification
by: Zhao, Yujian, et al.
Published: (2025)
by: Zhao, Yujian, et al.
Published: (2025)
Remember and Recall: Associative-Memory-based Trajectory Prediction
by: Guo, Hang, et al.
Published: (2024)
by: Guo, Hang, et al.
Published: (2024)
From Pixels to Gigapixels: Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba
by: Qiu, Zhongwei, et al.
Published: (2024)
by: Qiu, Zhongwei, et al.
Published: (2024)
Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos
by: Feng, Yuang, et al.
Published: (2025)
by: Feng, Yuang, et al.
Published: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
by: Lee, Yujian, et al.
Published: (2026)
by: Lee, Yujian, et al.
Published: (2026)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
by: Qiu, Tianheng, et al.
Published: (2025)
by: Qiu, Tianheng, et al.
Published: (2025)
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
by: Gao, Ruopeng, et al.
Published: (2023)
by: Gao, Ruopeng, et al.
Published: (2023)
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
by: Long, Lin, et al.
Published: (2025)
by: Long, Lin, et al.
Published: (2025)
SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
by: Ye, Guanting, et al.
Published: (2026)
by: Ye, Guanting, et al.
Published: (2026)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
by: Dai, Ziyun, et al.
Published: (2025)
by: Dai, Ziyun, et al.
Published: (2025)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
by: Chen, Beitao, et al.
Published: (2025)
by: Chen, Beitao, et al.
Published: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
by: Peng, Kaixin, et al.
Published: (2025)
by: Peng, Kaixin, et al.
Published: (2025)
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
by: Yan, Bei, et al.
Published: (2025)
by: Yan, Bei, et al.
Published: (2025)
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur Prior
by: Wu, Haitao, et al.
Published: (2025)
by: Wu, Haitao, et al.
Published: (2025)
Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling
by: Zhang, Min, et al.
Published: (2024)
by: Zhang, Min, et al.
Published: (2024)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
by: Nie, Sen, et al.
Published: (2025)
by: Nie, Sen, et al.
Published: (2025)
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
by: Lei, Jingyu, et al.
Published: (2025)
by: Lei, Jingyu, et al.
Published: (2025)
Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
by: Yu, Chengzhi, et al.
Published: (2025)
by: Yu, Chengzhi, et al.
Published: (2025)
S2HPruner: Soft-to-Hard Distillation Bridges the Discretization Gap in Pruning
by: Lin, Weihao, et al.
Published: (2024)
by: Lin, Weihao, et al.
Published: (2024)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Similar Items
-
NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval
by: Gao, Peng, et al.
Published: (2025) -
MeInTime: Bridging Age Gap in Identity-Preserving Face Restoration
by: Song, Teer, et al.
Published: (2026) -
Dynamic Identity-Guided Attention Network for Visible-Infrared Person Re-identification
by: Gao, Peng, et al.
Published: (2024) -
Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
by: Marouf, Imad Eddine, et al.
Published: (2025) -
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
by: Zhang, Jiacheng, et al.
Published: (2026)