ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yujun, Aniri, Bi, Jinhe, Pirk, Soeren, Ma, Yunpu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
di: Bi, Jinhe, et al.
Pubblicazione: (2024)
di: Bi, Jinhe, et al.
Pubblicazione: (2024)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
di: Hassan, Umair
Pubblicazione: (2025)
di: Hassan, Umair
Pubblicazione: (2025)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
di: Françani, André O., et al.
Pubblicazione: (2024)
di: Françani, André O., et al.
Pubblicazione: (2024)
CylinderPlane: Nested Cylinder Representation for 3D-aware Image Generation
di: Jia, Ru, et al.
Pubblicazione: (2025)
di: Jia, Ru, et al.
Pubblicazione: (2025)
PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
di: Jiang, Wei, et al.
Pubblicazione: (2026)
di: Jiang, Wei, et al.
Pubblicazione: (2026)
BID: Boundary-Interior Decoding for Unsupervised Temporal Action Localization Pre-Trainin
di: Fang, Qihang, et al.
Pubblicazione: (2024)
di: Fang, Qihang, et al.
Pubblicazione: (2024)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
di: Han, Haoxuan, et al.
Pubblicazione: (2026)
di: Han, Haoxuan, et al.
Pubblicazione: (2026)
TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery
di: Wu, Yanan, et al.
Pubblicazione: (2026)
di: Wu, Yanan, et al.
Pubblicazione: (2026)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
di: Rao, Penghao, et al.
Pubblicazione: (2025)
di: Rao, Penghao, et al.
Pubblicazione: (2025)
Robust Perspective Correction for Real-World Crack Evolution Tracking in Image-Based Structural Health Monitoring
di: Sun, Xinxin, et al.
Pubblicazione: (2025)
di: Sun, Xinxin, et al.
Pubblicazione: (2025)
ADLDA: A Method to Reduce the Harm of Data Distribution Shift in Data Augmentation
di: Wang, Haonan
Pubblicazione: (2024)
di: Wang, Haonan
Pubblicazione: (2024)
An inpainting approach to manipulate asymmetry in pre-operative breast images
di: Montenegro, Helena, et al.
Pubblicazione: (2025)
di: Montenegro, Helena, et al.
Pubblicazione: (2025)
QCFace: Image Quality Control for boosting Face Representation & Recognition
di: Doan-Ngo, Duc-Phuong, et al.
Pubblicazione: (2025)
di: Doan-Ngo, Duc-Phuong, et al.
Pubblicazione: (2025)
Real-time Object and Event Detection Service through Computer Vision and Edge Computing
di: Mendes, Marcos, et al.
Pubblicazione: (2025)
di: Mendes, Marcos, et al.
Pubblicazione: (2025)
SealD-NeRF: Interactive Pixel-Level Editing for Dynamic Scenes by Neural Radiance Fields
di: Huang, Zhentao, et al.
Pubblicazione: (2024)
di: Huang, Zhentao, et al.
Pubblicazione: (2024)
Gaussian-Constrained LeJEPA Representations for Unsupervised Scene Discovery and Pose Consistency
di: Mostafa, Mohsen
Pubblicazione: (2026)
di: Mostafa, Mohsen
Pubblicazione: (2026)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
di: Michalkiewicz, Mateusz, et al.
Pubblicazione: (2024)
di: Michalkiewicz, Mateusz, et al.
Pubblicazione: (2024)
Depth Priors in Removal Neural Radiance Fields
di: Guo, Zhihao, et al.
Pubblicazione: (2024)
di: Guo, Zhihao, et al.
Pubblicazione: (2024)
MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
di: Cheng, Ruting, et al.
Pubblicazione: (2025)
di: Cheng, Ruting, et al.
Pubblicazione: (2025)
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
di: Richings, Jack, et al.
Pubblicazione: (2025)
di: Richings, Jack, et al.
Pubblicazione: (2025)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
di: Cerqueira, Matheus A., et al.
Pubblicazione: (2024)
di: Cerqueira, Matheus A., et al.
Pubblicazione: (2024)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
di: Tu, Shuyuan, et al.
Pubblicazione: (2024)
di: Tu, Shuyuan, et al.
Pubblicazione: (2024)
Bench2FreeAD: A Benchmark for Vision-based End-to-end Navigation in Unstructured Robotic Environments
di: Peng, Yuhang, et al.
Pubblicazione: (2025)
di: Peng, Yuhang, et al.
Pubblicazione: (2025)
Single-Scanline Relative Pose Estimation for Rolling Shutter Cameras
di: Hruby, Petr, et al.
Pubblicazione: (2025)
di: Hruby, Petr, et al.
Pubblicazione: (2025)
Minimal Solvers for Full DoF Motion Estimation from Asynchronous Tracks
di: Hruby, Petr, et al.
Pubblicazione: (2025)
di: Hruby, Petr, et al.
Pubblicazione: (2025)
Efficient Solution of Point-Line Absolute Pose
di: Hruby, Petr, et al.
Pubblicazione: (2024)
di: Hruby, Petr, et al.
Pubblicazione: (2024)
PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
di: Bi, Jinhe, et al.
Pubblicazione: (2025)
di: Bi, Jinhe, et al.
Pubblicazione: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
CNNtention: Can CNNs do better with Attention?
di: Kapila, Nikhil, et al.
Pubblicazione: (2024)
di: Kapila, Nikhil, et al.
Pubblicazione: (2024)
ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans
di: Abouagour, Mohamed, et al.
Pubblicazione: (2025)
di: Abouagour, Mohamed, et al.
Pubblicazione: (2025)
Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
di: Zheng, Xiaoji, et al.
Pubblicazione: (2024)
di: Zheng, Xiaoji, et al.
Pubblicazione: (2024)
Z-Order Transformer for Feed-Forward Gaussian Splatting
di: Wang, Can, et al.
Pubblicazione: (2026)
di: Wang, Can, et al.
Pubblicazione: (2026)
Representation Paradigms in AI-based 3D Radiological Image Reconstruction: A Systematic Review
di: Yang, Yuezhe, et al.
Pubblicazione: (2025)
di: Yang, Yuezhe, et al.
Pubblicazione: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution
di: Li, Xingyuan, et al.
Pubblicazione: (2025)
di: Li, Xingyuan, et al.
Pubblicazione: (2025)
GP-GS: Gaussian Processes Densification for 3D Gaussian Splatting
di: Guo, Zhihao, et al.
Pubblicazione: (2025)
di: Guo, Zhihao, et al.
Pubblicazione: (2025)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
di: Shin, Suyeon, et al.
Pubblicazione: (2024)
di: Shin, Suyeon, et al.
Pubblicazione: (2024)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
di: Rakib, Mazharul Islam, et al.
Pubblicazione: (2025)
di: Rakib, Mazharul Islam, et al.
Pubblicazione: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
di: Bi, Jinhe, et al.
Pubblicazione: (2024) -
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
di: Hassan, Umair
Pubblicazione: (2025) -
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
di: Françani, André O., et al.
Pubblicazione: (2024) -
CylinderPlane: Nested Cylinder Representation for 3D-aware Image Generation
di: Jia, Ru, et al.
Pubblicazione: (2025) -
PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines
di: Jiang, Wei, et al.
Pubblicazione: (2026)