Representation Understanding via Activation Maximization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hongbo, Cangelosi, Angelo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Noise-Free Explanation for Driving Action Prediction
by: Zhu, Hongbo, et al.
Published: (2024)
by: Zhu, Hongbo, et al.
Published: (2024)
Attributes-aware Visual Emotion Representation Learning
by: Maharjan, Rahul Singh, et al.
Published: (2025)
by: Maharjan, Rahul Singh, et al.
Published: (2025)
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
by: Tavella, Federico, et al.
Published: (2025)
by: Tavella, Federico, et al.
Published: (2025)
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
by: Xie, Haodong, et al.
Published: (2026)
by: Xie, Haodong, et al.
Published: (2026)
The Safety Challenge of World Models for Embodied AI Agents: A Review
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
by: Wei, Zhihua, et al.
Published: (2026)
by: Wei, Zhihua, et al.
Published: (2026)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
by: Liang, Yue, et al.
Published: (2026)
by: Liang, Yue, et al.
Published: (2026)
HexPlane Representation for 3D Semantic Scene Understanding
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
by: Benhammou, Yassir, et al.
Published: (2025)
by: Benhammou, Yassir, et al.
Published: (2025)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Visualizing and Controlling Cortical Responses Using Voxel-Weighted Activation Maximization
by: Shinkle, Matthew W., et al.
Published: (2025)
by: Shinkle, Matthew W., et al.
Published: (2025)
Data Pruning by Information Maximization
by: Tan, Haoru, et al.
Published: (2025)
by: Tan, Haoru, et al.
Published: (2025)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
Introducing 3D Representation for Medical Image Volume-to-Volume Translation via Score Fusion
by: Zhu, Xiyue, et al.
Published: (2025)
by: Zhu, Xiyue, et al.
Published: (2025)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
by: Murlidaran, Shravan, et al.
Published: (2026)
by: Murlidaran, Shravan, et al.
Published: (2026)
Hybrid Primal Sketch: Combining Analogy, Qualitative Representations, and Computer Vision for Scene Understanding
by: Forbus, Kenneth D., et al.
Published: (2024)
by: Forbus, Kenneth D., et al.
Published: (2024)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
Cut2Next: Generating Next Shot via In-Context Tuning
by: He, Jingwen, et al.
Published: (2025)
by: He, Jingwen, et al.
Published: (2025)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion
by: Zhu, Jichao, et al.
Published: (2026)
by: Zhu, Jichao, et al.
Published: (2026)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024)
by: Wei, Shicai, et al.
Published: (2024)
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
by: Belmecheri, Nassim, et al.
Published: (2025)
by: Belmecheri, Nassim, et al.
Published: (2025)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
by: Suo, Yucheng, et al.
Published: (2025)
by: Suo, Yucheng, et al.
Published: (2025)
LVD-GS: Gaussian Splatting SLAM for Dynamic Scenes via Hierarchical Explicit-Implicit Representation Collaboration Rendering
by: Zhu, Wenkai, et al.
Published: (2025)
by: Zhu, Wenkai, et al.
Published: (2025)
VISD: Enhancing Video Reasoning via Structured Self-Distillation
by: Lin, Hao, et al.
Published: (2026)
by: Lin, Hao, et al.
Published: (2026)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
by: Kawasaki, Haruka, et al.
Published: (2026)
by: Kawasaki, Haruka, et al.
Published: (2026)
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
by: Mei, Xiaodong, et al.
Published: (2025)
by: Mei, Xiaodong, et al.
Published: (2025)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
by: Pham, Trong-Thang, et al.
Published: (2026)
by: Pham, Trong-Thang, et al.
Published: (2026)
LGCA: Enhancing Semantic Representation via Progressive Expansion
by: Cao, Thanh Hieu, et al.
Published: (2025)
by: Cao, Thanh Hieu, et al.
Published: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
by: Han, Jiaming, et al.
Published: (2025)
by: Han, Jiaming, et al.
Published: (2025)
Restoring Real-World Images with an Internal Detail Enhancement Diffusion Model
by: Xiao, Peng, et al.
Published: (2025)
by: Xiao, Peng, et al.
Published: (2025)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
by: Li, Bozheng, et al.
Published: (2025)
by: Li, Bozheng, et al.
Published: (2025)
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning
by: Bai, Hongbo, et al.
Published: (2026)
by: Bai, Hongbo, et al.
Published: (2026)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
by: Zhang, Rongyu, et al.
Published: (2024)
by: Zhang, Rongyu, et al.
Published: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Contrastive Representation Distillation via Multi-Scale Feature Decoupling
by: Wang, Cuipeng, et al.
Published: (2025)
by: Wang, Cuipeng, et al.
Published: (2025)
Similar Items
-
Noise-Free Explanation for Driving Action Prediction
by: Zhu, Hongbo, et al.
Published: (2024) -
Attributes-aware Visual Emotion Representation Learning
by: Maharjan, Rahul Singh, et al.
Published: (2025) -
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
by: Tavella, Federico, et al.
Published: (2025) -
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
by: Xie, Haodong, et al.
Published: (2026) -
The Safety Challenge of World Models for Embodied AI Agents: A Review
by: Baraldi, Lorenzo, et al.
Published: (2025)