EL-VIT: Probing Vision Transformer with Interactive Visualization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Hong, Zhang, Rui, Lai, Peifeng, Guo, Chaoran, Wang, Yong, Sun, Zhida, Li, Junjie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024)
by: Xie, Hongxia, et al.
Published: (2024)
VIT-Ped: Visionary Intention Transformer for Pedestrian Behavior Analysis
by: Elkammar, Aly R., et al.
Published: (2026)
by: Elkammar, Aly R., et al.
Published: (2026)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
by: Karadag, Ozge Oztimur
Published: (2023)
by: Karadag, Ozge Oztimur
Published: (2023)
CrossVIT-augmented Geospatial-Intelligence Visualization System for Tracking Economic Development Dynamics
by: Bai, Yanbing, et al.
Published: (2024)
by: Bai, Yanbing, et al.
Published: (2024)
SMT(LIA) Sampling with High Diversity
by: Lai, Yong, et al.
Published: (2025)
by: Lai, Yong, et al.
Published: (2025)
Probing Perceptual Constancy in Large Vision-Language Models
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
OneLatent: Single-Token Compression for Visual Latent Reasoning
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
Structured Click Control in Transformer-based Interactive Segmentation
by: Xu, Long, et al.
Published: (2024)
by: Xu, Long, et al.
Published: (2024)
Improving Dialogue Discourse Parsing through Discourse-aware Utterance Clarification
by: Fan, Yaxin, et al.
Published: (2025)
by: Fan, Yaxin, et al.
Published: (2025)
ICR: Iterative Clarification and Rewriting for Conversational Search
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search
by: Cao, Zhiyu, et al.
Published: (2026)
by: Cao, Zhiyu, et al.
Published: (2026)
OpenPort Protocol: A Security Governance Specification for AI Agent Tool Access
by: Zhu, Genliang, et al.
Published: (2026)
by: Zhu, Genliang, et al.
Published: (2026)
Probing Mechanical Reasoning in Large Vision Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Sparse Visual Thought Circuits in Vision-Language Models
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
Action with Visual Primitives
by: Guo, Weilong, et al.
Published: (2026)
by: Guo, Weilong, et al.
Published: (2026)
MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models
by: Zhao, Zhengyi, et al.
Published: (2025)
by: Zhao, Zhengyi, et al.
Published: (2025)
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
Visual Distraction Undermines Moral Reasoning in Vision-Language Models
by: Yang, Xinyi, et al.
Published: (2026)
by: Yang, Xinyi, et al.
Published: (2026)
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
by: Ding, Rui, et al.
Published: (2024)
by: Ding, Rui, et al.
Published: (2024)
InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models
by: Yin, Xiaofei, et al.
Published: (2025)
by: Yin, Xiaofei, et al.
Published: (2025)
A LSTM-Transformer Model for pulsation control of pVADs
by: E, Chaoran, et al.
Published: (2025)
by: E, Chaoran, et al.
Published: (2025)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
by: Zhou, Chenyue, et al.
Published: (2025)
by: Zhou, Chenyue, et al.
Published: (2025)
Geometric Point Attention Transformer for 3D Shape Reassembly
by: Li, Jiahan, et al.
Published: (2024)
by: Li, Jiahan, et al.
Published: (2024)
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis
by: Fan, Yaxin, et al.
Published: (2023)
by: Fan, Yaxin, et al.
Published: (2023)
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
by: Deng, Zehao, et al.
Published: (2026)
by: Deng, Zehao, et al.
Published: (2026)
AdaMotif: Graph Simplification via Adaptive Motif Design
by: Zhou, Hong, et al.
Published: (2024)
by: Zhou, Hong, et al.
Published: (2024)
From Static to Interactive: Authoring Interactive Visualizations via Natural Language
by: Liu, Can, et al.
Published: (2026)
by: Liu, Can, et al.
Published: (2026)
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
by: Theodoridis, Nikos, et al.
Published: (2026)
by: Theodoridis, Nikos, et al.
Published: (2026)
Dissecting Query-Key Interaction in Vision Transformers
by: Pan, Xu, et al.
Published: (2024)
by: Pan, Xu, et al.
Published: (2024)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
Large Language Models Enhanced Hyperbolic Space Recommender Systems
by: Cheng, Wentao, et al.
Published: (2025)
by: Cheng, Wentao, et al.
Published: (2025)
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
Cardinality Estimation for High Dimensional Similarity Queries with Adaptive Bucket Probing
by: Chen, Zhonghan, et al.
Published: (2026)
by: Chen, Zhonghan, et al.
Published: (2026)
Two-stage Incomplete Utterance Rewriting on Editing Operation
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance Augmentation
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
by: Sun, Haoxiang, et al.
Published: (2026)
by: Sun, Haoxiang, et al.
Published: (2026)
Similar Items
-
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024) -
VIT-Ped: Visionary Intention Transformer for Pedestrian Behavior Analysis
by: Elkammar, Aly R., et al.
Published: (2026) -
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
by: Karadag, Ozge Oztimur
Published: (2023) -
CrossVIT-augmented Geospatial-Intelligence Visualization System for Tracking Economic Development Dynamics
by: Bai, Yanbing, et al.
Published: (2024) -
SMT(LIA) Sampling with High Diversity
by: Lai, Yong, et al.
Published: (2025)