LookupViT: Compressing visual information to a limited number of tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Koner, Rajat, Jain, Gagan, Jain, Prateek, Tresp, Volker, Paul, Sujoy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025)
by: Goyal, Sahil, et al.
Published: (2025)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024)
by: Jain, Gagan, et al.
Published: (2024)
Provably Better Explanations with Optimized Aggregation of Feature Attributions
by: Decker, Thomas, et al.
Published: (2024)
by: Decker, Thomas, et al.
Published: (2024)
ViPRA: Video Prediction for Robot Actions
by: Routray, Sandeep, et al.
Published: (2025)
by: Routray, Sandeep, et al.
Published: (2025)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
Pixel Embedding: Fully Quantized Convolutional Neural Network with Differentiable Lookup Table
by: Tokunaga, Hiroyuki, et al.
Published: (2024)
by: Tokunaga, Hiroyuki, et al.
Published: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
by: Kuhar, Sachit, et al.
Published: (2023)
by: Kuhar, Sachit, et al.
Published: (2023)
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
by: Sun, Jian, et al.
Published: (2023)
by: Sun, Jian, et al.
Published: (2023)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
by: Chen, Shuo, et al.
Published: (2023)
by: Chen, Shuo, et al.
Published: (2023)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
Improved Algorithm for Deep Active Learning under Imbalance via Optimal Separation
by: Nuggehalli, Shyam, et al.
Published: (2023)
by: Nuggehalli, Shyam, et al.
Published: (2023)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)
by: Farid, Karim, et al.
Published: (2025)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
by: Shrivastava, Ayush, et al.
Published: (2026)
by: Shrivastava, Ayush, et al.
Published: (2026)
MTCNET: Multi-task Learning Paradigm for Crowd Count Estimation
by: Kumar, Abhay, et al.
Published: (2019)
by: Kumar, Abhay, et al.
Published: (2019)
Sub-token ViT Embedding via Stochastic Resonance Transformers
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Visual Question Decomposition on Multimodal Large Language Models
by: Zhang, Haowei, et al.
Published: (2024)
by: Zhang, Haowei, et al.
Published: (2024)
Causal Representation Learning with Observational Grouping for CXR Classification
by: Rasal, Rajat, et al.
Published: (2025)
by: Rasal, Rajat, et al.
Published: (2025)
LLM Augmented LLMs: Expanding Capabilities through Composition
by: Bansal, Rachit, et al.
Published: (2024)
by: Bansal, Rachit, et al.
Published: (2024)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
ScriptViT: Vision Transformer-Based Personalized Handwriting Generation
by: Acharya, Sajjan, et al.
Published: (2025)
by: Acharya, Sajjan, et al.
Published: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Multimodal Pragmatic Jailbreak on Text-to-image Models
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations
by: Ameperosa, Ezra, et al.
Published: (2024)
by: Ameperosa, Ezra, et al.
Published: (2024)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
FastCLIPstyler: Optimisation-free Text-based Image Style Transfer Using Style Representations
by: Suresh, Ananda Padhmanabhan, et al.
Published: (2022)
by: Suresh, Ananda Padhmanabhan, et al.
Published: (2022)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
TraSCE: Trajectory Steering for Concept Erasure
by: Jain, Anubhav, et al.
Published: (2024)
by: Jain, Anubhav, et al.
Published: (2024)
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
by: Jain, Anubhav, et al.
Published: (2024)
by: Jain, Anubhav, et al.
Published: (2024)
ViTime: Foundation Model for Time Series Forecasting Powered by Vision Intelligence
by: Yang, Luoxiao, et al.
Published: (2024)
by: Yang, Luoxiao, et al.
Published: (2024)
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
by: Li, Zhengang, et al.
Published: (2024)
by: Li, Zhengang, et al.
Published: (2024)
Similar Items
-
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025) -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024) -
Provably Better Explanations with Optimized Aggregation of Feature Attributions
by: Decker, Thomas, et al.
Published: (2024) -
ViPRA: Video Prediction for Robot Actions
by: Routray, Sandeep, et al.
Published: (2025) -
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)