Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Young, Shawn, Zeng, Xingyu, Xu, Lijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Chest X-ray Representation Learning via Semantic-Partitioned Contrastive Learning
von: Feng, Wangyu, et al.
Veröffentlicht: (2026)
von: Feng, Wangyu, et al.
Veröffentlicht: (2026)
XrayClaw: Cooperative-Competitive Multi-Agent Alignment for Trustworthy Chest X-ray Diagnosis
von: Young, Shawn, et al.
Veröffentlicht: (2026)
von: Young, Shawn, et al.
Veröffentlicht: (2026)
TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
von: Chen, Zhuo, et al.
Veröffentlicht: (2026)
von: Chen, Zhuo, et al.
Veröffentlicht: (2026)
Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
von: He, Landi, et al.
Veröffentlicht: (2026)
von: He, Landi, et al.
Veröffentlicht: (2026)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
The Model Knows Which Tokens Matter: Automatic Token Selection via Noise Gating
von: He, Landi, et al.
Veröffentlicht: (2026)
von: He, Landi, et al.
Veröffentlicht: (2026)
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
von: Guo, Yangyang, et al.
Veröffentlicht: (2023)
von: Guo, Yangyang, et al.
Veröffentlicht: (2023)
Multimodal Model for Computational Pathology:Representation Learning and Image Compression
von: Wu, Peihang, et al.
Veröffentlicht: (2026)
von: Wu, Peihang, et al.
Veröffentlicht: (2026)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
XNet v2: Fewer Limitations, Better Results and Greater Universality
von: Zhou, Yanfeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yanfeng, et al.
Veröffentlicht: (2024)
Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion
von: Li, Lijian
Veröffentlicht: (2025)
von: Li, Lijian
Veröffentlicht: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
ZeroSense:How Vision matters in Long Context Compression
von: Gao, Yonghan, et al.
Veröffentlicht: (2026)
von: Gao, Yonghan, et al.
Veröffentlicht: (2026)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
von: Li, Kevin Y., et al.
Veröffentlicht: (2024)
von: Li, Kevin Y., et al.
Veröffentlicht: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2026)
TokenSeg: Efficient 3D Medical Image Segmentation via Hierarchical Visual Token Compression
von: Zeng, Sen, et al.
Veröffentlicht: (2026)
von: Zeng, Sen, et al.
Veröffentlicht: (2026)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xinliang, et al.
Veröffentlicht: (2025)
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
On the Role of Discrete Tokenization in Visual Representation Learning
von: Du, Tianqi, et al.
Veröffentlicht: (2024)
von: Du, Tianqi, et al.
Veröffentlicht: (2024)
MST: Adaptive Multi-Scale Tokens Guided Interactive Segmentation
von: Xu, Long, et al.
Veröffentlicht: (2024)
von: Xu, Long, et al.
Veröffentlicht: (2024)
Scaling Language-Free Visual Representation Learning
von: Fan, David, et al.
Veröffentlicht: (2025)
von: Fan, David, et al.
Veröffentlicht: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
von: Gong, Chao, et al.
Veröffentlicht: (2025)
von: Gong, Chao, et al.
Veröffentlicht: (2025)
Soft Tail-dropping for Adaptive Visual Tokenization
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
von: Chen, Zeyuan, et al.
Veröffentlicht: (2026)
Replacement Learning: Training Neural Networks with Fewer Parameters
von: Zhang, Yuming, et al.
Veröffentlicht: (2026)
von: Zhang, Yuming, et al.
Veröffentlicht: (2026)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
Efficient Autoregressive Shape Generation via Octree-Based Adaptive Tokenization
von: Deng, Kangle, et al.
Veröffentlicht: (2025)
von: Deng, Kangle, et al.
Veröffentlicht: (2025)
ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration
von: Zhang, Leheng, et al.
Veröffentlicht: (2026)
von: Zhang, Leheng, et al.
Veröffentlicht: (2026)
RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
von: Zhang, Xiangjun, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangjun, et al.
Veröffentlicht: (2025)
AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
von: Chen, Zilin, et al.
Veröffentlicht: (2025)
von: Chen, Zilin, et al.
Veröffentlicht: (2025)
Efficient Prototype Consistency Learning in Medical Image Segmentation via Joint Uncertainty and Data Augmentation
von: Li, Lijian, et al.
Veröffentlicht: (2025)
von: Li, Lijian, et al.
Veröffentlicht: (2025)
TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
From Snapshots to Symphonies: The Evolution of Protein Prediction from Static Structures to Generative Dynamics and Multimodal Interactions
von: Chen, Jingzhi, et al.
Veröffentlicht: (2026)
von: Chen, Jingzhi, et al.
Veröffentlicht: (2026)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
von: Zeng, Weili, et al.
Veröffentlicht: (2025)
von: Zeng, Weili, et al.
Veröffentlicht: (2025)
Expansive Supervision for Neural Radiance Field
von: Zhang, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhang, Weixiang, et al.
Veröffentlicht: (2024)
Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies
von: Zhu, Yuehua, et al.
Veröffentlicht: (2020)
von: Zhu, Yuehua, et al.
Veröffentlicht: (2020)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Chest X-ray Representation Learning via Semantic-Partitioned Contrastive Learning
von: Feng, Wangyu, et al.
Veröffentlicht: (2026) -
XrayClaw: Cooperative-Competitive Multi-Agent Alignment for Trustworthy Chest X-ray Diagnosis
von: Young, Shawn, et al.
Veröffentlicht: (2026) -
TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
von: Chen, Zhuo, et al.
Veröffentlicht: (2026) -
Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
von: He, Landi, et al.
Veröffentlicht: (2026) -
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)