Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Quan, Xiao, Yanru, Shen, Yuhao, Wang, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
by: Qian, Jiahe, et al.
Published: (2025)
by: Qian, Jiahe, et al.
Published: (2025)
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models
by: Kojima, Yuto, et al.
Published: (2025)
by: Kojima, Yuto, et al.
Published: (2025)
ForgeryTTT: Zero-Shot Image Manipulation Localization with Test-Time Training
by: Liu, Weihuang, et al.
Published: (2024)
by: Liu, Weihuang, et al.
Published: (2024)
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
by: Xing, Bohao, et al.
Published: (2025)
by: Xing, Bohao, et al.
Published: (2025)
DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
by: Gu, Wenhao, et al.
Published: (2025)
by: Gu, Wenhao, et al.
Published: (2025)
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
by: Bi, Xiuli, et al.
Published: (2024)
by: Bi, Xiuli, et al.
Published: (2024)
TTT3R: 3D Reconstruction as Test-Time Training
by: Chen, Xingyu, et al.
Published: (2025)
by: Chen, Xingyu, et al.
Published: (2025)
Med-TTT: Vision Test-Time Training model for Medical Image Segmentation
by: Xu, Jiashu
Published: (2024)
by: Xu, Jiashu
Published: (2024)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
by: Kong, Quan, et al.
Published: (2026)
by: Kong, Quan, et al.
Published: (2026)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning
by: Peng, Hongxiang, et al.
Published: (2026)
by: Peng, Hongxiang, et al.
Published: (2026)
NC-TTT: A Noise Contrastive Approach for Test-Time Training
by: Osowiechi, David, et al.
Published: (2024)
by: Osowiechi, David, et al.
Published: (2024)
ReC-TTT: Contrastive Feature Reconstruction for Test-Time Training
by: Colussi, Marco, et al.
Published: (2024)
by: Colussi, Marco, et al.
Published: (2024)
SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object Detection
by: Yu, Zhenni, et al.
Published: (2025)
by: Yu, Zhenni, et al.
Published: (2025)
TTT-KD: Test-Time Training for 3D Semantic Segmentation through Knowledge Distillation from Foundation Models
by: Weijler, Lisa, et al.
Published: (2024)
by: Weijler, Lisa, et al.
Published: (2024)
MotionTTT: 2D Test-Time-Training Motion Estimation for 3D Motion Corrected MRI
by: Klug, Tobit, et al.
Published: (2024)
by: Klug, Tobit, et al.
Published: (2024)
Efficient Test-Time Prompt Tuning for Vision-Language Models
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
TTT-Unet: Enhancing U-Net with Test-Time Training Layers for Biomedical Image Segmentation
by: Zhou, Rong, et al.
Published: (2024)
by: Zhou, Rong, et al.
Published: (2024)
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
by: Feng, Ze, et al.
Published: (2025)
by: Feng, Ze, et al.
Published: (2025)
ViT$^3$: Unlocking Test-Time Training in Vision
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training
by: Wang, Mengmeng, et al.
Published: (2026)
by: Wang, Mengmeng, et al.
Published: (2026)
Test-Time Conditioning with Representation-Aligned Visual Features
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
by: Tian, Kaibin, et al.
Published: (2024)
by: Tian, Kaibin, et al.
Published: (2024)
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
by: Debnath, Soumyaratna, et al.
Published: (2026)
by: Debnath, Soumyaratna, et al.
Published: (2026)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
by: Yang, Yantai, et al.
Published: (2025)
by: Yang, Yantai, et al.
Published: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
by: Zang, Zehua, et al.
Published: (2026)
by: Zang, Zehua, et al.
Published: (2026)
Efficient Test-Time Adaptation of Vision-Language Models
by: Karmanov, Adilbek, et al.
Published: (2024)
by: Karmanov, Adilbek, et al.
Published: (2024)
GISE-TTT:A Framework for Global InformationSegmentation and Enhancement
by: Hao, Fenglei, et al.
Published: (2025)
by: Hao, Fenglei, et al.
Published: (2025)
Flatness Guided Test-Time Adaptation for Vision-Language Models
by: Li, Aodi, et al.
Published: (2025)
by: Li, Aodi, et al.
Published: (2025)
Test3R: Learning to Reconstruct 3D at Test Time
by: Yuan, Yuheng, et al.
Published: (2025)
by: Yuan, Yuheng, et al.
Published: (2025)
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
by: Su, Zhuo, et al.
Published: (2024)
by: Su, Zhuo, et al.
Published: (2024)
Multi-Grained Contrast for Data-Efficient Unsupervised Representation Learning
by: Shen, Chengchao, et al.
Published: (2024)
by: Shen, Chengchao, et al.
Published: (2024)
Efficient Visual Representation Learning with Heat Conduction Equation
by: Zhang, Zhemin, et al.
Published: (2024)
by: Zhang, Zhemin, et al.
Published: (2024)
EVLM: An Efficient Vision-Language Model for Visual Understanding
by: Chen, Kaibing, et al.
Published: (2024)
by: Chen, Kaibing, et al.
Published: (2024)
EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems
by: Ma, Xiao, et al.
Published: (2025)
by: Ma, Xiao, et al.
Published: (2025)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
Similar Items
-
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
by: Qian, Jiahe, et al.
Published: (2025) -
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models
by: Kojima, Yuto, et al.
Published: (2025) -
ForgeryTTT: Zero-Shot Image Manipulation Localization with Test-Time Training
by: Liu, Weihuang, et al.
Published: (2024) -
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
by: Xing, Bohao, et al.
Published: (2025) -
DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
by: Gu, Wenhao, et al.
Published: (2025)